Pith. sign in

Paper Citation Record · LEDGER

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.19595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19595 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:55.714499Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:52.829136Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:15:55.827429Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb7fc26a-ef85-4b17-9cd1-4f8198f10e5c · outbound

This paper cites an unresolved cited work.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:16:01.965263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:52.748091Z digest=sha256:41c35bba9fdbd1adda8f0ee60c3017defac22196ef6cbaad6ad33e2454cfcf97

Observation 7613c17a-498e-42a3-8d17-dd44120f1d7c · outbound

This paper cites Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:15:55.878979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:52.829136Z digest=sha256:b667fae48788e8e37ed96813df19508dc9296213cb9cd52bd57d23917df87305

Observation 01dc9593-8e57-4e1e-a048-8da73aba0b7c · outbound

This paper cites Dataset We train our model on the LibriTTS [31] dataset, a multi- speaker English corpus containing approximately 585 hours of read speech sampled at 24 kHz.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Dataset We train our model on the LibriTTS [31] dataset, a multi- speaker English corpus containing approximately 585 hours of read speech sampled at 24 kHz

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.813339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:52.919656Z digest=sha256:cc728b4332d28d4b0e19f3e67fb0a0173b043576cce67fe7bdf232b854f2e9b2

Observation 702f54ac-37a2-4adc-a8b5-3a746a42d4ac · outbound

This paper cites speech align.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment speech align

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.618499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.049253Z digest=sha256:c4760927da5ca713341be94c8345d6042917945e3965e9efaaf07f4da80f8126

Observation 0faadb96-b8ce-493a-b09d-7274d9b0a12a · outbound

This paper cites This ap- proach significantly accelerates model convergence while en- hancing the quality of generated speech.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment This ap- proach significantly accelerates model convergence while en- hancing the quality of generated speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.472750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.162304Z digest=sha256:5e488311b8b8586a1802945ceaa30e88c1f1b67c7ea05006343ea249161bbc77

Observation 98f04b34-3288-47ce-a147-dd04790540d3 · outbound

This paper cites In future work, we plan to further op- timize the framework by reducing model complexity and accel- erating inference.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment In future work, we plan to further op- timize the framework by reducing model complexity and accel- erating inference

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.307516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.252705Z digest=sha256:385c920cda3418b8e36a63985ffbd04a51b7c510868353b2652e44096338ec5c

Observation 258d8797-2481-4966-a231-dbe012515b96 · outbound

This paper cites Choi, J.-H.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Choi, J.-H

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:01.152594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.388149Z digest=sha256:c0e055b040b2ab4a2cda70da07cd7dc927d51b164201f37a1d228b33663f49a3

Observation 6cc1c89c-7940-405c-8b17-9a854be9677f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.471613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.471613Z digest=sha256:465fd8c8c502214b5d73431ce25406bbfed53bc3f06e86f08e1d6977e32191b7

Observation 1d4d6018-6c07-457a-8bcd-920d8e4b9f3c · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.546922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.546922Z digest=sha256:24270923d173580aa8325afe4731f974d12c8adf77c2b96053e863147cd15dc5

Observation d2e23561-cb4c-4429-b577-f92f37f4fae7 · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V oicecraft: Zero-shot speech editing and text-to-speech in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.980493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.654282Z digest=sha256:c8cf573c44cfd107450094a7b6ab44147d7387e5d590fd809c47e10d09dcb535

Observation e6714b00-7472-44e4-8dd7-2f684454f6c5 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.775073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.775073Z digest=sha256:a45be066720f44a1e5798ee93b33a5e980e7f7cf3626a69034b67fb866fc21a8

Observation 94639417-d75f-4193-b59b-dd326c03cc02 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Autoregressive Speech Synthesis without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.849366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.849366Z digest=sha256:a973205a9677b978df6ae03f7df3431d9f2e288b7e53143f59231860e284bcc7

Observation e41f997a-0485-4e23-98ed-f51d6bd54884 · outbound

This paper cites V oice- box: Text-guided multilingual universal speech generation at scale,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V oice- box: Text-guided multilingual universal speech generation at scale,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.771048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.908404Z digest=sha256:d6dedf0514ac966551e966afe26e759a7334fb8cf136a4be71c6997a7bfa7274

Observation 27bab5ff-4fe7-4713-8a79-6587a2a54bbc · outbound

This paper cites Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.611342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.958470Z digest=sha256:7d20053b70a67552c3f5aa17b54e24155fe34fc039e69653b07137fde2e8cc82

Observation 5c1f7394-f16b-432a-8e9c-fe282429a3cf · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.429665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.966196Z digest=sha256:5f8b2249fea7bc7741e19e26bbe90f1e7a6a436c409a30f8256415293b8ab39b

Observation 9596f8ff-8a8d-4e9a-8684-6b308bf5e031 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.976201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.976201Z digest=sha256:ab6b60861c10fd0b83c9b3deb402714ab6d696bf3f7eadda2182015b75c65dbe

Observation c507b2b2-efe6-46d7-8b42-232f6c1a8f0c · outbound

This paper cites Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.150992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:53.988216Z digest=sha256:f594eadbf1399f8645cb6f66cb28c7aa574be37f4426df507776eca8fd59417f

Observation 5f3b1eba-920e-48f8-b4c0-9c14da64bfbe · outbound

This paper cites Accelerating codec-based speech synthesis with multi- token prediction and speculative decoding,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating codec-based speech synthesis with multi- token prediction and speculative decoding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.913822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.007318Z digest=sha256:df34f22dab71dc287beeb30dc09a397814315f0ded20847c09eb1986e9239257

Observation 578d911b-13af-4ec7-899d-516c8e495232 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.026219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.026219Z digest=sha256:949f0b45b12226ffcc90056c3d9650b615c71f5f6d4fae6b1b8f8c774bef955e

Observation 37e878e4-81ed-4917-b258-be620ad63813 · outbound

This paper cites Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.607420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.054161Z digest=sha256:86b52fb66b9b4e706de6901ec3be8caea378e4f797dba48b52a11d1e5fba07bd

Observation e4889d69-d0b6-4e18-94b5-6db5a6357eba · outbound

This paper cites Flow matching for generative modeling,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Flow matching for generative modeling,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.088901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.088901Z digest=sha256:0103361e27cbe0cf18de15f3d1ce4d4973b3bef97da4f96d34cc93e117024888

Observation 8e022587-4e6e-4444-8ab2-63d0c33e41af · outbound

This paper cites Scalable diffusion models with transform- ers,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Scalable diffusion models with transform- ers,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.124824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.124824Z digest=sha256:ee02c043f68cbecea50116296a1b5676ec6b7261d42de1066b633c4a5d71b8b3

Observation bf7e6c63-0eae-4340-ae2d-04e1ecd0ec97 · outbound

This paper cites Efficient diffusion training via min-snr weighting strat- egy,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Efficient diffusion training via min-snr weighting strat- egy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.323953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.161703Z digest=sha256:62bff724212e412b0412795ce9a4fd59497651589275a7da6c1cd93d258e040c

Observation 7e709a39-c473-4410-b4e4-20bf79475049 · outbound

This paper cites Fasterdit: Towards faster diffusion transformers training without architecture modification,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Fasterdit: Towards faster diffusion transformers training without architecture modification,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.205251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.235356Z digest=sha256:270fdc8ed532a668d5e8f34b7ee598f9d80d37ad9b5c4d43c8631d1c54b7ce36

Observation e9e50506-c6bb-495b-98d0-e26ff9e6bf4d · outbound

This paper cites Representation alignment for generation: Training diffu- sion transformers is easier than you think,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Representation alignment for generation: Training diffu- sion transformers is easier than you think,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.124774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.329789Z digest=sha256:11feaf0da65d3d511f228b1fbb97e7022323707980ee6de537cd5f016c293c61

Observation c2174071-aecc-4622-8714-9a44830d98e4 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment U-net: Convolutional networks for biomedical image segmentation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.895075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.427902Z digest=sha256:dcd91bb1d1964c2175147812f7d9057b89e4cb0115420888ac047f430864f667

Observation 8d6618dd-846d-4943-9b46-f4f7c3b948a5 · outbound

This paper cites Attention is all you need,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.493555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.493555Z digest=sha256:77b76df62fbaddf248e57f3b704fd7f72f378c0556dc328518f9e2dedca962b1

Observation a3b9ff3e-2fcf-4bf6-b73e-6b2a7743199f · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.617541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.571933Z digest=sha256:15d380ff58ec4c337a058007c6c6d6ca310e0189401f2046de050e2a5cf54bee

Observation 9f3a332f-2942-4462-a4da-9d66d61fe7de · outbound

This paper cites Intermediate loss regularization for ctc- based speech recognition,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Intermediate loss regularization for ctc- based speech recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.310597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.646730Z digest=sha256:f5ac46c7be7807a8db9d00d398dcda11efdb01aaaa8d1616d25bd041724ee8db

Observation 64a576d3-3564-423d-a154-bf4847c85770 · outbound

This paper cites Audio-visual efficient conformer for robust speech recognition,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Audio-visual efficient conformer for robust speech recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.067644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.691350Z digest=sha256:f3177909eff4f29937fee2174b6f3945ee1dd7d4bd25b02d664848ef13e583aa

Observation 12af8411-e6e6-4a18-afc5-e109460d64a5 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.827126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.760242Z digest=sha256:862a958d454ea061bb65fc3480be5315168dba544c8478a33b30d7356ca53192

Observation dfce062a-4b05-4457-ad8e-f42ebb911f27 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.646947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.810329Z digest=sha256:712702230e7495f7f99de2cec6b192000acc5035416413cd59e3a65edca4bb78

Observation afad2a94-9925-4ab6-bb5f-c3a1503d52ce · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.497847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.866978Z digest=sha256:552325098e6661697f5cd6eb494908ab18603c14ba2638aab4f7963779cfa3c3

Observation 803692c3-015b-466b-8cbe-2b7465842a7f · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.264673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.956368Z digest=sha256:315f6df72ac51f8df954dde7c99aacd0be17bd6e69a316482ef602ac4de46b6b

Observation 9d8afcab-636c-4abe-8b1d-ef87767ad15e · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:57.123506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.056539Z digest=sha256:787f49b7ab99e597d2e92601ce19e7985c7275d07ee15e3af448e2342cac8918

Observation e3641322-96fe-4ba8-8d41-73fac6cef310 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.127465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.127465Z digest=sha256:0f42a71cb193b874083fb736d5640756dd1d65657df1ea06b941540234e4ba42

Observation 594d5d61-61ff-4221-becd-05815f436794 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.201763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.201763Z digest=sha256:655c616f5eeda352e40d235d52a3f72850283c0ebb09e9962847e912b893ffa7

Observation 35f5bfe1-7bfc-40ea-98c9-3231adacf766 · outbound

This paper cites Libritts: A corpus derived from librispeech for text- to-speech,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Libritts: A corpus derived from librispeech for text- to-speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.922439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.262774Z digest=sha256:49a1f5643f4e08a42405e74d0229b46e5670275bc58943663b91b316ef4ac3d0

Observation 80e28793-afc4-4762-8796-e84d592efcdd · outbound

This paper cites Lib- rispeech: An asr corpus based on public domain audio books,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Lib- rispeech: An asr corpus based on public domain audio books,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.795579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.319566Z digest=sha256:bb99aa83df276eb73bbd1d8e959a41b086f7219e3d7b00469a3a81517b13ae93

Observation 697a294c-cd83-46f6-bcda-2ebe01d793fb · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Convnext v2: Co-designing and scaling convnets with masked autoencoders,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.689787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.396629Z digest=sha256:fb666608ca94ccb91c751c4d8d892a80159a7f5871dfa6d0145a36633d260541

Observation 53c028ee-ac48-4ef9-be5f-9f625d05ba04 · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Decoupled weight decay regulariza- tion,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.469364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.469364Z digest=sha256:86a1a51da186b945e2b2f3ef104e794c0b78462eedbb9a3db08c4629b3c82aba

Observation f1a62001-924d-4777-95bb-7a89c40b6917 · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.516535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.566232Z digest=sha256:6d06808bb0ad5d774b11bc5659b2bb87e58663eb0447f7514dbd6f16fd79b229

Observation 03fb12a6-8584-435b-a51b-19c107223dbd · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.364968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.604506Z digest=sha256:055e675e11b003f8db28a53b412b2865c581b941d1b2347c5bbe85f303bc9d16

Observation f1d8e692-c73e-4f5f-b7e5-6a6bb1df7fa6 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Robust speech recognition via large-scale weak su- pervision,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.244826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.657348Z digest=sha256:62cca79a8b3f5b6c887b9c662b1ce5ac3b25b9db4ab045ef404f4240ebd0c070

Observation 9292437f-a354-4340-a186-9c6b2b4d6cf5 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022,.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Utmos: Utokyo-sarulab system for voicemos challenge 2022,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:56.033418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.714499Z digest=sha256:dd286aa2fe2b5166045ea10160bf956c644199dab02aabbfd68a976236f7db58

Pith citing papers

Observation 7613c17a-498e-42a3-8d17-dd44120f1d7c · inbound

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment cites this paper.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:15:55.878979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:52.829136Z digest=sha256:b667fae48788e8e37ed96813df19508dc9296213cb9cd52bd57d23917df87305