Pith. sign in

Paper Citation Record · LEDGER

FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 68 inbound Pith citation observations for arXiv:2006.04558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2006.04558 v8

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 68 of 68 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:26.891680Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T18:15:21.399873Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3ced301-f751-4fa9-a4b5-f08d66b409ff · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.432620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:5c4f0d8a2717df384ba8cc12e8c9d2669df2a697b0518f40d8d10b187a4954f9

Observation f9c429f0-e244-418f-b4c4-40b9260a9347 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.498563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:77a8c33b1d1368209bd6ddb27f04dfcc3e72407a9c18a36ca7af294ba1aff6af

Observation 7d176ddb-9b26-4ff6-935c-4a9d5e2baec5 · inbound

I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception cites this paper.

I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:37:25.329513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:37:25.329513Z digest=sha256:7b24d1d16664264a0482a039c404353cedb157426ed45d5d1fd25a36e70e27d5

Observation aab0afda-e84e-46d6-8d38-4f2dba8dd974 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.934932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.934932Z digest=sha256:9925861672eb0c4041c84a104ca87d93c54104c656feb388117c364d1804435c

Observation ff9c97c0-aaa2-497e-b7b9-dd602b6bb6bc · inbound

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model cites this paper.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.895990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.895990Z digest=sha256:7218b8c7e0424575e1674f682e957547ec33211e4c5745d3bcfeac232916ecde

Observation 2ec1ae0d-b63b-44fb-a6c7-f2372c92b755 · inbound

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration cites this paper.

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:26.922799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:26.922799Z digest=sha256:089e91c729958e9a6ceb0190be655e04ba49580e850517c7fe6807d94211d061

Observation 46ba7a1e-1aaf-4d62-9f4f-a52768b8babd · inbound

LatentSpeech: Latent Diffusion for Text-To-Speech Generation cites this paper.

LatentSpeech: Latent Diffusion for Text-To-Speech Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:16:12.922242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:16:12.922242Z digest=sha256:d6c42d334f484dacba37aeaf94720f25d6f62cc0d20adc57e2cda035b87017b3

Observation 4c17a08f-eb8d-4e28-a643-276f24aa57c7 · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:48.944939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:48.944939Z digest=sha256:f4ef1aec891f41aa6372f72a4183ec6863ad70fceb1ef812e458983b3b8aaa5a

Observation 2ebe90da-d97e-4d12-a128-be2ba4309274 · inbound

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data cites this paper.

Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:05:25.870697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:05:25.870697Z digest=sha256:e6c63460b29884e25a1de303fd04d61866d2a5ae5fb845bbfcaaa4bb30d00adc

Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · inbound

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction cites this paper.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.851676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.851676Z digest=sha256:9781e5aa26ed1b8ae1c1fe3a3c2d2cc6304aca0d2d3f5da98bc12ef783b70c45

Observation 68c217f8-b477-4d00-a6f6-1b230cea27e9 · inbound

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles cites this paper.

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:41:36.264122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:41:36.264122Z digest=sha256:cfbd846a441c82ce53cdf260ce8765651d4c4b089c2db8f1b0c88a6c229da288

Observation 304b3469-c214-49b8-b21e-bf8811bbf1c3 · inbound

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control cites this paper.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.949717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.949717Z digest=sha256:5fdd0af179a1562ab1743790acd6b47e781f8fdccafd981fde5483acffc0a43d

Observation 25f1d255-6d72-4087-87ff-f4a4ff22e5fa · inbound

A Non-autoregressive Model for Joint STT and TTS cites this paper.

A Non-autoregressive Model for Joint STT and TTS FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.530654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.530654Z digest=sha256:2b40717d60f0561c78f8d0e55c287ae4716d6c6be574cdaad84fca82ed04d381

Observation b3401b66-54ee-4800-a4b7-f3fed282a958 · inbound

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video cites this paper.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.265052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.265052Z digest=sha256:d79c81468a2d5cf8cbc2d42cf584c23fe82b1ea79c77329f90417644b99ec151

Observation 30595c40-9843-41ae-99a8-e0c97676db0f · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.853890Z digest=sha256:85860326a458e2dd031238030d648d8581da8b6114d18d57520a7fcdf3ba3b54

Observation cfe4fe30-e4b7-4e03-be0a-e17f861970d9 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.722514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.722514Z digest=sha256:783508d3aa426d998c27b9fcc74c0fafaea66b7fe03e3cf91e857e25c930360f

Observation c19b05be-bb84-4943-a068-019d9501fb15 · inbound

VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect cites this paper.

VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:34:12.985216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:34:12.985216Z digest=sha256:08bcea3a5df56b6c9630e4047a4085ccca25d76a71473055b3fec909b65ac4e2

Observation 275a9f07-8958-4fb7-8499-07f61f28cec4 · inbound

FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning cites this paper.

FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:26.891680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:26.891680Z digest=sha256:2743abbe954e75236158ca25e1a79a71f9da37e763f671e8f47548a6acddd753

Observation 4c31c211-b8d0-40f6-8ce2-44ce49450fd3 · inbound

Perceptual implications of automatic anonymization in pathological speech cites this paper.

Perceptual implications of automatic anonymization in pathological speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:01:54.538397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T17:58:15.380780Z digest=sha256:0d6d6ac342b7053b651df42ef4192bb9a31af7fb7d5550084e28e57648dc295f

Observation ea6bd5fc-2563-43c4-949b-1e27e10d4eeb · inbound

A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration cites this paper.

A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:41.105359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:41.105359Z digest=sha256:52dff6a7ab83d07dd0980025d48868cd10af47ee84e0581b60d01e959df2b8e1

Observation e6c22446-75de-408e-9bdf-9c1dddaee454 · inbound

Score-Based Training for Energy-Based TTS Models cites this paper.

Score-Based Training for Energy-Based TTS Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:15:05.415737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:15:05.415737Z digest=sha256:f5af3ff996fe6710767984ef45b2fe5748ed3f8de2f954cf89a0707c97e350fa

Observation b7501c1f-dd39-4816-9344-aa9aacd8daad · inbound

LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context cites this paper.

LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:16.526730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:16.526730Z digest=sha256:eb7f4e7c520f147293efe7434cc0041acdbb9a5e3e8cab4f41e597b28281a0c7

Observation d013a140-7618-4028-8f5d-28a9f72203de · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:38.162651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:38.162651Z digest=sha256:5e490257ef83a177209880065d6a763406f88702aa9956c84c7ceb4e3bee73dd

Observation bd0474c8-cad5-43b2-acc6-e2e05a9ca0d3 · inbound

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning cites this paper.

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:01.464367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:01.464367Z digest=sha256:dfd3a8d657a82b9b42d426d8c20f7276c6e0937a9762d3b69f052a84867a78c8

Observation 7fc88f8e-e561-4c13-89a6-35e5a654f82f · inbound

SpeakStream: Streaming Text-to-Speech with Interleaved Data cites this paper.

SpeakStream: Streaming Text-to-Speech with Interleaved Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.169600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.169600Z digest=sha256:a05de59992bda86774ee611615b625dd62c05674ff7c17d45fd46d9f8669e2a0

Observation bf85582c-7e86-4658-a5b5-f5e2cdf7c42d · inbound

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling cites this paper.

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:14.783720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:14.783720Z digest=sha256:801947b2e57f27dacf3b8ac571fe3f2dd852607b95cdcf4d7c0dc1adcd8cfabe

Observation fc3de9b6-9a65-4842-a3ef-4433dcddfc69 · inbound

Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models cites this paper.

Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:10.317617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:00:10.317617Z digest=sha256:dd08c7499093fc9e340aa1044ccecca75e1816f8612902f0a4ccdc7cb09ec43e

Observation e0cd1bd0-62d1-4a9c-8e8c-c581427ab987 · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.164512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.164512Z digest=sha256:3003480756a8169b34063122a43799ad804851e143a712f256a4aebb4c95026e

Observation 762799b2-c25c-48d3-be60-9f847c2a8a4d · inbound

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge cites this paper.

VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:53:42.127548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:53:42.127548Z digest=sha256:dbf42443b0f32370bcf9c090b0eaba92d1471ae7cd1a42c792b2f854212596a9

Observation d5e5367b-6e78-4ca4-b051-14f110120e45 · inbound

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems cites this paper.

InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T19:32:30.944300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:32:30.944300Z digest=sha256:938100b01408b2d2910da5e8fdbc75434eab036dbd6dd9b132d698fbc23e6eae

Observation 51258cdc-1a04-47bc-b851-77e66cfb4e41 · inbound

SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture cites this paper.

SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:28.366767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:28.366767Z digest=sha256:c1fe627f6983afbbef83b87b67ccfcf7ff8cadeb17286fd97f19d3dc93636e81

Observation b5600eec-68cb-4ea2-b308-f9cfa7d32d1f · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.001574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:7692cfe6fa9bcc13f370d37cb9d68491f35fa9a9fc0c667d2740123bdb5b0ae8

Observation 8341b7aa-fa69-4608-9424-d0d851b6f797 · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.831204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.831204Z digest=sha256:aa41e50c49565f590bb891747cbaa165ada1865dfff138fae6ab5444d7170d0f

Observation 5402b3f9-b82d-4aa5-aab3-8a802652788f · inbound

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning cites this paper.

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:35.420849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:35.420849Z digest=sha256:aa5c798d4ff2d1e74db05970a9d1125fd4861c78ae56a8a3b154e4db20d2ce2c

Observation 628de9bd-5f27-43bd-ab8d-750b22b24240 · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.287364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.287364Z digest=sha256:e05d0772962062789c438bb8be61371fec1772f83cc6bcd355e7517a4ea0b3f9

Observation 133feca1-a02c-4ef2-a45a-904b8aa3e946 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.194912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:9854c107e92db19ccc236a8bf35e75818156dddb56dd357208317ef7232abc79

Observation dbec8d13-d99a-4ef8-bbec-77b4844aaab6 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:30.485844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:30.485844Z digest=sha256:853ffe1fd1ae8c1ee473da2fbe68fdfb27c1b9e07a843eaeef3061a5c551ff79

Observation 2b9a4059-25c8-49d5-9df7-ee5e831806cc · inbound

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols cites this paper.

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 185

Resolution
unresolved
no resolver link, observed 2026-08-06T13:54:40.218211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:54:40.218211Z digest=sha256:de534b58c79cd34f04f2538175040f0b18cb6db3f96d9827a3c33a699cb876a3

Observation 10183caa-b808-4807-92b6-767070d9fb8e · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.523005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.523005Z digest=sha256:6194b2a09159592695ebcbddd33fefee4b57f385faafa6e99db4a2b2b50472ad

Observation 68004d0d-dddb-45cc-b8d5-53f7c625913f · inbound

Inference-time Scaling for Diffusion-based Audio Super-resolution cites this paper.

Inference-time Scaling for Diffusion-based Audio Super-resolution FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:27.864867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:05:27.864867Z digest=sha256:f20e4fd6633aa9111270c88e822977efe0a3e37b7e23af374aeeef431bb9a2cc

Observation 747a176c-8695-4ef4-a090-5eede55c3179 · inbound

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation cites this paper.

XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:17:38.247156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:17:38.247156Z digest=sha256:6f8b413b6cf75c3dcf3ec8b2bdde2f77c013b1eb95157483bde7a42197c7a748

Observation 7d0148a8-a446-4da9-8b35-6f0dd695f5e8 · inbound

SwiftF0: Fast and Accurate Monophonic Pitch Detection cites this paper.

SwiftF0: Fast and Accurate Monophonic Pitch Detection FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:31:05.578155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:31:05.578155Z digest=sha256:2c7f9d4c741e811caa30c14b37d9063eab6deeecfbb0355241ee53c50d7ae032

Observation 1257031f-77b7-42fc-a4b6-f15c4431988c · inbound

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model cites this paper.

MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:10.529671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:40:10.529671Z digest=sha256:b25e40581daa723b7d14ac6a9f168fb31ff161ba1d0b42f9c5341ae5848dc5c9

Observation e9e67638-e956-49a0-bc13-2c90fd26f737 · inbound

OLaPh: Optimal Language Phonemizer cites this paper.

OLaPh: Optimal Language Phonemizer FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.192861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T14:24:51.262622Z digest=sha256:c5189c696c112d9d07234c1ff83046f36cd3a477b3a3784099c23c881e4a22a9

Observation 80ba89bc-1bef-46a9-97a6-e5657b0b5713 · inbound

AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors cites this paper.

AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:51:37.338770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:51:37.338770Z digest=sha256:21bf4ec934c1413a1b50ba10ce8094655de51b438e5573b60cf7a426e18ba000

Observation 7dac97be-17ba-447e-ad96-257d4e781f18 · inbound

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion cites this paper.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.435469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.435469Z digest=sha256:d1e9f6d58193ecaf14fa50786d048387f2db6675ff6a983ac00cf38c028f085e

Observation e2bb892b-6864-48fd-8c99-9ea5d662fef8 · inbound

Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects cites this paper.

Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:30.232674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T08:11:31.181704Z digest=sha256:05977c5e181a964e9d8a655613251622047bcd6a43a4e8b95da59484cda7fbe8

Observation 795dc8f7-e99f-4731-a327-18aba2d1db48 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:16.003779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:16.003779Z digest=sha256:10a82aa97953acf67d0fe0ddfcc0a89293a641a6d9c7d455588312775ffe3612

Observation 759b864d-0c3b-408b-9987-b89f0e364d76 · inbound

MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline cites this paper.

MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:18.020832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T20:17:45.920234Z digest=sha256:d2cf12a7d44bd24078bd84192976fc167cbd36c8a8699ff583b311dde9b20618

Observation 8d0d5933-a97a-415d-99c4-b9cfd8414a70 · inbound

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models cites this paper.

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T18:35:56.164869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:35:56.164869Z digest=sha256:43802a75d6695dd199f3f5e289e2ccf35eb1535ead1927ec4f2e3fef32e74f62

Observation 6c288869-8085-4a6a-8da6-d7d552cceac1 · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.252636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T07:59:34.622096Z digest=sha256:dcdf46590ba21992ef0f60dc881c7ce4f18e926ff191d64e624b547427086965

Observation 9bf430db-384e-425a-a788-69cf379c3eb3 · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T05:29:35.464288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:29:35.464288Z digest=sha256:62e9ef0375bd615c8ef578e9fc58d57c2a7552a914ce475a5635c9ed59ab9c42

Observation 8c7c3b5d-04c4-40d9-9054-e0883607fd77 · inbound

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment cites this paper.

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.304537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T01:07:50.243903Z digest=sha256:5e7f623745214e250914d9567bd5ae762096b2958f5d4d6921a6dc03644baea2

Observation 7d3b91d2-5017-4284-a5a8-6182bb24e22d · inbound

MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech cites this paper.

MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:20:59.810611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:05:58.034000Z digest=sha256:37633d46bc979d13b5e82376991a8e89cd637b8cafa1abdc0095581289527d2f

Observation 82fed65f-3184-4ec3-8c2e-3ca5adf453b1 · inbound

HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation cites this paper.

HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:28.158629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T04:11:29.782569Z digest=sha256:6b3ca7b0f1d37a2f04c11c5473a8d3682ceae39ccde4191edf5ab3d1af47eab7

Observation d04f7cec-2598-4d24-b07d-362ee582d878 · inbound

N\"ushuVoice: Reviving the Voice of Endangered N\"ushu with Pitch-Aware Text-to-Speech cites this paper.

N\"ushuVoice: Reviving the Voice of Endangered N\"ushu with Pitch-Aware Text-to-Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:30.978541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T16:39:49.263497Z digest=sha256:ab237fb971244e628acb7b2ee38c1e3748934f3ae7d83d60a261df87ae2b204b

Observation abb0c836-e6d8-4c3e-b29c-9ba1b2ef9c9f · inbound

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech cites this paper.

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.574696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T19:09:33.605232Z digest=sha256:fd484e9769a3b1e4fcb80bdc55aadc5549ae15baca781bfbe9a0d3c0ee5e0fdf

Observation 8a76256e-c556-48ec-b6c9-e65bce9e280c · inbound

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation cites this paper.

DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:38.611490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:53:34.428544Z digest=sha256:200c0a154f1d07ecad96f10ad90de8d36af61d37bdba115f9f8adbb39e5d41e7

Observation f4de4787-a61c-49b9-8988-98a03e38c516 · inbound

Learning to Evade: Adaptive Attacks on Audio Watermarking cites this paper.

Learning to Evade: Adaptive Attacks on Audio Watermarking FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:19:43.962125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T10:10:18.351233Z digest=sha256:981ab8c54dcdf47894694c1c12844e93a00d6cb66e17e025853b35dec65dad3e

Observation 6f0c767e-c6a4-4b61-86fb-329d8473eeb5 · inbound

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling cites this paper.

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:24:12.484621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T03:20:05.125635Z digest=sha256:b17150add17aeb2e1f33add5c718c47f3b191d7b575991cbaddbf0a85a3175f3

Observation 401d0283-227c-44fd-924b-d3115d05a1fd · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.203091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:81bfdaf6fc96ac056d61d4374d1075be5315dbd26978b889f78e58b77b65ebf1

Observation 6fa43e09-9a85-4bbf-9982-bcb5baabf37d · inbound

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech cites this paper.

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T18:15:21.402116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T18:09:22.207379Z digest=sha256:737055120f2b4fe72de3bbb1696f1fe82553805d199218addf9969fd3bca8273

Observation 8b525661-3042-45e4-ad7a-d4be247ee97a · inbound

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling cites this paper.

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T06:27:06.683926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:27:06.683926Z digest=sha256:15460504ee3c585e700d30887dfc6b0c567ffc4724b16281f683d8e505694b2b

Observation a0d3e1bf-75f0-4f90-b74b-e46458c72f28 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.513735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.513735Z digest=sha256:0c67d94a4e21afb90dd2715a82f4e7f920e16c5c0fc4b490d2cc84e67f856eb8

Observation d0c683b4-48b1-4d44-8145-b7b42e8bd38a · inbound

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion cites this paper.

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:57:40.230200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:57:40.230200Z digest=sha256:d3f46507e8067a947467feb55caca82721eeb9244933819db9710547d0a740ba

Observation b241bb6a-636c-4942-a9b6-e30e605fe28d · inbound

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis cites this paper.

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T15:08:48.177277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:08:48.177277Z digest=sha256:4925fd287072661065e5e4e1e1967e9cfcc52bb708e30572f3b30acba0858269

Observation b272a6a6-74fa-42e1-a810-a34233ee9888 · inbound

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces cites this paper.

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:33.129335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:33.129335Z digest=sha256:d27bc328e0d3c9e350db34a347bea3779ea43a01871cdcec989b3921f824a219

Observation b3081ec9-7616-4677-9e7a-06de3c16e19b · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:42.145891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:42.145891Z digest=sha256:981ff3ecf0461f26e0ef98f3dedcafc777658fe7ec4fbd004ba03a989131bd52