Pith. sign in

Paper Citation Record · LEDGER

MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2502.18924.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18924 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:30:34.140760Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:50:11.241591Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cdbe0ec9-7038-4163-a31d-7bc3d19ebd5f · inbound

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis cites this paper.

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:34.140760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:34.140760Z digest=sha256:d09dcb3a96dab4ccd4922bf8381514b7177299690e24c010a5b1cfebc4ebe6fc

Observation 587953c3-962d-4b96-8ae5-507bfbef7791 · inbound

ACE-Step: A Step Towards Music Generation Foundation Model cites this paper.

ACE-Step: A Step Towards Music Generation Foundation Model MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.366580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.366580Z digest=sha256:898f4a2437aa780d11af2671420683c50cfc9b0ab299eaece64c461249b4cfd1

Observation 099d9eb2-9a58-43c6-b9b0-a7be49a96e24 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:55.278735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:55.278735Z digest=sha256:0ab648c8a52f241f30ed95640084a5fb3bac0c7397d7ceadd207a3ea8df58775

Observation f6bd46b6-78ae-4182-87a8-0c239981490a · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.053768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:ee76640419c4b2caa8853c4cbb3da04f3a0314567cc3585f386405847b2b6afa

Observation 03cecab6-b8f0-4506-8765-6f7727eddc1e · inbound

The Thin Line Between Comprehension and Persuasion in LLMs cites this paper.

The Thin Line Between Comprehension and Persuasion in LLMs MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:07:07.323651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T06:05:44.600781Z digest=sha256:c2b65c3a20edba576bfdddf1fa702e30dcd522c4ada107d610ebef92b40ca8c8

Observation a800ec70-a113-4f3c-8d23-8a6a1ba1b0f0 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:10.849757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:10.849757Z digest=sha256:883c2fbb9b816fbe0a080b344220523f57089e8e3894187335b2f1e79e711951

Observation 693a491d-ead1-4089-a240-50699ac46ec3 · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.524679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.524679Z digest=sha256:253ee5b99d2bd2899751c1838cec0ce499be90ea682a13497078bf94d5344031

Observation c783b992-901b-485c-ae29-8d885f5d5cd6 · inbound

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System cites this paper.

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-Text System MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:56.269739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:21:56.269739Z digest=sha256:157db84ca9d3822fcd783cf122db38e5141b7c85ca1b0a67bbfecab78d0d4a60

Observation 7dd14df9-6bc9-44e4-a47a-e65792cea62a · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.530961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.530961Z digest=sha256:2a171208cc89e027b054d2ae661617137ece8bb6494c66437b5c2f81d5de3bcf

Observation e4ed716a-13e0-4bc3-a677-661395c9a330 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.995304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.995304Z digest=sha256:aeb75dc462c3eea7d318b6c7873b6ddba7264527d3ecce961c2b3457a537d17f

Observation 9625e5cf-2220-4ca0-8e39-cb1d27fd56b5 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.776569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:bd8862d5aee9e7580f6447795f4ce5dee0db1f32c40255fb08045f4577aa2414

Observation 1662e086-ac3a-4e49-a673-016f13136596 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.172616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:0b46968b6f340fdc87fabdf862319b67a503480460ae941caf3869d9d4580870

Observation a4c39120-64ca-43b2-befd-ce6d88be8442 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:5463126031f9bb59b169671a8a9714ac617b58ded427323dd9d340ebea6f2948

Observation 04768d21-8e5a-4997-9c49-9a26dedad9ca · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:51:00.828282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:1edbb8c45e718cd5adb33ba479235a7dcc250a6d2c3c5563883b3e5741954cd1

Observation 4099f967-69ac-4231-b82a-512ba525982c · inbound

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction cites this paper.

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:59:50.716425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T07:59:22.416259Z digest=sha256:d86bd2d1268d19a0a2adaafd0f8efbf5e826d22d6c2fc576b7882e2386571ebb

Observation 2590a872-a3f4-425c-86d7-f6e915558813 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:00:58.998461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:532dc0c6ddb2d21dbf5eb52d4b6308849b798958fa674aeb3983d40db5a6823c

Observation 3ad9fbba-a588-4172-919f-eb6559d93ab0 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.318898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.318898Z digest=sha256:19bc294c2d25f8f4fbb1b61bc0b4d37cdacc7c702a3f9408583273680700570b

Observation f700083f-6c0b-4c7d-afe0-7f7c12ce8eba · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.109018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:bcab46188aca116c617c69140f451018b411fd86ffb2df6318d329859c665bc9

Observation 768e4f15-4a57-45e2-adb1-22e3139c3de8 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.706809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:bbbf8589e90081031f30423accff388ec39bb2164c33cbd4d531e096102e3bb6

Observation 111fa7d6-a449-4200-a13e-c0646cbd6737 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.748586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:20a0ec73e8e58e8e000b3b7fc732265cff5d7ea666a7753fc1f7fdb86d200131

Observation d62cfc91-7398-45ee-acda-3dc9648f1de1 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.051412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:9fbd3320203cc814f2339eca00e874d0cd2a1d7ec68bd0f275579078461c5135

Observation 6ee9bf96-8091-4db3-a043-d2fa1fc3cf9b · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.929572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:87f495c5626597feeec8332fdacee2ea040b50cf277f8e1d7ff77970f34360d1

Observation 2bcc5c2b-2ddf-4239-a728-12ede820e97e · inbound

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors cites this paper.

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.876078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T19:04:25.542629Z digest=sha256:70038ba3163e5e9addafbb479f6ae0fc937438a0d1662aeeb2a785e9c73da89e

Observation 9ab8c22f-96c7-43be-8eaa-9bd5e66b0d74 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.566017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:5773f4aac435702508abd91ae9ef1b055b47798cdd5e6309d8086f75fbd722c6

Observation 910e558f-fe07-4e39-b823-61df4e107646 · inbound

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS cites this paper.

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:11.243408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:35:52.631543Z digest=sha256:bba368b957c81e69ff0202d3785b3c1cc95e5ad178542a0676b2bcb98a2f05d9

Observation 763bbe90-8755-42a1-9e6e-4e16b0463491 · inbound

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS cites this paper.

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:40.076521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:07:48.053378Z digest=sha256:01e7d5090cbe71d89c9aca4499a38b22e1e7db7be32c1e796e9a78ec0a381d65

Observation 4d72292d-bc6e-491b-a2d7-affa35327ce3 · inbound

DETECT-3B-Omni is Agnostic of Content and Demographics cites this paper.

DETECT-3B-Omni is Agnostic of Content and Demographics MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T02:38:17.163398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:38:17.163398Z digest=sha256:d9abfd7b214311a3f8bcf123bb9b193b349cc789f187c9ba3f3e2965120656af

Observation 966fa3d7-be6a-434e-9f6b-365d6f941014 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:932ca9acee41aba7b005292c5cf5b8773dd33e3dbc464c98fdf6af5d2fe9bec8

Observation 1562ad06-8805-48b2-a8ad-4d7b16698ae6 · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:36.903349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:36.903349Z digest=sha256:c65774ebbd728a5fa02502d9a461bdee49fa4aceb719ea0efcd8210b62512f6e

Observation 5f8c3cba-3765-47a0-b2de-46647600cdaf · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:23.503407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:23.503407Z digest=sha256:2a12dd2d26600d2b478cafe4ba59b34a90190ad1a62edab973ba85824b28a76a

Observation bf3d068e-ceb0-4675-abbf-063188f005fa · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:44.716159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:44.716159Z digest=sha256:65fc803f41ab4a5b8559d90627bd08458ee517f8022697fb0f6c66c5b4b7ba36