Pith. sign in

Paper Citation Record · LEDGER

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

As of 19 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2508.11326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11326 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:04:51.686582Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:28:18.202974Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9738eb00-04f2-48bf-92c0-b27cd17b3e05 · outbound

This paper cites URL https://elevenlabs.io/.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://elevenlabs.io/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.748254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:47.469156Z digest=sha256:8c939ff10c78de2b9641c9d5387c91473452076b46a5df58eaa1c5d311a3bec8

Observation ce9e6f66-ac65-4b5c-9169-07ff32a861b1 · outbound

This paper cites URL https://www.minimaxi.com/.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://www.minimaxi.com/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.589041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:47.565640Z digest=sha256:6ab63b3157aa3ae35da103f7bcd13ce3364b4ec12ca2ff334188a5cd72bb1f3b

Observation 89f4c621-ff5f-4dcc-822d-b331d18b5808 · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Prompttts: Controllable text-to-speech with text descriptions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.428810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:47.689717Z digest=sha256:fcd87a755450c7b182774abf858464ec2b2dae0921d03b2bf4569af3ce5c3ecf

Observation 8c8023d0-3447-4634-b9fe-92b3eadf8ecc · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Textrolspeech: A text style control speech corpus with codec language text-to-speech models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.027737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.027737Z digest=sha256:98ec1a6b42c020db37126726ea8cb9f91fb0cb282d9d372536faf4d608f35bfa

Observation e089af01-3335-4723-9b80-2e7ee8217d75 · outbound

This paper cites Prompttts 2: Describing and generating voices with text prompt.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Prompttts 2: Describing and generating voices with text prompt

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.245317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:48.124723Z digest=sha256:98f138c6c4e4d8485e6fc2be3bf34410fc1b332bd4bb0d8f22a5240373f43956

Observation 1bdee4dd-a422-45a5-97ce-36ccaa271d3e · outbound

This paper cites Libritts-p: A corpus with speaking style and speaker identity prompts for text- to-speech and style captioning.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Libritts-p: A corpus with speaking style and speaker identity prompts for text- to-speech and style captioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.227297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.227297Z digest=sha256:8d9d41bfcb6437271dff4bfe38b153863475e336390e1421c2b46b2326b36880

Observation 7fc63fbb-c694-47ce-96fa-7e5e5aa45525 · outbound

This paper cites Speechcraft: A fine-grained expressive speech dataset with natural language description.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Speechcraft: A fine-grained expressive speech dataset with natural language description

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.356380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.356380Z digest=sha256:8da326d72b92fc0edd7bb7c36373a3f42c62e32a46ca81b94a06929aa733fbd2

Observation 604d43fb-5ff3-4590-b857-a3b1e3df899f · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.453562Z digest=sha256:1b6b5bf45bd6f5e4783292762a9f957228d9941b3b2a9139dbabd62503e8c370

Observation ad880798-f23f-401c-86a7-7da9281cad6e · outbound

This paper cites Scaling rich style-prompted text-to-speech datasets.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Scaling rich style-prompted text-to-speech datasets

Reference 10

Resolution
verified exact
doi, observed 2026-08-05T20:04:52.561894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:48.511222Z digest=sha256:8afb53cd63119fb9ef7731426259ac36bad51aee6e0955c09c1c0f2421504122

Observation da4489d1-bfb3-4e5a-88bc-db1fc786b8b6 · outbound

This paper cites GPT-4 Technical Report.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts GPT-4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.592233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.592233Z digest=sha256:f87969c785fb81e6412b72132e4a84863ac0fcef21ba2f818978774770ddb529

Observation b68494d3-07d7-4dba-a3df-bf733c33b7f2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.684230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.684230Z digest=sha256:c8fac8e20802eee0421ed8ac023521a0afb2e0fb0135d035c0d53ea1a4a24980

Observation f0be7c83-54b5-4e2c-b38c-2317bd581025 · outbound

This paper cites an unresolved cited work.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:04:55.112183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:48.772419Z digest=sha256:09f0152e74ee0614a5f946d3e82c80f3a44b446e9560362aba8364833b8d16af

Observation 971ec484-d792-4683-9284-c4dd997917f3 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.852163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.852163Z digest=sha256:aca69418077317fed6176df7d40a0fe79e1cbb28cace0c08eb638fc4fab32568

Observation e8030ec9-5c4e-4c8b-b9de-b51f1d09328a · outbound

This paper cites EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.933028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.933028Z digest=sha256:4eb51365e23ce086324ff26c7846594579d09da987d8afece91f6635cc305648

Observation cf596226-53b9-4921-ab0f-3282f7c3cfe8 · outbound

This paper cites Qwen3 Technical Report.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.009853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.009853Z digest=sha256:4cb09d4f04689c8c4fb91266b941a68b22a91a76a16f3361df115485c9559477

Observation bf9a4a57-a755-4743-aae5-8ba47bbd97e6 · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.104039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.104039Z digest=sha256:7a969c0f357e746cbad6ecc160c440715cade0b7ed54bbb2f33c16dc5754660f

Observation d8fb27d0-048f-43b8-8811-f6ae30b3d965 · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.957933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:49.260735Z digest=sha256:7871a284792a777d6d8edc59339085b8ee26589a9b8f0ea275a8bac614776aa4

Observation 8b2ff2d1-6018-4309-95ab-bf44defa8f83 · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.161915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.161915Z digest=sha256:11a6fb3c51fd341094e335ee81f35d4a9f440ac1e84c8a8db9564f8027575c0a

Observation 43dfd901-5a48-435a-97c5-6aa2d561ba95 · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.706910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:49.447667Z digest=sha256:6f73e97547b815e489666895b561eb632476fb7e4c0c74d3afa4245a6e66667e

Observation 6313292f-a74b-4422-952c-043a5e25126a · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.342135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.342135Z digest=sha256:c5702bbf68507aedcef16f61bf71eb43d18428b38ff2c9792f29aba4f287a6c2

Observation b0af524c-c2bd-43bc-b1e0-e4d2c969f240 · outbound

This paper cites Simbert: Integrating retrieval and generation into bert.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Simbert: Integrating retrieval and generation into bert

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.508410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:49.727665Z digest=sha256:a61a7466e0597178c754636454f78527fabfe4d7c37193c326114f81e96ed8ed

Observation 813005c5-8fe4-4834-b72a-2f5f23fc6ff8 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.511799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.511799Z digest=sha256:56f49e5e9f085f8be6fb6434cea4281b56a1935a88f66c407522d2c9f91ca620

Observation 695d6b7c-a62b-4a93-bf27-bbd25450c258 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Fastspeech: Fast, robust and controllable text to speech

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.168169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:49.987765Z digest=sha256:1fee3ef98618ccd479aef1a2e721e8bb9d6790dc594a5cffe2e26e0e5d84d373

Observation faa15446-d0e2-479a-84e9-03b0b46f337b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.064063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.064063Z digest=sha256:427c1d4fb655164654ee288498c35a529908d3f9da93bcd9107e4742f281957a

Observation ac9cca37-f696-416f-8db7-12761b0f7f4a · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts BERT: pre-training of deep bidirectional transformers for language understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.327115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:49.809205Z digest=sha256:b03c64d68698655524a4979c948a0011006f4efe191e096f390feb71ee3d1f47

Observation 2c2df928-bc5f-48f1-b625-13fb9153f588 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.260659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.260659Z digest=sha256:6dd6197fd56297ab08deceb7ebf20a5352ec47c8a10b79da99d9650e4e218b93

Observation 88f43fff-0582-44ed-bff3-0f42b3fb910a · outbound

This paper cites Scaling vision-language models with sparse mixture of experts.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Scaling vision-language models with sparse mixture of experts

Reference 28

Resolution
verified exact
doi, observed 2026-08-05T20:04:52.080129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:50.387922Z digest=sha256:513960c9441f97a79e107a2ce89a6798cc39007ca8d9861d6cbe52d1d5e4df45

Observation dae82b05-3ee8-421c-9566-a108eea12dd3 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.463688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.463688Z digest=sha256:1ea3e67e02cc467800e5c096ba0381cc6038160048f457e55efd416faaff9229

Observation 653e739d-550c-4d36-abff-8a836d6e1634 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.143999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.143999Z digest=sha256:74ad3c8ff9db974f32a8f2d57aba7d523e34563030b7b6213fb4a7f7a463e670

Observation 2bc8c221-afe1-4c67-97d8-90be3bb98c38 · outbound

This paper cites The Llama 3 Herd of Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.680474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.680474Z digest=sha256:180647c6009610512296b8e43f4bc7d7610d387d3f1f71dc498ace797575cd98

Observation d17d0c2a-3b7c-4648-827f-3ffda6d2e73f · outbound

This paper cites Orpheus tts: Towards human-sounding tts.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Orpheus tts: Towards human-sounding tts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.997908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:50.781373Z digest=sha256:df7904b7d683709aaa3313696874bd642476b0d57d4bd675c3cd23c4bfbcbd29

Observation fe17e825-4a1a-4234-a1b4-7f458d45e298 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.870799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:50.872744Z digest=sha256:eaec6c716b744d44931342c0843714d37c3aafdc6c4119976da068f4084e2cae

Observation cf0c18c0-5f8a-4972-b220-3ebce6c1707b · outbound

This paper cites Qwen2.5-1M Technical Report.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Qwen2.5-1M Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.570486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.570486Z digest=sha256:cd3edeafe6be63fc7b3fe29f299cac38924ae2e960b05fdf76c7cd09fb944382

Observation 66e3afe6-3357-4c14-8fde-7939efa6c934 · outbound

This paper cites Neural discrete representation learning.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Neural discrete representation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.559273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:51.113315Z digest=sha256:cb75d8a5dfed2a4cf7a321d3583acc80c37d3b0ca3a5f6c93cb3af65c4989c7f

Observation dc0ecd74-90e6-4083-b7bc-ac851aba7832 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Soundstream: An end-to-end neural audio codec

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.210481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.210481Z digest=sha256:4da4dc329b845a9ccbc0386d9f877ff2ae51bfd8e964c9d396c3dc703633c523

Observation 4ed51057-2fdc-4987-a663-7e456b2448e6 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts High-fidelity audio compression with improved RVQGAN

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.446296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:51.289926Z digest=sha256:e9f66cca859bd84cc3045141cf04e7490c916a1a67609a7b19a135ce8c59e720

Observation 9accb710-3a08-439c-b5ad-e7e7fbbba6c0 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.369330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.369330Z digest=sha256:cebbbd214f036147fc36e1c5cc8e27f6a76dd7a356e9aa9bee414038b78affd8

Observation bdc1c6f7-2d9a-4f7d-9bd7-3bf548064f4d · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Self-supervised learning with random-projection quantizer for speech recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.769779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:51.037030Z digest=sha256:e4a519d44ea8d63243b42d7395f843a41e027de72b41ce13f0b4dd9c41045d0c

Observation ecd9b8f4-018b-449c-80fd-7bc44ad28dfb · outbound

This paper cites Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.505861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.505861Z digest=sha256:c376ad24d94c4d3472e9110f96108f8d575ed3a3033bacdf9e6c99b6081cfd61

Observation 4dc488d9-eef8-4a29-9e0e-fef17261ee57 · outbound

This paper cites Albergo, Nicholas M.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Albergo, Nicholas M

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.590164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.590164Z digest=sha256:6ee4caf1292bead547317ea5785e369f6942535081ac5636178a9b4b215880bc

Observation a6b3cc7b-26e8-46e9-9a56-07636ec06f1b · outbound

This paper cites Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.686582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.686582Z digest=sha256:52b9a28093a6bc588f93e365790afdcd6e02d491b83db6c2fd763bf84aba39a0

Observation be96fe05-0c7d-4492-9fbe-fc61f0d13e93 · outbound

This paper cites Elucidating the de- sign space of diffusion-based generative models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Elucidating the de- sign space of diffusion-based generative models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.311325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T20:04:51.448310Z digest=sha256:aad3393a74987ca995126fb6ce863aa7b93b2d37941ef31ecdf8b15b7fc31ec0

Observation 8854b402-7093-4e14-9a29-82a68a516603 · outbound

This paper cites URL https://doi.org/10.1109/TASLP.2021.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://doi.org/10.1109/TASLP.2021

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.973979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.973979Z digest=sha256:0817b06c9bbb1677c17d593aaef2bbcd4d43cbb48f94702cc464aac0f1550ee2

Observation 5d35de97-2743-4564-959d-2c91eed98a3a · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.612420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.612420Z digest=sha256:7011834f3d3447b7a0b80823ab440e4df51e15142bc1663fbda7e1cd6c657a31

Observation 1a804652-6dfb-4a4a-9f30-ad5deddc4abd · outbound

This paper cites URL https://doi.org/10.1109/ ICASSP49357.2023.10096285.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://doi.org/10.1109/ ICASSP49357.2023.10096285

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:47.813576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:47.813576Z digest=sha256:d7c8b095e58a645a0f4b092480cab5665716ab95dbb48f5f6c7a95cc3b546b1d

Observation e5145ede-99f9-4cef-bc1e-8e700f33f581 · outbound

This paper cites doi: 10.18653/V1/N19-1423.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts doi: 10.18653/V1/N19-1423

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.884453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.884453Z digest=sha256:958460877f71613d0e67f5fe9f926d277fc809c3e67d117b5e8a688778cc4497

Pith citing papers

Observation 5c51c5e4-1523-44d3-a286-74cf8ac7a6d0 · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.871334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:8b0487582d7c4abe2be59e47403e6ea17de663241c57f8dc9acbc8ad4eccc1c9