Pith. sign in

Paper Citation Record · LEDGER

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

As of 7 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2508.11326.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11326 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:04:51.686582Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:28:18.202974Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9738eb00-04f2-48bf-92c0-b27cd17b3e05 · outbound

This paper cites URL https://elevenlabs.io/.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://elevenlabs.io/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.748254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:47.469156Z digest=sha256:5dfb4a477103e272d80660cf6181106a13e93ffececedf4a9ba8dadf4650a5f9

Observation ce9e6f66-ac65-4b5c-9169-07ff32a861b1 · outbound

This paper cites URL https://www.minimaxi.com/.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://www.minimaxi.com/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.589041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:47.565640Z digest=sha256:ef877df2599284ee8d46698ab8c68257ca02a7ffdf213b94829e7ee1cc1ac805

Observation 89f4c621-ff5f-4dcc-822d-b331d18b5808 · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Prompttts: Controllable text-to-speech with text descriptions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.428810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:47.689717Z digest=sha256:35489ed44ef55b6aaab12a4577fa39aabd588da9f4a884f5d391baa494a564cd

Observation 8c8023d0-3447-4634-b9fe-92b3eadf8ecc · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Textrolspeech: A text style control speech corpus with codec language text-to-speech models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.027737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.027737Z digest=sha256:f4aafea6d42034199c29cf99adcd8a89e89e99f6de7b12aeb7af1b82afe00405

Observation e089af01-3335-4723-9b80-2e7ee8217d75 · outbound

This paper cites Prompttts 2: Describing and generating voices with text prompt.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Prompttts 2: Describing and generating voices with text prompt

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:55.245317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:48.124723Z digest=sha256:59b5b6663913b5f48aa87971f3e921e350b07ff297ac05f8359b1c040d9f8e58

Observation 1bdee4dd-a422-45a5-97ce-36ccaa271d3e · outbound

This paper cites Libritts-p: A corpus with speaking style and speaker identity prompts for text- to-speech and style captioning.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Libritts-p: A corpus with speaking style and speaker identity prompts for text- to-speech and style captioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.227297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.227297Z digest=sha256:e1c3377b7940b5490f8e313860c8d638ad7b8cd37b8f5c57c7f4b6e4d95e61f9

Observation 7fc63fbb-c694-47ce-96fa-7e5e5aa45525 · outbound

This paper cites Speechcraft: A fine-grained expressive speech dataset with natural language description.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Speechcraft: A fine-grained expressive speech dataset with natural language description

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.356380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.356380Z digest=sha256:1701741894e4c4e1364377ed39162c9d04ed142cb84577d5ee7ebea87afa2758

Observation 604d43fb-5ff3-4590-b857-a3b1e3df899f · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.453562Z digest=sha256:c1ceb1875282589eda890625bb8a07acca4097e8a6b0320afe8a50b512baddbc

Observation ad880798-f23f-401c-86a7-7da9281cad6e · outbound

This paper cites Scaling rich style-prompted text-to-speech datasets.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Scaling rich style-prompted text-to-speech datasets

Reference 10

Resolution
verified exact
doi, observed 2026-08-05T20:04:52.561894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:48.511222Z digest=sha256:aa805016eb277c77b591cdd0726138b9878551adb474df506ab37351f3863a3a

Observation da4489d1-bfb3-4e5a-88bc-db1fc786b8b6 · outbound

This paper cites GPT-4 Technical Report.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts GPT-4 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.592233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.592233Z digest=sha256:bc198a5bd01f6de5c7f6559f99778bf58e662c681e4dc3bee848396a527f1418

Observation b68494d3-07d7-4dba-a3df-bf733c33b7f2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.684230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.684230Z digest=sha256:970a9e22f4ab3e43d3b2265732539a47e4ae31035cd48355eea7d3e90f22d257

Observation f0be7c83-54b5-4e2c-b38c-2317bd581025 · outbound

This paper cites an unresolved cited work.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:04:55.112183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:48.772419Z digest=sha256:efaabd8c0d2091ae2d0898ef4a95035c4d5a6bd185c94b6cb71833707d918938

Observation 971ec484-d792-4683-9284-c4dd997917f3 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.852163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.852163Z digest=sha256:99815ac8be6a480c5dc79ee1e8e62c17a145d75e2ae27a1f11014d1c455a62f3

Observation e8030ec9-5c4e-4c8b-b9de-b51f1d09328a · outbound

This paper cites EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:48.933028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:48.933028Z digest=sha256:4d1ea22c892826f28c2d715b35ba8758313358b3be68e9c7e2bc92ca34e28a41

Observation cf596226-53b9-4921-ab0f-3282f7c3cfe8 · outbound

This paper cites Qwen3 Technical Report.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.009853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.009853Z digest=sha256:bd06494523d41fc45d83880f23ccd967f63e78a43b84c06592b4ee4e4f17ab55

Observation bf9a4a57-a755-4743-aae5-8ba47bbd97e6 · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.104039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.104039Z digest=sha256:583ef3d39b97d9fa1ac9a7aa45f0fdeac21edf806c6d3b6c64fbcc61419a3382

Observation d8fb27d0-048f-43b8-8811-f6ae30b3d965 · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.957933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:49.260735Z digest=sha256:ad0901cbd5205baacab7abc8fab0a035e4d39dda6d15d2b16a6673b12a60b64a

Observation 8b2ff2d1-6018-4309-95ab-bf44defa8f83 · outbound

This paper cites Investigating the Catastrophic Forgetting in Multimodal Large Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Investigating the Catastrophic Forgetting in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.161915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.161915Z digest=sha256:9ccf3a0d47222037023a118ee590f6cb0376c012e211bab24569c540c8c944a7

Observation 43dfd901-5a48-435a-97c5-6aa2d561ba95 · outbound

This paper cites Vlmo: Unified vision- language pre-training with mixture-of-modality-experts.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Vlmo: Unified vision- language pre-training with mixture-of-modality-experts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.706910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:49.447667Z digest=sha256:473f2af551fedafd09594849d3d0a8033de6106b71d367db5ab9056d5b123d8e

Observation 6313292f-a74b-4422-952c-043a5e25126a · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.342135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.342135Z digest=sha256:6749147f6b56864e67e43259b5033d6d73f56050498dd8de0c826b201a0382c7

Observation b0af524c-c2bd-43bc-b1e0-e4d2c969f240 · outbound

This paper cites Simbert: Integrating retrieval and generation into bert.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Simbert: Integrating retrieval and generation into bert

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.508410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:49.727665Z digest=sha256:59e0680978e18c5b33605ce40d6aea3037b39199539d37435f5ec974d706a2d9

Observation 813005c5-8fe4-4834-b72a-2f5f23fc6ff8 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.511799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.511799Z digest=sha256:571d6608b39f92bd99702f83a584173759d9357569b088231541443d15fc9173

Observation 695d6b7c-a62b-4a93-bf27-bbd25450c258 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Fastspeech: Fast, robust and controllable text to speech

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.168169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:49.987765Z digest=sha256:4bff0d98879525023ea0476012f2e9f0da8c147f43d960c7477580f25b5b0899

Observation faa15446-d0e2-479a-84e9-03b0b46f337b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.064063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.064063Z digest=sha256:6b2bef6ada38fdeebb61a857c5aa4936a4340242877dbf7ed13820db985ca6ac

Observation ac9cca37-f696-416f-8db7-12761b0f7f4a · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts BERT: pre-training of deep bidirectional transformers for language understanding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:54.327115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:49.809205Z digest=sha256:934f5428be0bd6e1e78ff499162e7ca315350f44c371b7f39c2fd25eece219a6

Observation 2c2df928-bc5f-48f1-b625-13fb9153f588 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.260659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.260659Z digest=sha256:c1e6e56fc4467dfe65a33abbb7046cfcf5233a004d47c0a9e185617d969e6566

Observation 88f43fff-0582-44ed-bff3-0f42b3fb910a · outbound

This paper cites Scaling vision-language models with sparse mixture of experts.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Scaling vision-language models with sparse mixture of experts

Reference 28

Resolution
verified exact
doi, observed 2026-08-05T20:04:52.080129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:50.387922Z digest=sha256:88a0bd25c8c57bbe0c3b8751533e38f3387e26567e39d05a2125afdbad7119c3

Observation dae82b05-3ee8-421c-9566-a108eea12dd3 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.463688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.463688Z digest=sha256:5264bc5e90721a4ff2cbedec8a44ac23b7a12b5c5cf9c8524b515a0c05615c25

Observation 653e739d-550c-4d36-abff-8a836d6e1634 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.143999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.143999Z digest=sha256:7e7d00e936c43654350062d69fb789ddcdec0782988614dbf41b3ce5847cfa35

Observation 2bc8c221-afe1-4c67-97d8-90be3bb98c38 · outbound

This paper cites The Llama 3 Herd of Models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.680474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.680474Z digest=sha256:0f0de33a2256cd8a1eab8370dacbc36c37503a62cf191d4f9a64a99449a9f8c9

Observation d17d0c2a-3b7c-4648-827f-3ffda6d2e73f · outbound

This paper cites Orpheus tts: Towards human-sounding tts.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Orpheus tts: Towards human-sounding tts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.997908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:50.781373Z digest=sha256:99b148668e1e1562a8fa58817d273176e00b56fb7a19a2c7d0231915d012abeb

Observation fe17e825-4a1a-4234-a1b4-7f458d45e298 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.870799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:50.872744Z digest=sha256:f09fac0ebfefa35102af594277a409cb40ec3729a98969a16361fa4b2002e6b8

Observation cf0c18c0-5f8a-4972-b220-3ebce6c1707b · outbound

This paper cites Qwen2.5-1M Technical Report.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Qwen2.5-1M Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.570486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.570486Z digest=sha256:4bbf444f50d8743f4d01e4fc19ce9b1a49239f240a3a188f4772246ed78210f1

Observation 66e3afe6-3357-4c14-8fde-7939efa6c934 · outbound

This paper cites Neural discrete representation learning.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Neural discrete representation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.559273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:51.113315Z digest=sha256:f40a50087e4a8d891c0c60d9fbac7f6be18822012f96c0604a13884ebc446fb7

Observation dc0ecd74-90e6-4083-b7bc-ac851aba7832 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Soundstream: An end-to-end neural audio codec

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.210481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.210481Z digest=sha256:7ce3aea77923f3df61d5b32abc4f1cfd803f03b29fd1789ef33c52f95493050e

Observation 4ed51057-2fdc-4987-a663-7e456b2448e6 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts High-fidelity audio compression with improved RVQGAN

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.446296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:51.289926Z digest=sha256:d00d055467a38cdd57787bffe6118ef02bbfc1024b134cbe1bc32e87839d0b32

Observation 9accb710-3a08-439c-b5ad-e7e7fbbba6c0 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.369330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.369330Z digest=sha256:825d2ba47b41148f2a33ed41e57de40117cf2c97de1588b972bbc313f652fcd9

Observation bdc1c6f7-2d9a-4f7d-9bd7-3bf548064f4d · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Self-supervised learning with random-projection quantizer for speech recognition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.769779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:51.037030Z digest=sha256:afe43e2f41921b4e54208b3e917de92935f951b0b1aabfc476ca4d0e107bbf9e

Observation ecd9b8f4-018b-449c-80fd-7bc44ad28dfb · outbound

This paper cites Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.505861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.505861Z digest=sha256:e9c98e84f7ad97092809abafcf83a88ff3ef55f15690044d3c2960f3ddd1ac51

Observation 4dc488d9-eef8-4a29-9e0e-fef17261ee57 · outbound

This paper cites Albergo, Nicholas M.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Albergo, Nicholas M

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.590164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.590164Z digest=sha256:800a8a38f2fb19d6d485c65396c9a7a23f6f14a1d2c0a15edb5e3c4b7d43d652

Observation a6b3cc7b-26e8-46e9-9a56-07636ec06f1b · outbound

This paper cites Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Emilia: A large-scale, extensive, multilingual, and diverse dataset for speech generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:51.686582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:51.686582Z digest=sha256:11d0a0e12718faebae279d5df187546d0a5012494898e5ffd3c07a0bf88e9da8

Observation be96fe05-0c7d-4492-9fbe-fc61f0d13e93 · outbound

This paper cites Elucidating the de- sign space of diffusion-based generative models.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Elucidating the de- sign space of diffusion-based generative models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:04:53.311325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T20:04:51.448310Z digest=sha256:53700988b67cbe19bdc95c22b8b7c424cc87166eedef0e80c66057acd298ede9

Observation 8854b402-7093-4e14-9a29-82a68a516603 · outbound

This paper cites URL https://doi.org/10.1109/TASLP.2021.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://doi.org/10.1109/TASLP.2021

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.973979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.973979Z digest=sha256:f343db353429cb665cbcc98a9b4bfd9d55e3a32a852d6600b7fb6cf2534c0ffc

Observation 5d35de97-2743-4564-959d-2c91eed98a3a · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.612420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.612420Z digest=sha256:fd0d73a37ccd243e7e451be555287f0b17ed465f2103460d2ecf5086952bc69a

Observation 1a804652-6dfb-4a4a-9f30-ad5deddc4abd · outbound

This paper cites URL https://doi.org/10.1109/ ICASSP49357.2023.10096285.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts URL https://doi.org/10.1109/ ICASSP49357.2023.10096285

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:47.813576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:47.813576Z digest=sha256:046a54a6d0e69279f450888d104a6cae668fd10320e8cd7766d071c4c111ad58

Observation e5145ede-99f9-4cef-bc1e-8e700f33f581 · outbound

This paper cites doi: 10.18653/V1/N19-1423.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts doi: 10.18653/V1/N19-1423

Reference 4186

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:49.884453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:49.884453Z digest=sha256:b71dad06363eacb92a5e46b8861f251b789f9c5cf94e197c7c0cf9894faafb0c

Pith citing papers

Observation 5c51c5e4-1523-44d3-a286-74cf8ac7a6d0 · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.871334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:6f1e8d8c717ded4e5e0be1122a322683b77f1a061e767d31627bbde4c08b54a8