Pith. sign in

Paper Citation Record · LEDGER

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

As of 17 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 21 inbound Pith citation observations for arXiv:2505.07916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07916 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:16:17.331703Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:40:24.938990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:49:24.877685Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9659b86-5361-4832-a840-652caf34245c · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.212140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.212140Z digest=sha256:66ffdb88535ac872148e2018b058f613a957f16708c4974dcd7135c2e163a9fc

Observation 44f878d8-1be9-4f38-b1e1-9f7b288489c2 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.232012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.232012Z digest=sha256:7e6a788053b800400171e423fde91c161eaede2ed4b3588ff88f89c02dc101e6

Observation 0199b579-b207-4108-8ed2-ca7e01bc1c82 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.237166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.237166Z digest=sha256:3c810cd678e00a6979635c6168b9117fe27d1cecfb1b0c3cfe3275dc52cd0384

Observation 319a8379-7973-4d43-b338-8f9a01d550ea · outbound

This paper cites NICE: non-linear independent components estima- tion.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder NICE: non-linear independent components estima- tion

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.651495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.241924Z digest=sha256:3a5e9171b516ad54d5556b3aeda8da43c9ad5fe0c52ca11f5fe212004af4c025

Observation 4f29587a-cba9-463a-8dfd-7d45c3405628 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.255201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.255201Z digest=sha256:fcd69c14c3024c96bd33f165737b0d0447e4a65a380b5d5b75b298e46bf5e627

Observation 450f5f95-ce57-49e9-964b-de015244eaee · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.265205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.265205Z digest=sha256:0f28366ddd6e906d47faa9c118bc93e7ddfe7244e2dad8bccb2bcc6b191a0d25

Observation e9baf5bf-116a-44c0-bdac-a77a3eb8da28 · outbound

This paper cites Efficient neural audio synthesis.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Efficient neural audio synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.610377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.275583Z digest=sha256:37e524cec98d4c5f5a3975d09a7198436253b03cce2130cbf1101186980a22a4

Observation 1dca5944-1989-40eb-ac7d-4630ce216ed9 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.285213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.285213Z digest=sha256:239798a2a134f6306b30cbe9b8a3d5e773dbede4c4ce75acee66ecdcc20d89dc

Observation 6c915cf2-1caa-4b07-a112-1a39cf7b00d7 · outbound

This paper cites Prefix-tuning: Optimizing continuous prompts for generation.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Prefix-tuning: Optimizing continuous prompts for generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.596252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.289702Z digest=sha256:88debc904824936e2505207ad0611816eeee59327482070c442f205ef80fa186

Observation 0fa9f953-ec73-49b9-9e8c-9fce775e225d · outbound

This paper cites P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.299028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.299028Z digest=sha256:e7996e0b50d78d494426741395e30fe3968809d7ec170245c33e0902980a4222

Observation 0bc4f18e-81d9-46e4-b9ee-da0a55f31ab2 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.Trans.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Improving and generalizing flow-based generative models with minibatch optimal transport.Trans

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.558354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.309149Z digest=sha256:e40eeb70fd230f1b50a62e280140d67e76537b5529a9d8dfade56aab68e1ce91

Observation 7c38ca61-f338-42de-a774-da32b2a21d15 · outbound

This paper cites LPCNet: Improving Neural Speech Synthesis Through Linear Prediction.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder LPCNet: Improving Neural Speech Synthesis Through Linear Prediction

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.544386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.313665Z digest=sha256:41a57f3a047f17aa23ccc03be0112ad18d53de164baac7940d79f19822aefaf4

Observation 054fb0be-2c45-4631-828c-597e6a09762c · outbound

This paper cites Senior, and Koray Kavukcuoglu.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Senior, and Koray Kavukcuoglu

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.530296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.318569Z digest=sha256:966d58f33ea9c42dfca3072ce879afb50fdd0dee4ff286fb9f20e0d926666a0e

Observation f7e47149-1103-499c-a18b-9ce6eb91b9d7 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.323155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.323155Z digest=sha256:694234758435a04b7575f2742306f161d0c72dacc7a2bf009cd236003bbb9a3b

Observation 5e71f19e-d50f-4ae4-8075-535d96840a6a · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.327211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.327211Z digest=sha256:533d9ba9ab38b3a666c5ed021d1b4256679867854a2288f0812f4ecd0759a608

Observation 07ee884e-10a3-469e-a0a5-68388c88c031 · outbound

This paper cites SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.331703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.331703Z digest=sha256:dfd4978cff340edd80303fcfdbdccd2c97ff813fb2d904bc75678fea68923642

Observation a27fcba3-2f7d-4b05-bc50-8f4ac196424e · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Matcha-tts: A fast tts architecture with conditional flow matching

Reference 1993

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.570339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.303637Z digest=sha256:437e0d53d962fe05d602df03c7d2281594493ecf36c63a9e53b01cf318640b8d

Observation 991c71c5-7b09-4c38-ab7a-bca578d30acd · outbound

This paper cites Density estimation using real NVP.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Density estimation using real NVP

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.638841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.246090Z digest=sha256:75264bed97b1800ed8aca2976f0fd7cdf909646d58243540f259dc0f7fcba649

Observation b9696730-0d52-4928-8d7e-1f3be68c4383 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.250085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.250085Z digest=sha256:f1cde1fd545cf295bc859931c8bfe48d769494fb32c8aa17b84ba22c189a9a42

Observation bce397b9-eff3-45d0-9698-6fbd1393201e · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.280466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.280466Z digest=sha256:97d403bcbb9c3b7715e0d58463a9db7abaf77a3bbcdda620a744ecfec15a3007

Observation b5baeffe-e4bf-42ca-88c4-73a767608be8 · outbound

This paper cites Better speech synthesis through scaling.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Better speech synthesis through scaling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.223139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.223139Z digest=sha256:99b62aaafd15710a29ecf219c1efcbad6fbc67c2ff9bf93e4b5432ac836df3a7

Observation d5c1561c-2a0c-4545-960c-524e0c4fb7fc · outbound

This paper cites an unresolved cited work.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:16:17.582304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.293839Z digest=sha256:7789b8675cb77d1e000a00facc410957f611100f2c44367e2a027364c7222baa

Observation 6a39e9de-aed9-4235-83d5-8ba3fefcf736 · outbound

This paper cites TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.260074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.260074Z digest=sha256:574d09de2f39c60c4c190175589f7fc6e8b1d5d1921d6da9bcf969abe18b887c

Observation 729c2f50-15f7-4a43-8ccb-7b99c27c46fa · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.227895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.227895Z digest=sha256:f8af8431ccad28f4abb8b438927ef9996cc827299f7ab47f93f6d9021d3add0a

Observation dc69c9c4-130b-4270-8cef-bcae2bb8a690 · outbound

This paper cites Tyers, and Gregor Weber.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Tyers, and Gregor Weber

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.673703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.217729Z digest=sha256:94d53ae85022735d25d5ec3efab4d594d464747634f447266afdcd1812e6c557

Observation 4cac08b3-7e89-4ec4-a475-9b3fbb1c28b6 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:16:17.625734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:16:17.271222Z digest=sha256:367df09937c0a799a5e1510c7468729c4b1e117d4e0a7389ca9e2c6e6f2aa833

Pith citing papers

Observation c6acac66-d6dc-4b43-b1fe-df73c9954fc7 · inbound

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models cites this paper.

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:16.347461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:16.347461Z digest=sha256:46bc43cf793dad48b5b171835dd663b1e27e1a059bd40baa223d669227eea8ce

Observation 5a717487-dfee-48e8-8cad-7c7e4e17f5a7 · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.720088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.720088Z digest=sha256:d9ce08f048f60eeeb62a46a6d2437006de715cc988cc16e8e20e87b6886d9ec7

Observation 5c555d22-4377-4e1e-b64f-0cc68308a95a · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.111664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:926df84bde7fe569556a6df4467a9e04435d4d9265e903ecf2d2c2b9f4c3ebe7

Observation bc5e8aef-029b-4609-8889-5fef279eb141 · inbound

TTS-1 Technical Report cites this paper.

TTS-1 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.891456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.891456Z digest=sha256:704714ec0f379dfc0bcca3e7077d416de9530396543c79f33e46d0f8a1517220

Observation 2c2df928-bc5f-48f1-b625-13fb9153f588 · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.260659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.260659Z digest=sha256:9e21eb88b911032913be09d45f711786bbcb3dc10c4c2d6d271eed498cacd5c0

Observation f404586a-99f0-435a-8b3f-92b8deed627d · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:38.178397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:c95ca1da97e823da964091fcedbe65336abc985b98398b70c4593ea0ec7f14f4

Observation 8bb46fa4-e9f3-4ddb-b5ca-0fde45dc48ab · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.108210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:ea51386b31ccb353ca410224d537afdecca329e19a97cbea0c314097dee77960

Observation 0337724a-9c59-481c-9ea8-1557a43548e9 · inbound

Voxtral TTS cites this paper.

Voxtral TTS MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.876040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:7a5d043b078f5a9e77db4fb438829f1e733ccfed98b217ae8d6ac6a6f5aa153f

Observation bf3233e1-8325-4ee6-bf24-e01632e14d75 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.769453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:e1adad108fecf5abc52891bd402fe24162c913ed5f4c57941a8d26932b889477

Observation a265c1c4-9ce2-46b3-87c6-f95636f518df · inbound

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control cites this paper.

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 4

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T18:51:06.774413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T13:49:09.386368Z digest=sha256:c884173e5f18cdbb095f339a042ca3628e23a80a98b66ef1b0cd700aa3fbf725

Observation bbcf6cd4-b43c-452a-bc91-d6fc1e178eba · inbound

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling cites this paper.

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:14.063829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T06:44:36.000353Z digest=sha256:9c82a4a3e9f5e546d2c890d8e970e99ac5672668a0ad99501b21cadad0b4dda7

Observation 73bb91e9-1be7-423b-b39d-a6e3c88be29a · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:00:58.977726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:6b27e30cbd5f1192fad33dfa7fc95bf72462b045758805cdc08a489074195770

Observation f036e8b3-a8aa-40dc-9b07-c05582b7026b · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.404566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.404566Z digest=sha256:cf9a32bb1e2a7a8932887732b50e8040cf01a37ae31fe93b5820ea983a795fbd

Observation 7a9da86e-b539-440f-b7d0-d589bed79e3e · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:b80db500e8bebcb6de2d6273d16cd0fc370afa36b3b514d914d477188522136e

Observation 15cdf836-4114-4b60-96f9-fb7dbc455bfa · inbound

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cites this paper.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.874719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:f8953e355ff0dd8f4a1c6ecfbc09b9e049ef4791fab0121c8ed71402e7144f04

Observation 69e13e97-23ab-42e5-8719-a38d336cc674 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.778528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:888d497bffa5aad072259c2802093c2888e6dcb82bb41f3765b9e82279ebe0fc

Observation 50e092b2-dc59-4516-9d83-940688b066fd · inbound

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors cites this paper.

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.879690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T19:04:25.542629Z digest=sha256:42abad15319a7f3d03d7a1b2352a29d2be5fd6c4c5a6131e96844996fede7a96

Observation 68de5465-d55c-4082-8340-0bfc92d14875 · inbound

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis cites this paper.

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.528359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T17:01:13.972071Z digest=sha256:cd1a0675cdd0d5e5148a092d926d3e4e3137df2710ff91c2214133d532a003e2

Observation 7dd17738-c6e4-496b-8b23-6700317c5213 · inbound

Luna-TTS Family Technical Report cites this paper.

Luna-TTS Family Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.938990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.938990Z digest=sha256:84397440da9be67c7c2bfaa41cb1b8fc3dd14615192e4cecb1ecaca974ed5409

Observation a4408303-cfc5-49ef-bf47-55239e67808e · inbound

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder cites this paper.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.760400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.760400Z digest=sha256:69d4cc8de9f00d738da309a9e5ea9add606a023e48f02202b9453a7c05ba9d2f

Observation 4b8ed3be-f1d0-4fc9-9555-fdace64326e2 · inbound

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching cites this paper.

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:30:28.369540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:30:28.369540Z digest=sha256:ec91dc338bcfe9c13828a4125efa472f6515d0aef037201e92674f89527e4cdc