Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:16:17.331703Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 21 inbound Pith citation observations for arXiv:2505.07916.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:16:17.331703Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:40:24.938990Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T02:49:24.877685Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a9659b86-5361-4832-a840-652caf34245c · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f878d8-1be9-4f38-b1e1-9f7b288489c2 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0199b579-b207-4108-8ed2-ca7e01bc1c82 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 319a8379-7973-4d43-b338-8f9a01d550ea · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder NICE: non-linear independent components estima- tion
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4f29587a-cba9-463a-8dfd-7d45c3405628 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450f5f95-ce57-49e9-964b-de015244eaee · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9baf5bf-116a-44c0-bdac-a77a3eb8da28 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Efficient neural audio synthesis
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1dca5944-1989-40eb-ac7d-4630ce216ed9 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c915cf2-1caa-4b07-a112-1a39cf7b00d7 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Prefix-tuning: Optimizing continuous prompts for generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0fa9f953-ec73-49b9-9e8c-9fce775e225d · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc4f18e-81d9-46e4-b9ee-da0a55f31ab2 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Improving and generalizing flow-based generative models with minibatch optimal transport.Trans
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c38ca61-f338-42de-a774-da32b2a21d15 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder LPCNet: Improving Neural Speech Synthesis Through Linear Prediction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 054fb0be-2c45-4631-828c-597e6a09762c · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Senior, and Koray Kavukcuoglu
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f7e47149-1103-499c-a18b-9ce6eb91b9d7 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e71f19e-d50f-4ae4-8075-535d96840a6a · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ee884e-10a3-469e-a0a5-68388c88c031 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a27fcba3-2f7d-4b05-bc50-8f4ac196424e · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Matcha-tts: A fast tts architecture with conditional flow matching
Reference 1993
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 991c71c5-7b09-4c38-ab7a-bca578d30acd · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Density estimation using real NVP
Reference 2015
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b9696730-0d52-4928-8d7e-1f3be68c4383 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce397b9-eff3-45d0-9698-6fbd1393201e · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5baeffe-e4bf-42ca-88c4-73a767608be8 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Better speech synthesis through scaling
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5c1561c-2a0c-4545-960c-524e0c4fb7fc · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a39e9de-aed9-4235-83d5-8ba3fefcf736 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729c2f50-15f7-4a43-8ccb-7b99c27c46fa · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc69c9c4-130b-4270-8cef-bcae2bb8a690 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Tyers, and Gregor Weber
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4cac08b3-7e89-4ec4-a475-9b3fbb1c28b6 · outbound
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6acac66-d6dc-4b43-b1fe-df73c9954fc7 · inbound
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a717487-dfee-48e8-8cad-7c7e4e17f5a7 · inbound
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c555d22-4377-4e1e-b64f-0cc68308a95a · inbound
Step-Audio 2 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bc5e8aef-029b-4609-8889-5fef279eb141 · inbound
TTS-1 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2df928-bc5f-48f1-b625-13fb9153f588 · inbound
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f404586a-99f0-435a-8b3f-92b8deed627d · inbound
Qwen3-Omni Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8bb46fa4-e9f3-4ddb-b5ca-0fde45dc48ab · inbound
Qwen3-TTS Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0337724a-9c59-481c-9ea8-1557a43548e9 · inbound
Voxtral TTS MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bf3233e1-8325-4ee6-bf24-e01632e14d75 · inbound
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a265c1c4-9ce2-46b3-87c6-f95636f518df · inbound
MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bbcf6cd4-b43c-452a-bc91-d6fc1e178eba · inbound
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73bb91e9-1be7-423b-b39d-a6e3c88be29a · inbound
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f036e8b3-a8aa-40dc-9b07-c05582b7026b · inbound
RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a9da86e-b539-440f-b7d0-d589bed79e3e · inbound
Raon-Speech Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15cdf836-4114-4b60-96f9-fb7dbc455bfa · inbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69e13e97-23ab-42e5-8719-a38d336cc674 · inbound
VoxCPM2 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 50e092b2-dc59-4516-9d83-940688b066fd · inbound
Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 68de5465-d55c-4082-8340-0bfc92d14875 · inbound
EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7dd17738-c6e4-496b-8b23-6700317c5213 · inbound
Luna-TTS Family Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4408303-cfc5-49ef-bf47-55239e67808e · inbound
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b8ed3be-f1d0-4fc9-9555-fdace64326e2 · inbound
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.