Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.819238Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 7 inbound Pith citation observations for arXiv:2502.03128.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.819238Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:07.252044Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T23:19:02.895734Z
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 078fdd36-19a2-4805-bf67-b57f4715db1c · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa735e97-4d1b-4b88-8167-547e681c3894 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874190da-82b6-4c2a-b0a5-6489711d148e · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2974013-dc9e-4bfe-ba11-fa1f12cffff2 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02251048-9c56-4018-b48c-6bc4fd013b8d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Common Voice: A Massively-Multilingual Speech Corpus
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e738f8d8-31f3-4e88-ad1c-18dcd74638b8 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Voice Conversion With Just Nearest Neighbors
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f342eb-cef5-4a8a-9753-3bf6a6872a55 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training BEiT: BERT Pre-Training of Image Transformers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e187116-59ac-4d0d-acb4-a1785930ef6c · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Better speech synthesis through scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49244d92-a743-4a79-a993-d510d3eefd72 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiolm: a language modeling approach to audio generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e5b5ae6-4a89-4b0c-acfe-8aa21b53d497 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SoundStorm: Efficient Parallel Audio Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4addb919-8823-4a23-bd6d-f0e0d6cfc1ba · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b887fea-28bc-4be0-aa05-17577a1bd9b4 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab7713c-374a-45dc-a085-e8a845ca2ad3 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf0e655-85f2-4d3f-9b8b-37945160f7b0 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf1b8ddc-77ca-4774-98c5-0d95aae76a1d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18c0bd10-8677-4536-8701-5c0ea894574c · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2963741-253b-434f-a6a8-99f3d43adfd4 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Self-supervised learning with random-projection quantizer for speech recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26033474-b91d-4c1c-bfeb-983bcc84e4b6 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Diff-hiervc: Diffusion-based hierarchical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18e1d699-ac5c-432a-9943-863877db5bdd · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Intelligible Lip-to-Speech Synthesis with Speech Units
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15cd84d7-9ff7-483f-9f7f-bd2a928f4861 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81187307-ea8b-4b5f-8e4b-937abbb286a9 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training LibriMix: An Open-Source Dataset for Generalizable Speech Separation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c42fac-eceb-42b9-86e1-70d678caea7a · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training High Fidelity Neural Audio Compression
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76096e5b-58f1-4b4b-9742-2e2851733e26 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbda1223-8e06-4cd4-88a2-fdc2d3aaf3c5 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training I., Waldner, F., Caccetta, P., and Wu, C
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e5f573-73db-400a-898c-7f3fb56cab33 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c089544-da15-42bf-8f53-fb11bb675a1f · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Icassp 2023 deep noise suppression challenge
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebedc275-c7ac-45db-af72-95d012609230 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Taming transformers for high-resolution image synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50c2960-5819-4513-a7f5-8bbc77fc416d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b89e85a-0dd9-495a-aa05-62feb8f82cbd · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b04887-867e-437a-b4fc-f3046be5db61 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b86198-926d-4bcc-a65a-e4a8158e5670 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3bba317-7182-4ecc-b6e8-88d69cfe9217 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Didispeech: A large scale mandarin speech corpus
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b976f6-b36a-401e-b381-059329f04c31 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68af5c54-7455-483a-86c5-30ee8222e7c7 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Masked autoencoders are scalable vision learners
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cb90a7-ee37-44c2-a596-65417a424889 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Classifier-Free Diffusion Guidance
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7798344-d706-4998-bd82-1599446ffc3b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 385a03f1-8709-46ec-9beb-0ec9b66dbd8b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training LoRA: Low-Rank Adaptation of Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10011a1-d3a8-4e72-8096-586ca8304808 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ca5bb1-b748-4c44-98be-e0d2b0eca1f9 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a955d75-8abb-488e-a61a-1808cfe3b27e · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Libri-light: A benchmark for asr with limited or no supervision
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 200296fa-2d2e-4df5-bf2a-f248c7c6cadc · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Libriheavy: a 50,000 hours asr corpus with punctuation casing and context
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce94c5b-f42c-488f-961a-8679e1f5255e · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925a626e-3539-420d-affa-301cb4551749 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d261d34-ae66-4e1c-bf42-20f2e8d4d545 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec50def-0479-43b4-8ac4-996f63c4a057 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training C., Lo, W.-Y., et al
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d7fed14-5f30-48e9-9628-4c842d792cd0 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training L., and Khudanpur, S
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46674634-6c73-4a31-a61a-2cbc1dc3ec08 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training High-fidelity audio compression with improved rvqgan
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8727cb72-1997-498d-bed3-a1f90fd36d10 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0583204-a10d-4fd0-b938-e4f94ba2ec2b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Crafting papers on machine learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d80736c7-4322-4487-8fbc-7300dd6c239e · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Voicebox: Text-guided multilingual universal speech generation at scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e1a0fcb-be09-4d73-934a-3e235b15ca80 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6ebac2-1ee9-4832-ae63-c5bfa326fb8a · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Improved masked image generation with token-critic
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03fa471d-c8a8-4bcb-b453-e794ff1e7d6b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Freevc: Towards high-quality text-free one-shot voice conversion
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02c1cd28-f134-41d4-a561-9010244950b7 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MaskSR: Masked Language Model for Full-band Speech Restoration
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1398b445-13fe-4c92-8269-ed1bbf4d2ed1 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Flow Matching for Generative Modeling
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50b2fdc-66d7-482f-9061-2a2cde300b3e · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Generative Pre-training for Speech with Flow Matching
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 391d55a6-a39b-46b9-b727-7a21b7806637 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceFixer: Toward General Speech Restoration with Neural Vocoder
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ea4b77-fc04-487a-b32a-736cfaecefbe · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a087ae35-4672-47b7-bd73-89d62abaedcd · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e878bbb0-5d97-44ef-8ab1-fe796331a090 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2d25a93-677f-48a3-8304-bbdc21413872 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Decoupled Weight Decay Regularization
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f75a97a-e785-4eee-912f-3b2a6d4a7866 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Auto-avsr: Audio-visual speech recognition with automatic labels
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c12aaf43-7de6-45ee-adfd-7cfa2152ef9d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f804ab5c-5ea1-4216-b6c9-23861bec9d05 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 23b77483-089a-4353-8464-a610bf0b36fa · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoxCeleb: a large-scale speaker identification dataset
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cb1fde5-ae9d-4721-af48-f8dcd95acb9d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8fe11650-a364-41f5-bac3-25ff1735f9b3 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training and Waibel, A
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e474235b-a6d9-48c8-9a29-0a95687b954b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Librispeech: an asr corpus based on public domain audio books
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 646efddb-8507-47f0-85e5-4bec0e3bd88d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SEGAN: Speech Enhancement Generative Adversarial Network
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b862f44-ab43-4145-833d-738fbad9bb76 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c8aa96-a383-4fc9-8102-2622bbb8e64d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MLS: A Large-Scale Multilingual Dataset for Speech Research
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ab8c129-6bf2-48bb-bc43-804503660c70 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Autovc: Zero-shot voice style transfer with only autoencoder loss
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d82cf2df-1836-4b1e-b2e7-87646d61f838 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training OpenVoice: Versatile Instant Voice Cloning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a6d39fa-513f-49d6-97d8-0004b35f4d2b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Language models are unsupervised multitask learners
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785e26aa-b645-468d-b18b-168d408cd971 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff9627c-32c5-4e6e-b5ca-7bfa5483616f · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training K., Gopal, V., and Cutler, R
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fafeeac0-60a1-4cd0-b55a-60900ba645e5 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Fastspeech: Fast, robust and controllable text to speech
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfe4fe30-e4b7-4e03-be0a-e17f861970d9 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6315151a-bfd4-449a-b68c-b0427656b917 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5ca28d-17ae-426b-a5b3-3936b1de4e81 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae922d9d-19b8-4f8b-bc0e-12295ace2f4f · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491cae55-699b-4b4c-8fac-bd585b4a2d6b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd0efa9e-8430-4bf0-97ec-4797e7591bbf · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960b469c-553f-4d20-98e6-7ecb07302aac · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Neural discrete representation learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83537652-d13a-446b-bc96-a0070a2c828d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Attention is all you need
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34be66d-e229-47ff-a52c-457d02c9aa2c · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 163393c2-c867-4746-b125-fe627ac4467d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44d84d4-925a-4bfe-9cde-451247cfef80 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387d996f-4d08-4dcf-861a-6178b55bc93d · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96c1a35-2d2e-476a-9f0b-dc1474778e8c · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e51422-5846-4ab5-a9af-00b0ac99a83f · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training E., Chen, S., Tang, M., Liu, S., Li, J., and Yoshioka, T
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f4d8a91e-e789-4371-b7d5-103e43f7811c · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7cd87c-be18-42aa-a4cf-c96eed0bddec · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Lm-vc: Zero-shot voice conversion via speech generation based on language models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a95679b0-4c62-457f-bd4d-1bd5ed94b882 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Selm: Speech enhancement using discrete tokens and language models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 273d08e5-d991-4608-a813-f5c7d9dd72f2 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Tf-gridnet: Integrating full-and sub-band modeling for speech separation
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82ca947b-6ee3-4c70-bb17-d2c49b3a1a0b · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training WHAM!: Extending Speech Separation to Noisy Environments
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487492ea-0bfc-4fc3-974b-06fc15d72545 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78a21a8-abbf-4ac3-9259-c77b9d0d1bf3 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54158eb-a70f-4b1b-8d1a-3de6e8727181 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training G., Yang, M.-H., Hao, Y., Essa, I., et al
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75875b51-e68a-4244-8ca0-e4b82c4592d3 · outbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb056b2-5e5a-4357-a714-17287e0219e7 · inbound
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59be1fe7-50bc-4cd0-9420-a896d77f6338 · inbound
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be49d1f-980f-499b-93b9-695450da0e11 · inbound
Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe196cd4-f51f-4a5a-9ce6-6f3312196439 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a6be6d39-106f-4b77-9a27-073f67aa5b06 · inbound
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cddd88ed-459f-40af-b8f9-d782ac39fceb · inbound
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31f23d98-8ad1-4700-a8d9-f6089db8771c · inbound
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.