Pith. sign in

Paper Citation Record · LEDGER

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

As of 10 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 7 inbound Pith citation observations for arXiv:2502.03128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03128 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.819238Z

measured 107 of 107 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:07.252044Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:19:02.895734Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved89
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 078fdd36-19a2-4805-bf67-b57f4715db1c · outbound

This paper cites write newline.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.378414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.378414Z digest=sha256:ed85c64c1e013d0c96b7f3c73925a2a133944364d6ed79f165bb1efd6fe03c6b

Observation aa735e97-4d1b-4b88-8167-547e681c3894 · outbound

This paper cites GPT-4 Technical Report.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.384234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.384234Z digest=sha256:3a81cb958be4acc0c5bc8c662f266b983c5c2a48b53922f89e8df720b2003673

Observation 874190da-82b6-4c2a-b0a5-6489711d148e · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.389057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.389057Z digest=sha256:56b1e51fc1fc41821345aae72965a26e0f3727dcb74a35ab9a33271db7b642c8

Observation c2974013-dc9e-4bfe-ba11-fa1f12cffff2 · outbound

This paper cites VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.393875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.393875Z digest=sha256:f591058f73eb6c73281591cb2e23874fa4e33dff8d4cad00c6f6f598cccea68c

Observation 02251048-9c56-4018-b48c-6bc4fd013b8d · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Common Voice: A Massively-Multilingual Speech Corpus

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.398909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.398909Z digest=sha256:6117add1bee8a25e17ba12112c8fb253d8f5f084eadb2baa639da1610b93cc5e

Observation e738f8d8-31f3-4e88-ad1c-18dcd74638b8 · outbound

This paper cites Voice Conversion With Just Nearest Neighbors.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Voice Conversion With Just Nearest Neighbors

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.403711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.403711Z digest=sha256:32842f1e53a6c9fcdb1904f8362dd3dc073c0e93c1ea278e2fa87687e6bc8fdf

Observation 85f342eb-cef5-4a8a-9753-3bf6a6872a55 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training BEiT: BERT Pre-Training of Image Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.408360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.408360Z digest=sha256:cef5b25ca930ed3333bb1a4571d53ec55fbf28993dbda6299073e5d89917a153

Observation 7e187116-59ac-4d0d-acb4-a1785930ef6c · outbound

This paper cites Better speech synthesis through scaling.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Better speech synthesis through scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.413462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.413462Z digest=sha256:5e7b2c52b047cbbe900a98b5270efceb77da2a08464937ce808c39d0d4263aee

Observation 49244d92-a743-4a79-a993-d510d3eefd72 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiolm: a language modeling approach to audio generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.418329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.418329Z digest=sha256:793d663d14b2e5638a162b4f259c37868414316ff8f9762cc234b2895825ccac

Observation 0e5b5ae6-4a89-4b0c-acfe-8aa21b53d497 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SoundStorm: Efficient Parallel Audio Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.422665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.422665Z digest=sha256:1d94a9f294ce01d6ce4fdb581aa5bbca2ea78c36604b2d442a5f656fbbaca237

Observation 4addb919-8823-4a23-bd6d-f0e0d6cfc1ba · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.427049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.427049Z digest=sha256:ad40f4ceab4dcd9690885ece9ed1db6f590c6366f80cca948e349a3f952726c7

Observation 3b887fea-28bc-4be0-aa05-17577a1bd9b4 · outbound

This paper cites an unresolved cited work.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.431531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.431531Z digest=sha256:8cd3dfdab4eb7c1fc57dd58b5721a8a72793ac83491d8bf3ae63f48846e542bd

Observation 8ab7713c-374a-45dc-a085-e8a845ca2ad3 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.435628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.435628Z digest=sha256:2d9772efea15e8164ae03488202e59c44a39dc11af738a63db58cd17b550a7e5

Observation 3bf0e655-85f2-4d3f-9b8b-37945160f7b0 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.440103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.440103Z digest=sha256:b0f963322722415a4240a7f8675247cf4710c311dc27974a3aaeebfccc957ee0

Observation bf1b8ddc-77ca-4774-98c5-0d95aae76a1d · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.444712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.444712Z digest=sha256:fbe9f65212e80785f7b921a678d87c845295c5df66991c3e0697c68248444f92

Observation 18c0bd10-8677-4536-8701-5c0ea894574c · outbound

This paper cites Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.448981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.448981Z digest=sha256:4975cfc300a926368d0ad78fdc1ef8b5298f3e99763974cf4610ce112c14a8a9

Observation e2963741-253b-434f-a6a8-99f3d43adfd4 · outbound

This paper cites Self-supervised learning with random-projection quantizer for speech recognition.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Self-supervised learning with random-projection quantizer for speech recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.453446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.453446Z digest=sha256:79c11233c620b129f551f50a2eb000ef7559ce15a507d15ff1b227e52f4789cb

Observation 26033474-b91d-4c1c-bfeb-983bcc84e4b6 · outbound

This paper cites Diff-hiervc: Diffusion-based hierarchical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Diff-hiervc: Diffusion-based hierarchical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.457801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.457801Z digest=sha256:22daebc5bca275b3d3b181de8a7bbbbc789aca33a1eeeac8ac9aa823faf565fd

Observation 18e1d699-ac5c-432a-9943-863877db5bdd · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.462539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.462539Z digest=sha256:ea4cf960466153adfb6f3147c4649854deb381ed83d71123a4169557275069ee

Observation 15cd84d7-9ff7-483f-9f7f-bd2a928f4861 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.466972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.466972Z digest=sha256:d0c9b8d212ec1ec96c81a1102aac35e05c1ffb49d080f454f0e0e8eab6078686

Observation 81187307-ea8b-4b5f-8e4b-937abbb286a9 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.471270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.471270Z digest=sha256:ca09571cfc184e08bec61d3d68510aba26e072f3c7010470ed2ffd7bc66eb02c

Observation 25c42fac-eceb-42b9-86e1-70d678caea7a · outbound

This paper cites High Fidelity Neural Audio Compression.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training High Fidelity Neural Audio Compression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.475968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.475968Z digest=sha256:13dffa17c3fc47713f974dbde0b48725ec04bd6e9b3de3237b8144288c1de662

Observation 76096e5b-58f1-4b4b-9742-2e2851733e26 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.480389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.480389Z digest=sha256:4137be18e3213d79da5fcb71d331f88ed89163a85b4d3daf8f8c85cf98a73f3c

Observation bbda1223-8e06-4cd4-88a2-fdc2d3aaf3c5 · outbound

This paper cites I., Waldner, F., Caccetta, P., and Wu, C.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training I., Waldner, F., Caccetta, P., and Wu, C

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.484930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.484930Z digest=sha256:997a6c46be92a9ffdc64aea9c89dc7b7c2adc23acd3e1cccc44a687dcb205446

Observation e4e5f573-73db-400a-898c-7f3fb56cab33 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.489311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.489311Z digest=sha256:58c7af574a7389c3cabc1fb32bb6b8d77e56b298a537d8e5caacc927c7343674

Observation 3c089544-da15-42bf-8f53-fb11bb675a1f · outbound

This paper cites Icassp 2023 deep noise suppression challenge.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Icassp 2023 deep noise suppression challenge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.493798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.493798Z digest=sha256:1dd1c7dcb1a4c41381d688b98031ca6ad26988e904360d9782ae96515fd49fe3

Observation ebedc275-c7ac-45db-af72-95d012609230 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Taming transformers for high-resolution image synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.498038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.498038Z digest=sha256:3cbc62eb6ba3b84cd17d2f678b56a29ed13ee3d26b3056c3319c835daa80d553

Observation f50c2960-5819-4513-a7f5-8bbc77fc416d · outbound

This paper cites MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.502297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.502297Z digest=sha256:eed5e809a9f5bf4c6c3c6f9ee55dd40420cef50b15ff5743550daf941f033df8

Observation 1b89e85a-0dd9-495a-aa05-62feb8f82cbd · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.506600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.506600Z digest=sha256:2d0f78d50561ea7943a75dcc17916d2567fe9d231ab17cd3045db05f7852a620

Observation e4b04887-867e-437a-b4fc-f3046be5db61 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.511185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.511185Z digest=sha256:0d0c01126b4a14fa9703f455babbeddab0e6a515cb92b6294af318a38eb52662

Observation 16b86198-926d-4bcc-a65a-e4a8158e5670 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.515495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.515495Z digest=sha256:d38baa92711aa4890b23df21267c7623b1c5decd809153aabdcb57146f145b67

Observation e3bba317-7182-4ecc-b6e8-88d69cfe9217 · outbound

This paper cites Didispeech: A large scale mandarin speech corpus.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Didispeech: A large scale mandarin speech corpus

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.519972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.519972Z digest=sha256:645084867d6a57da386651a90b5bd74aa6da1aefcd133db6a0585f1ca902dc0b

Observation 12b976f6-b36a-401e-b381-059329f04c31 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.525240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.525240Z digest=sha256:2c9f7b1cb969e2282f47f88842ceda10b3d859fab65abaea165947c88eba57b6

Observation 68af5c54-7455-483a-86c5-30ee8222e7c7 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Masked autoencoders are scalable vision learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.529842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.529842Z digest=sha256:73edf929c5f8180f339462b222a2c32bdfebc112a2a40bd2cc9b222135abc2e8

Observation 21cb90a7-ee37-44c2-a596-65417a424889 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Classifier-Free Diffusion Guidance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.534140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.534140Z digest=sha256:3172355be2eff47992dbf8c4beffb907a9d4a970266becee28277f9d54c01083

Observation e7798344-d706-4998-bd82-1599446ffc3b · outbound

This paper cites H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.538720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.538720Z digest=sha256:8b185a0a0f2f6bb1c097c7391136f4f678dcb9dfee3e9d9f1c71defaa0f4bfc6

Observation 385a03f1-8709-46ec-9beb-0ec9b66dbd8b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training LoRA: Low-Rank Adaptation of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.543003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.543003Z digest=sha256:2e6b950b68d522585c86d1c0e1f7222694018a413ccdc427350a2eecf1e193e6

Observation d10011a1-d3a8-4e72-8096-586ca8304808 · outbound

This paper cites Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.547292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.547292Z digest=sha256:f29958b6225c528f439900c885b501cc78ed101d2e8576401193c974b844e1fc

Observation e6ca5bb1-b748-4c44-98be-e0d2b0eca1f9 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.551748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.551748Z digest=sha256:9cf5e44fa6d356ff48e8abcfd5eda649c4d1b1d752c86b907b8dd5654f4aa8e5

Observation 2a955d75-8abb-488e-a61a-1808cfe3b27e · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Libri-light: A benchmark for asr with limited or no supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.556103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.556103Z digest=sha256:b48a56be38911e902525a29043f1bb22e299dbb29f302276e8c643cfda898a17

Observation 200296fa-2d2e-4df5-bf2a-f248c7c6cadc · outbound

This paper cites Libriheavy: a 50,000 hours asr corpus with punctuation casing and context.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Libriheavy: a 50,000 hours asr corpus with punctuation casing and context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.560652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.560652Z digest=sha256:d4780e829897c74b5c0a23d30a463fd91fb0c0eac94f577ee92499657bd93d69

Observation 7ce94c5b-f42c-488f-961a-8679e1f5255e · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Speak, read and prompt: High-fidelity text-to-speech with minimal supervision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.564817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.564817Z digest=sha256:5cbe7a2970c25da2f2e20e2fa962e4caf6cf4ec022dbb9a67d6f57fde3fcbf22

Observation 925a626e-3539-420d-affa-301cb4551749 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.568816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.568816Z digest=sha256:0b7387b5102a95935f1a3e166e31040a6f2af0ff05cfb27a8784814491593c49

Observation 0d261d34-ae66-4e1c-bf42-20f2e8d4d545 · outbound

This paper cites an unresolved cited work.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.573066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.573066Z digest=sha256:74bd864858add314cb4d68e970d37642cf45a9d41423f83859562243b29571a0

Observation 7ec50def-0479-43b4-8ac4-996f63c4a057 · outbound

This paper cites C., Lo, W.-Y., et al.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training C., Lo, W.-Y., et al

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.577254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.577254Z digest=sha256:282c4e714ad89d869b084807fe8f4cd4f615949ec27389c9f7a0fcaba7438676

Observation 8d7fed14-5f30-48e9-9628-4c842d792cd0 · outbound

This paper cites L., and Khudanpur, S.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training L., and Khudanpur, S

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.581433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.581433Z digest=sha256:25d18165afa6760e17584a494e650a27c5797224f338965761e9f616578512ea

Observation 46674634-6c73-4a31-a61a-2cbc1dc3ec08 · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training High-fidelity audio compression with improved rvqgan

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.585757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.585757Z digest=sha256:a34bf0a957e398db2a5066acf8caa2372ff4729d848352d4d65f198dd7e62290

Observation 8727cb72-1997-498d-bed3-a1f90fd36d10 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.590124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.590124Z digest=sha256:38d1485b4a168a5290beb3ab398fa9a4bac2c522fc4f8ec29b592854f7576f5f

Observation a0583204-a10d-4fd0-b938-e4f94ba2ec2b · outbound

This paper cites Crafting papers on machine learning.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Crafting papers on machine learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.594841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.594841Z digest=sha256:3df8a7209bd2bf72af1354c41a06d9e16f376470375ad137df309877a2849a4d

Observation d80736c7-4322-4487-8fbc-7300dd6c239e · outbound

This paper cites Voicebox: Text-guided multilingual universal speech generation at scale.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Voicebox: Text-guided multilingual universal speech generation at scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.599012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.599012Z digest=sha256:4bfa499148d4450c05ab8c2b54587ea8ab0e8a280a722ba668396aeaf2c2f5db

Observation 1e1a0fcb-be09-4d73-934a-3e235b15ca80 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.603332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.603332Z digest=sha256:88215a3a74eb81304799cbb21e11123a77bdd927aa0a334d6ebb4b262e44da82

Observation de6ebac2-1ee9-4832-ae63-c5bfa326fb8a · outbound

This paper cites Improved masked image generation with token-critic.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Improved masked image generation with token-critic

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.607943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.607943Z digest=sha256:1fa888cc85515ff478688bbd87842c54b5029f995d7e3ac7b2c7d44ee306ec48

Observation 03fa471d-c8a8-4bcb-b453-e794ff1e7d6b · outbound

This paper cites Freevc: Towards high-quality text-free one-shot voice conversion.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Freevc: Towards high-quality text-free one-shot voice conversion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.968189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.612106Z digest=sha256:034aa7a61ed075a5125a01eda3bb462dedbb890e42bf786dea046eeb1736bfe8

Observation 02c1cd28-f134-41d4-a561-9010244950b7 · outbound

This paper cites MaskSR: Masked Language Model for Full-band Speech Restoration.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MaskSR: Masked Language Model for Full-band Speech Restoration

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.616170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.616170Z digest=sha256:17eb87c84637f6bb7795ae40d5e69a19d59bac3d7a0e46bd164471820e993fca

Observation 1398b445-13fe-4c92-8269-ed1bbf4d2ed1 · outbound

This paper cites Flow Matching for Generative Modeling.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Flow Matching for Generative Modeling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.620605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.620605Z digest=sha256:58089e59a626f5aaa89d1c1f9d2c0ae8f03713cbfd819a7cb334272ade36a7f2

Observation e50b2fdc-66d7-482f-9061-2a2cde300b3e · outbound

This paper cites Generative Pre-training for Speech with Flow Matching.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Generative Pre-training for Speech with Flow Matching

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.625064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.625064Z digest=sha256:e7cfd931b1d7839ceb9237a47660bf762ef0b57664cf48923e7ba8ba4c763b02

Observation 391d55a6-a39b-46b9-b727-7a21b7806637 · outbound

This paper cites VoiceFixer: Toward General Speech Restoration with Neural Vocoder.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceFixer: Toward General Speech Restoration with Neural Vocoder

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.629836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.629836Z digest=sha256:8c4e8fcacb7d65f89a371e030075661fc56fc3793bc6152cd8f5e9ee5a285fa2

Observation 16ea4b77-fc04-487a-b32a-736cfaecefbe · outbound

This paper cites VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceFixer: A Unified Framework for High-Fidelity Speech Restoration

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.634360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.634360Z digest=sha256:f4c2cad7d80bdeb4059ad521190059d50a0ed9a718b5d9d4dc59d32757b214d2

Observation a087ae35-4672-47b7-bd73-89d62abaedcd · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.639029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.639029Z digest=sha256:6a09133179dbf7d53bfcb86c649e3af76b822ab86d9c93b547cd5cb8d5a3f53c

Observation e878bbb0-5d97-44ef-8ab1-fe796331a090 · outbound

This paper cites an unresolved cited work.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.643499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.643499Z digest=sha256:3c35509b1b328fcd082ad35c76281d7e9b4921985bb1473e53aa363e640dff08

Observation a2d25a93-677f-48a3-8304-bbdc21413872 · outbound

This paper cites Decoupled Weight Decay Regularization.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Decoupled Weight Decay Regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.647739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.647739Z digest=sha256:806aec443cbd027ee53b3ddb64a2c4b246dd502b17d360ad70daffd12e84630e

Observation 2f75a97a-e785-4eee-912f-3b2a6d4a7866 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.652221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.652221Z digest=sha256:55a90b686a0a7787c3fc0feaf864127e490c82b38d39b738216d871b5c8ec2c5

Observation c12aaf43-7de6-45ee-adfd-7cfa2152ef9d · outbound

This paper cites NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.656523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.656523Z digest=sha256:715c877ea3644f5d11e8a396f5518ec4507f2ad855ca6994e57768824e714c03

Observation f804ab5c-5ea1-4216-b6c9-23861bec9d05 · outbound

This paper cites an unresolved cited work.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-09T05:54:19.936454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.661197Z digest=sha256:f27dd9fd1c4e29b2eb41f51d4614c750306970bc55318ccf43968acbcf3932e0

Observation 23b77483-089a-4353-8464-a610bf0b36fa · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoxCeleb: a large-scale speaker identification dataset

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.665957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.665957Z digest=sha256:be375d585c59a42f439cca700f49d5cba36b1d2fc26a2d6ace5179b35de808ed

Observation 6cb1fde5-ae9d-4721-af48-f8dcd95acb9d · outbound

This paper cites SelfVC: Voice Conversion With Iterative Refinement using Self Transformations.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SelfVC: Voice Conversion With Iterative Refinement using Self Transformations

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:54:19.190650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.670425Z digest=sha256:03e1aeca461b84008b8ff5c9727e11026323f512515f1701b4a6b6f04289c961

Observation 8fe11650-a364-41f5-bac3-25ff1735f9b3 · outbound

This paper cites and Waibel, A.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training and Waibel, A

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.922188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.675023Z digest=sha256:9576eb3ce6512f154859df0258e68a0d9c583a049b2f9362f76d4dd00f248763

Observation e474235b-a6d9-48c8-9a29-0a95687b954b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Librispeech: an asr corpus based on public domain audio books

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.679338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.679338Z digest=sha256:20784b2bd4265ace389da8fecaee8e370ab15c37c7928d57ea905e95dff3e8e7

Observation 646efddb-8507-47f0-85e5-4bec0e3bd88d · outbound

This paper cites SEGAN: Speech Enhancement Generative Adversarial Network.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training SEGAN: Speech Enhancement Generative Adversarial Network

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.683587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.683587Z digest=sha256:4c3f5847bfcaea70b4e8727ea2db78f558065f1f8ade43f4d8c52836c5080a79

Observation 2b862f44-ab43-4145-833d-738fbad9bb76 · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.688104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.688104Z digest=sha256:cf2e616fc5273bd20e60b25372fa22a78a8cb2e19a831e2e8af93c2a1a604546

Observation c8c8aa96-a383-4fc9-8102-2622bbb8e64d · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.692824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.692824Z digest=sha256:ade637dee90db278fba3bd9977bcabe18b110d11d161194a0789383e97eeba5d

Observation 5ab8c129-6bf2-48bb-bc43-804503660c70 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Autovc: Zero-shot voice style transfer with only autoencoder loss

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.899303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.697103Z digest=sha256:c1db0ee5b4773608c803a7c7cb9645b0788edfbdb2252a86919431eb449e7f93

Observation d82cf2df-1836-4b1e-b2e7-87646d61f838 · outbound

This paper cites OpenVoice: Versatile Instant Voice Cloning.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training OpenVoice: Versatile Instant Voice Cloning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.701474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.701474Z digest=sha256:99e483daf40736600241163d804ef1db610eb6169861ba6a79f337e9a94d7757

Observation 8a6d39fa-513f-49d6-97d8-0004b35f4d2b · outbound

This paper cites Language models are unsupervised multitask learners.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Language models are unsupervised multitask learners

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.705838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.705838Z digest=sha256:5924539cc69af3e1f5030cfd875d7d06c23f9aad73adf34a95006091e3756ec2

Observation 785e26aa-b645-468d-b18b-168d408cd971 · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.709976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.709976Z digest=sha256:3fac83da72e3a7fc2b7a9da38bc0eea7ecf3f8854771662927503c8480e0bace

Observation 4ff9627c-32c5-4e6e-b5ca-7bfa5483616f · outbound

This paper cites K., Gopal, V., and Cutler, R.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training K., Gopal, V., and Cutler, R

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.866997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.714080Z digest=sha256:bb9f74c3540ad4aa0b69db0e4196bda4443e1a7c3ce1901d5d689a543d780450

Observation fafeeac0-60a1-4cd0-b55a-60900ba645e5 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Fastspeech: Fast, robust and controllable text to speech

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.851663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.718356Z digest=sha256:20e4f9ac55cefec6dd9c3d2673c9ce5624f7053653ef9a3f7af07c12a791e11f

Observation cfe4fe30-e4b7-4e03-be0a-e17f861970d9 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.722514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.722514Z digest=sha256:7ebb88e87495c8d6b364c5753746487e0f8ac58469448f1846dd6a29901e0c9e

Observation 6315151a-bfd4-449a-b68c-b0427656b917 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.726951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.726951Z digest=sha256:f9ce180826aa6766926a62b3a5a78136adf0d324bb21beb8028b98472e9a2bc7

Observation 7d5ca28d-17ae-426b-a5b3-3936b1de4e81 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.731349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.731349Z digest=sha256:34b3248410ebf8d0ee25d49b881528f4bd4ae03bbac4da69d8bfc905b977c110

Observation ae922d9d-19b8-4f8b-bc0e-12295ace2f4f · outbound

This paper cites TSELM: Target Speaker Extraction using Discrete Tokens and Language Models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training TSELM: Target Speaker Extraction using Discrete Tokens and Language Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.735975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.735975Z digest=sha256:ed43c660a5cd4900fdd6a0be1c353cb4379128cacbe4b5943759c90337fb8563

Observation 491cae55-699b-4b4c-8fac-bd585b4a2d6b · outbound

This paper cites The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.740373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.740373Z digest=sha256:511362901cfe345300e766e0ee99793dcca7f62c4c858bf1a5579aff1d2d4b94

Observation fd0efa9e-8430-4bf0-97ec-4797e7591bbf · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.744660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.744660Z digest=sha256:07fe5458eb4e3953c6adc7d4dbf106ce8f627732c9a888bfd85189de068c7a7c

Observation 960b469c-553f-4d20-98e6-7ecb07302aac · outbound

This paper cites Neural discrete representation learning.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Neural discrete representation learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.749227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.749227Z digest=sha256:99c3c89e7eb4e4423b54d2e15aec0f1d80c1771d0d9c31df6df3e2487ec8c5b0

Observation 83537652-d13a-446b-bc96-a0070a2c828d · outbound

This paper cites Attention is all you need.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Attention is all you need

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.753297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.753297Z digest=sha256:cb02ab3abdc26410272d119d5436a33a64471df7ce678097bcc9a0735bbf0824

Observation d34be66d-e229-47ff-a52c-457d02c9aa2c · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.808818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.757666Z digest=sha256:a3601e17eadd800fd04b50d1e26bb871ed7d26e9acffa318860983602c2585a3

Observation 163393c2-c867-4746-b125-fe627ac4467d · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.761884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.761884Z digest=sha256:971533c7f63cb447c72c95e8843c0cade35ebceae993d680b201b153b89cb324

Observation b44d84d4-925a-4bfe-9cde-451247cfef80 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.766328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.766328Z digest=sha256:ec32c4888e879b9484400fe06a503d1c5995fe98b2ffc07aa52b3ceed07d03ce

Observation 387d996f-4d08-4dcf-861a-6178b55bc93d · outbound

This paper cites VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.770915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.770915Z digest=sha256:ae4546b973813302d24ee10448a7e22f7f86331335008071738384eb3de949d6

Observation d96c1a35-2d2e-476a-9f0b-dc1474778e8c · outbound

This paper cites WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.775414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.775414Z digest=sha256:cf6f51849b193612c64a4cd26875afa88eee3206462f6dd63ace1372b0876311

Observation c7e51422-5846-4ab5-a9af-00b0ac99a83f · outbound

This paper cites E., Chen, S., Tang, M., Liu, S., Li, J., and Yoshioka, T.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training E., Chen, S., Tang, M., Liu, S., Li, J., and Yoshioka, T

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.793957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.779739Z digest=sha256:693cfdfbc3e231f2628fffe3662c81e0ac427a5531b787468d44c83727d9dd36

Observation f4d8a91e-e789-4371-b7d5-103e43f7811c · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.783977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.783977Z digest=sha256:c811761635ec84b4f44bfe3e0051a960ea59465a801906807d6258f1d216862e

Observation 2f7cd87c-be18-42aa-a4cf-c96eed0bddec · outbound

This paper cites Lm-vc: Zero-shot voice conversion via speech generation based on language models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Lm-vc: Zero-shot voice conversion via speech generation based on language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.779298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.788526Z digest=sha256:321750bfbe01c401687aa0ece7fd23cf65608fd53e64eaa3da5a9c3ca119697b

Observation a95679b0-4c62-457f-bd4d-1bd5ed94b882 · outbound

This paper cites Selm: Speech enhancement using discrete tokens and language models.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Selm: Speech enhancement using discrete tokens and language models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.764011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.792763Z digest=sha256:1b2e51313a330f9ca96d90e1e36034661fdc35cf0ea6695430815bd1a8b82178

Observation 273d08e5-d991-4608-a813-f5c7d9dd72f2 · outbound

This paper cites Tf-gridnet: Integrating full-and sub-band modeling for speech separation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Tf-gridnet: Integrating full-and sub-band modeling for speech separation

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:54:19.749522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:54:18.796775Z digest=sha256:d88e843b267519517187bddf56823bdc2b44d2d4ce71ae0ed700039e7896c02a

Observation 82ca947b-6ee3-4c70-bb17-d2c49b3a1a0b · outbound

This paper cites WHAM!: Extending Speech Separation to Noisy Environments.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training WHAM!: Extending Speech Separation to Noisy Environments

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.801094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.801094Z digest=sha256:69bd531b4aa8b6940daa82ef115a42cf1ec0ca2b323dbe870d65476dbde97e87

Observation 487492ea-0bfc-4fc3-974b-06fc15d72545 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.805912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.805912Z digest=sha256:bed9eded786710bc556f3fa307a15c6714c1cc19e6529036007b1122df01cb27

Observation a78a21a8-abbf-4ac3-9259-c77b9d0d1bf3 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.810390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.810390Z digest=sha256:83a19431f7feca69f9e9e373815d7e7b35261e0c678c81ab500a73a6bb2bb8b4

Observation d54158eb-a70f-4b1b-8d1a-3de6e8727181 · outbound

This paper cites G., Yang, M.-H., Hao, Y., Essa, I., et al.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training G., Yang, M.-H., Hao, Y., Essa, I., et al

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.815060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.815060Z digest=sha256:7b965de00a53b790ee50057cbec8119d09dada10f2c126cbee994019bcfeb42b

Observation 75875b51-e68a-4244-8ca0-e4b82c4592d3 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.819238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.819238Z digest=sha256:8529aad9086e288e46d1bb8e224dfa2aeb879ac441dde59ed343dbadfddf30c3

Pith citing papers

Observation ffb056b2-5e5a-4357-a714-17287e0219e7 · inbound

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline cites this paper.

SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:07.252044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:07.252044Z digest=sha256:accba614bea84bf38ec8612a00165379d779213a8f4e4ce7dbf0dd8a865fb62c

Observation 59be1fe7-50bc-4cd0-9420-a896d77f6338 · inbound

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model cites this paper.

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:49.151323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:49.151323Z digest=sha256:b15f89184b5592d32bc5c08c726591ca143b9a63c6afd66a00476e46cae3c09f

Observation 5be49d1f-980f-499b-93b9-695450da0e11 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.898806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:3aa6e1668b192c86fc1fed9c926bad8c89a8b7f9562ff787851c6e83e602f76e

Observation fe196cd4-f51f-4a5a-9ce6-6f3312196439 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 203

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.149531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:ccac9e7b891d2cb8bcc427792980c5debfbea3437f202ea82343b0d24b51cd7e

Observation a6be6d39-106f-4b77-9a27-073f67aa5b06 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:aef0d54995300218e09964fee95d659343542b57600ee915e12ed9228fe9e559

Observation cddd88ed-459f-40af-b8f9-d782ac39fceb · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T06:45:43.330341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:45:43.330341Z digest=sha256:a8cfc77edd12766c95b81d55d6a044284af3c98247716d5ef32ce9fbb93ad203

Observation 31f23d98-8ad1-4700-a8d9-f6089db8771c · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance Metis: A Foundation Speech Generation Model with Masked Generative Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T07:05:05.797563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:05:05.797563Z digest=sha256:8d8fc0ade309a62bae1ecbbc7a3783caa495da174cdeec9e00c8a0a15406fb8b