Pith. sign in

Paper Citation Record · LEDGER

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

As of 16 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 0 inbound Pith citation observations for arXiv:2608.08667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08667 v1

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:32:48.158348Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

78 of 78 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2b59e89-d15c-4fee-bf0d-ad21db2c58f6 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies SoundStream: An End-to-End Neural Audio Codec

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.614160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:47.839562Z digest=sha256:5ee53a7cc3b4cf651a3c502e17b6b9d692fce9e8b4b6b10dbf219e5269d2de8c

Observation cfb1e325-496f-4b72-ac7c-cf182fc1f6fb · outbound

This paper cites High Fidelity Neural Audio Compression.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies High Fidelity Neural Audio Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.844577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.844577Z digest=sha256:1b782bf29cfef04c9ee02aa554d5df4d7aab6ce3ad05cc0f9552abcc5833be76

Observation eedf5e07-1a5f-4115-ba13-f91b73f997d6 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.848465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.848465Z digest=sha256:1f44a9819e2f3ec2e00e992cf4817a4137637accf2184d08bcc48146d3cf7bba

Observation 1452ac3a-e68b-402c-a1ba-a499c7043767 · outbound

This paper cites Simple and Controllable Music Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Simple and Controllable Music Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.852011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.852011Z digest=sha256:4423ace9b2174cbfff24aff97842c2161a79e4c9eae7a21db373f45401dc4d5c

Observation 7e491632-8991-4cc2-ad95-7b3ec1c2ca39 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.857262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.857262Z digest=sha256:0831368e74e60ce3cfb0bc4556f766204ac77c629205134b350cea67199191a4

Observation e8ddbbf9-7549-4140-81f6-76fd7ee31168 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Moshi: a speech-text foundation model for real-time dialogue

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.861238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.861238Z digest=sha256:e2f478261f80c9fec71c1925c505495684d999eaca49e507b28e7cd9d3ebf8ff

Observation b6e4639d-ccf9-4e20-b76d-7e15b40a3572 · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies AudioLM: a Language Modeling Approach to Audio Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.866569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.866569Z digest=sha256:34a4ab8da95a3a5bf03926e5a3528670cfa91d647608b08a9a7d18059b7ebe47

Observation 680aedd0-803b-4fae-9647-1e979396ba00 · outbound

This paper cites DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.871352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.871352Z digest=sha256:2a70bc2dc4bcb27b325eb8596c3bc8a4e472e571fb85a9d89b97a5a89b1eff52

Observation bba90d26-910c-41ad-b9f9-90b46d3d7161 · outbound

This paper cites Continuous Audio Language Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Continuous Audio Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.874790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.874790Z digest=sha256:330d8c57a6b5136549465531f917c3b8414a1ca8ee2179d35140b7ceda53c257

Observation afa56b14-ca97-4c73-934b-d52173cce56f · outbound

This paper cites Cover and Joy A.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Cover and Joy A

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.878575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.878575Z digest=sha256:f738ff851ced2f1a15bccefd7fedd314739561feac121571748a29eb7e0a7503

Observation c7c51330-02ac-4a8d-b75f-af6363632c44 · outbound

This paper cites UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.882061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.882061Z digest=sha256:85807d0cb40b8974aaaf950634fd8863bd35ebb20bb2c065bd0baac784266b61

Observation aa80f50f-65e4-460b-88ca-a6960e769793 · outbound

This paper cites The Information Bottleneck Method.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies The Information Bottleneck Method

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.491528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:47.886557Z digest=sha256:861aeb89d1bd3894269143dffe053bfc35b684e22853f1359a5faba267eafa23

Observation 1eead16d-e552-4c8d-8eb4-92aaee82e972 · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies High-Fidelity Audio Compression with Improved RVQGAN

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.482247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:47.890022Z digest=sha256:a4aa57df19d0779f7770cdf194a1bf62208cf64072b1952dc7e3db88c64fe9fe

Observation 3fbad04c-68e5-4345-b72c-69ed12b0b7cb · outbound

This paper cites Discrete Audio Tokens: More Than a Survey! https://arxiv.org/abs/2506.10274, 2025.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Discrete Audio Tokens: More Than a Survey! https://arxiv.org/abs/2506.10274, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.893232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.893232Z digest=sha256:6afbc7f30965d82da6623caf5827199b64e57a341cf08235ce9a67cd25cb5b55

Observation a8b297a1-93e5-4c46-bbc7-961cfe1ca123 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.896329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.896329Z digest=sha256:5047ec0755fc2f09be388543dd668e1a71fc6e572373e9e939cb3e3769542329

Observation f76e66a9-2689-4c01-92e5-6c5640ddb264 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.899732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.899732Z digest=sha256:e353fbcd3e28fc6fa180ea53e1e041df9556203279d611ee5c043f3d9d3f30f7

Observation 418150a3-dc47-4c54-890e-6f658235067d · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.903185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.903185Z digest=sha256:985e2c40f8792ffab1f29d8728a046125a701124b1537268927a9499c2534809

Observation e241e6a3-07c9-4bc7-a7d6-21f9e446c9f6 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Finite Scalar Quantization: VQ-VAE Made Simple

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.906647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.906647Z digest=sha256:1ee598eddeeb04bd59b9bbe140c63355a9b391541fe4887024a4759d2491374e

Observation 6133a4c6-3614-43f3-8784-4bcd3ebf1084 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies SoundStorm: Efficient Parallel Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.910377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.910377Z digest=sha256:b4af3ea7306a7dbff62a0ac2f09b6e5a550f4d629f9c646bd131f8dcf3e50972

Observation de8e945f-6794-4d3c-8279-d13c0ea79703 · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.913437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.913437Z digest=sha256:6604be556688606efa28526ba8ae0586daed12da0e2f70b5c4401f0a3c54349a

Observation ce1932c2-443a-457c-b121-0823992a3ee5 · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.916888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.916888Z digest=sha256:d585994e6296f24cf788df6a9df920295ff56a64d3026016d9dabf849654b7f9

Observation ce5fab92-d7ab-452b-8292-797e3f8cf7d9 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.920580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.920580Z digest=sha256:0563139e0652ffb72c0c93c18e866091c81cb8ada370bf4274692d016a8920e5

Observation 2fab2696-b580-4295-94d8-0f055aad8f63 · outbound

This paper cites SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:32:49.061753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:47.923864Z digest=sha256:383c5dc9aaa69563c65a8b503b66ac512ce63f6c5164a8c90bc0ceb3deff0eff

Observation a1c8901e-0f0d-4fb5-821c-77d2bf1fb045 · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.927340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.927340Z digest=sha256:ed32d93d3aa3e043208514c02bad6d39e648300363c6b2248286fe1859c1fd00

Observation fe5adde9-4640-45d3-b736-5a4b7939f957 · outbound

This paper cites Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.933751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.933751Z digest=sha256:b2f6cbd377e71b55f4f69054606df0288c518e8fb7ab93937afb5de5cb0d4bf9

Observation d49c1e6b-732d-4b39-9099-322895f7ec7b · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.948785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.948785Z digest=sha256:a01f11844ac57b9f5cc7a1bf4f9fb08286cb742b5bb6b2d5e8d805e1c1b92bac

Observation 0bde1496-b9f9-4427-b04e-13ff8dd373a2 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.952821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.952821Z digest=sha256:1ff8a2bd6bd027251505a0cbb8608fed1c221803fef31a70e48c55ab503d7758

Observation fc6cf0c3-0130-4677-bd7b-856b8021eaa1 · outbound

This paper cites MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.957531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.957531Z digest=sha256:2f12e9ab17ab7c6acfc0c865a4bf1ea80403cd74f9d4f8a88371c3ad3d477636

Observation 5c749d68-1735-49ff-8fa5-439b9cdb195d · outbound

This paper cites ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.960877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.960877Z digest=sha256:0da255cbbe96d1a362c0d7924419a963dea4806f1f8b0ffe7b4e60fc0bc3dba8

Observation 8cc305a7-31ff-49e2-aa1c-8c4aac5e2a48 · outbound

This paper cites Fewer-token Neural Speech Codec with Time-invariant Codes.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Fewer-token Neural Speech Codec with Time-invariant Codes

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.472237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:47.964418Z digest=sha256:adc7cb205eee4cb07ade69a60779cbd3a7736821c839b5a4657a80b77b989110

Observation 710342e8-7526-40e8-b051-35247d38dce7 · outbound

This paper cites Learning Source Disentanglement in Neural Audio Codec.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Learning Source Disentanglement in Neural Audio Codec

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.967536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.967536Z digest=sha256:f9b60f754d064fef305a167ce58764ba89152069f0f7303862f9f6858f3ac120

Observation 0d93c519-337d-4f78-aefc-0f3d941c7728 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.972389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.972389Z digest=sha256:e55810269e3154e5c8addc9c64e00df882fcb8fb760c2a57bcaf7e3ba992b9c6

Observation 7b7d3346-a2e6-4982-a0b7-84125d828276 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.977209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.977209Z digest=sha256:ce53c0494fed9f469618c467fc23056dc568a375fa4f3c872a8e2ddc6eee7bd1

Observation 3c7dffe6-19b0-40ab-ab93-6df36700500a · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.982218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.982218Z digest=sha256:6eea61ad0243b609f7761a31419c62151372601682aa4f1e86c04d162ba78184

Observation 033e1daf-d353-4952-8593-68c7cffb7fd9 · outbound

This paper cites Stable Audio Open.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Stable Audio Open

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.986595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.986595Z digest=sha256:3aba89a9dfd70739d4fbfd69d31c47e8abb832120e3d6eaff51a3353de73254f

Observation 74619dc8-a967-49db-b4d2-5365e11688ff · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.991814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.991814Z digest=sha256:7b05f115b4ce8a7ba2b28c0f8b379695017258fcdf02d028fea25a3fe9924d32

Observation 02d37584-975e-4631-9d6e-7f81ba866280 · outbound

This paper cites Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.995431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.995431Z digest=sha256:2f47a441598bea4afbd6b9c090810badc61ea29c02ac4aa5cc24eb695f0b6ceb

Observation 26ae9252-f018-422c-9de0-40b23401c289 · outbound

This paper cites DashengTokenizer: One layer is enough for unified audio understanding and generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies DashengTokenizer: One layer is enough for unified audio understanding and generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.998963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.998963Z digest=sha256:8bad01486e3915190ffa123a03b265dd6b67339fdb30632bce6732a10f710928

Observation aed3a8ba-eff8-4564-9320-224800260784 · outbound

This paper cites WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:32:48.784735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.002062Z digest=sha256:86a7150e6b2ab4301b5d1f924831660d3fd13eea7fad57659e049deb21ea3bac

Observation 2eecf9e9-4ac7-4739-aef3-27d356176657 · outbound

This paper cites Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.005428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.005428Z digest=sha256:85bd1fd2816604a8566d6592b074c00250d4165d5522216c63cfa71834086f7f

Observation 40e40133-7778-4238-a3be-cf5ab78ed5ca · outbound

This paper cites LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:32:48.735854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.008835Z digest=sha256:41cdeec52cd15498257f32a08bdce773a1bbe50040f86eca85a6ed79c0d1b77f

Observation 7b5a52d9-2c0b-48f0-9dff-202af3ada8be · outbound

This paper cites Flow Matching for Generative Modeling.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Flow Matching for Generative Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.013264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.013264Z digest=sha256:f4983406e0e2ddb6598ba725dd6ba3c41d91721c54ff2ce9b1023482b6a2a224

Observation bbde2eef-b8a9-41e2-9698-44f2f76c5e36 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.017504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.017504Z digest=sha256:c4d81d36bb9a68b3fc6fb51684c94f0a9383baea55a9fc4907b9897ea8f11d56

Observation ee5b1124-7fb4-4fe3-b88b-661b823849cb · outbound

This paper cites GIVT: Generative Infinite-V ocabulary Transformers.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies GIVT: Generative Infinite-V ocabulary Transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.459255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.021121Z digest=sha256:d9380c281be0f43a21db5e48d8b55177098421f8b652f7116ac673c2aea735bd

Observation 68a66f64-d066-4465-8ea1-a405c4fcd8c5 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Autoregressive Image Generation without Vector Quantization

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.450189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.024464Z digest=sha256:b88fcfbefda52a8ae99e5b01223a6b2dcd4cef63716811cc6a3d6e7af80c1725

Observation 91e01e7c-9497-4c35-9fc3-4ad627ebc00b · outbound

This paper cites Hyperspherical Latents Improve Continuous-Token Autoregressive Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Hyperspherical Latents Improve Continuous-Token Autoregressive Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.027469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.027469Z digest=sha256:57184a69fab9537eb7fb8aec339b61233144a00f1b8221f5f17b45ea877dad69

Observation e0b217a0-0fd8-4229-af2f-fb3244eee52d · outbound

This paper cites MaskGIT: Masked Generative Image Transformer.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies MaskGIT: Masked Generative Image Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.037218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.037218Z digest=sha256:d1ca9772b4bf71141d8f58d85dd2126447ea019311ced6bcb1122248ee248da5

Observation 0cb337cd-8f82-4908-8d8b-8a0ec5dd4e6e · outbound

This paper cites Denoising Diffusion Probabilistic Models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Denoising Diffusion Probabilistic Models.Advances in Neural Information Processing Systems, 33:6840–6851, 2020

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.046526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.046526Z digest=sha256:744e5b191455a89cf5ecc3ca5f3a1838bb9b08c650dc4746b917f9fc132709c6

Observation a2e3b1db-8c91-4c9d-b78e-28678519f5bb · outbound

This paper cites Generative Spoken Language Modeling from Raw Audio.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Generative Spoken Language Modeling from Raw Audio

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.049500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.049500Z digest=sha256:234cbcb3bb3aeee4a99838887a13f2deb027187fb00c2ccbdba15ed90f0dd94c

Observation 9c6e7a43-219d-442a-a9be-318988926923 · outbound

This paper cites Textually Pretrained Speech Language Models.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Textually Pretrained Speech Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.053372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.053372Z digest=sha256:7f531fdd7af22ce28cff54d56c8efc8a3f627849bc1bbb729d1a153981b2766b

Observation 7af9cea4-6af8-4ac6-8bbd-8c18142e4867 · outbound

This paper cites MusicLM: Generating Music From Text.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies MusicLM: Generating Music From Text

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.056851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.056851Z digest=sha256:cc9b5c7d43c38cdcfb702281788259052b4503cef42fde884f45a85fa83d65b2

Observation 919576c7-bfbd-4e15-811f-af3c294d908c · outbound

This paper cites Diffsound: Discrete Diffusion Model for Text-to-sound Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Diffsound: Discrete Diffusion Model for Text-to-sound Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.060493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.060493Z digest=sha256:46ac5a0d7ccde3e9dccdb819f6af93f9649781ab6f655addf9bb3e7f2c455cee

Observation 992b07fb-5507-4222-ac29-6a397499279c · outbound

This paper cites InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.063945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.063945Z digest=sha256:edae58f2d61628735ffa845354ffe25db3ed99c8d4ce948c7f723e93d84928b6

Observation 0e6decfd-b6b7-4b0d-9cd1-16704eff6bf9 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.067554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.067554Z digest=sha256:c892c272bb716af0fd2c206c07a1cd0925eccf92a5bbad49a608669084d9734a

Observation a0bde76f-c503-4a1f-9568-0732554a0d40 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.072180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.072180Z digest=sha256:b2ace947902ed1c40a4028f919b9aa741c1d008593592f72613f88448900a7f4

Observation 3592ff85-663b-4f84-b3af-417b3b73eb0b · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Autoregressive Speech Synthesis without Vector Quantization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.434571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.075627Z digest=sha256:175716a30a6a15231635cb79e9fec12c3b2822be44f0380ab659151aa2e96c19

Observation afc55e3d-a429-488d-a0ce-a9232926b084 · outbound

This paper cites VibeVoice Technical Report.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies VibeVoice Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.078597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.078597Z digest=sha256:db721fe71b3df198734ba55707317de25cef3b2b10fba9934731b7bb12af6119

Observation 14f27a22-bbf6-415a-801c-83a19178a63f · outbound

This paper cites DASB - Discrete Audio and Speech Benchmark.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies DASB - Discrete Audio and Speech Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.082086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.082086Z digest=sha256:0a650a36ccd665dbbde39956c83365caa7840598685054d13f0475ba0678231f

Observation 695ee055-b283-4777-965a-4ecb5687dd40 · outbound

This paper cites TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.085412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.085412Z digest=sha256:df3ea8997c3c29e786e1802a52cae2c9266705ea3b119cbd4233447f71d2b593

Observation 586f54b3-7fb0-4dbc-9409-73fcce1625d8 · outbound

This paper cites FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.088518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.088518Z digest=sha256:ad5eb130141308114be4f8eae782415b882d4a291505439ba061b2544b45d959

Observation 3dac3314-cccf-4292-8650-377336324588 · outbound

This paper cites Signal Estimation from Modified Short-Time Fourier Transform, 1984.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Signal Estimation from Modified Short-Time Fourier Transform, 1984

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.426137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.091664Z digest=sha256:eae59fa535de5ecf510211c6c1a19fc462c939b98554bdd15fc6d0c28d889496

Observation 90ad8a82-fbf4-4465-9083-21d17a99227b · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies WaveNet: A Generative Model for Raw Audio

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.094769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.094769Z digest=sha256:7117c8493156fabed90786669df0c393cc97848a888d92da6f67e597f294ba29

Observation fedec146-001e-4fd3-a522-7f5592eb1cef · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.098151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.098151Z digest=sha256:ef93a090553bb2794d6696bc5ffad34aca27e75a1b31b9a24ccbbea14fe8e86f

Observation 12333eaf-8c9b-4265-8eb5-e5a2924e9724 · outbound

This paper cites Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.101601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.101601Z digest=sha256:2917b044100ecbe2a4e63a1020cca4f76e48b908a5484a941ff4671ecbf4b35c

Observation d3082628-f791-459d-a9f3-4393e90dbd7d · outbound

This paper cites Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.105338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.105338Z digest=sha256:0c0106fd2f320fc0039ee6fe5d20a6056997752198421083e073a1ea9f8c4b29

Observation 3ebdef5a-ff7f-4971-8a60-b3fcdc44870e · outbound

This paper cites A Scale for the Measurement of the Psychological Magnitude Pitch.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies A Scale for the Measurement of the Psychological Magnitude Pitch

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.108713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.108713Z digest=sha256:c86c9a43e960eb9ee4e9a45e89fb95795719eda2cc975249f094932529d6c555

Observation 6a6fa749-8f53-4560-bb0e-e2523e919cb4 · outbound

This paper cites Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Comparison of Parametric Representations for Monosyllabic Word Recognition in Continuously Spoken Sentences

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.114746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.114746Z digest=sha256:ce6092c37d998a2687c357816589b9905bc6de57e43acb610aba0606c54189fa

Observation b8d7505e-0465-4e84-a12b-19c646290737 · outbound

This paper cites WaveGrad: Estimating Gradients for Waveform Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies WaveGrad: Estimating Gradients for Waveform Generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.417145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.118320Z digest=sha256:c8d12c8938d9db4643e0b1c3d311752a4381565fa1e9418ed738ccc77e3c4946

Observation 5b0935a2-50dc-4572-a1b2-e851afb857f5 · outbound

This paper cites DiffWave: A Versatile Diffusion Model for Audio Synthesis.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies DiffWave: A Versatile Diffusion Model for Audio Synthesis

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.407635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.121489Z digest=sha256:b2b0878e4be7ad9838f4dd7c1011787d836bc8436d04469e7b67a40ce0b1d98b

Observation a72bbb14-e7c8-4620-a0b5-4e100feb541b · outbound

This paper cites dMel: Speech Tokenization made Simple.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies dMel: Speech Tokenization made Simple

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.124613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.124613Z digest=sha256:3901473ef6039668da888f7e4354aa4ce72ef3945ff0ceea5690cbdb03d4db7a

Observation b3fbabab-de1d-447a-af47-08297b6e5b09 · outbound

This paper cites Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.127905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.127905Z digest=sha256:6188e9e9ad4cb35a997f64d6a35d356c983b5edce9a0fa5625d5125b8293cf6a

Observation 9781d73b-e82d-432b-877d-83190c2d1675 · outbound

This paper cites DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.131468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.131468Z digest=sha256:b1a6ee78286f59ae08ca4ae35adef6cf34697b8feed63a684bff6cdb3a465b2c

Observation e800c94e-5a30-402a-bb39-9e4aaf69d7ca · outbound

This paper cites Sound Texture Perception via Statistics of the Auditory Periphery: Evidence from Sound Synthesis.https://doi.org/10.1016/j.neuron.2011.06.032, 2011.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Sound Texture Perception via Statistics of the Auditory Periphery: Evidence from Sound Synthesis.https://doi.org/10.1016/j.neuron.2011.06.032, 2011

Reference 73

Resolution
verified exact
doi, observed 2026-08-14T04:32:48.194458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.135457Z digest=sha256:9c756361235498a9655db244c80ca1971e08033c209dc301dc64e37dafb17b6a

Observation 454ad516-7a76-4550-b4c9-12049203e106 · outbound

This paper cites Neural Discrete Representation Learning.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Neural Discrete Representation Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.140267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.140267Z digest=sha256:64ae93885336a892059082596b058ceada1c225be2a4534e2d24bfdbdb6aeda8

Observation 4bc48978-6e4c-4923-b962-f8b1e1fb50c9 · outbound

This paper cites SNAC: Multi-Scale Neural Audio Codec.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies SNAC: Multi-Scale Neural Audio Codec

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.145931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.145931Z digest=sha256:c1aef979d1e681690adba3271b2075379e0ea3bd6bf7808bcfc77e6701d7a640

Observation 78417bdc-e7b6-4707-afd5-1877e1372cdd · outbound

This paper cites An Image is Worth 32 Tokens for Reconstruction and Generation.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:32:49.398032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T04:32:48.151323Z digest=sha256:8ef80e9a0d41ef21b46000da89b9de2ae8396c9a82afba0bbcc4a57773cd5917

Observation 9600fbcc-9f63-44c7-a359-952eb896998d · outbound

This paper cites FlexTok: Resampling Images into 1D Token Sequences of Flexible Length.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies FlexTok: Resampling Images into 1D Token Sequences of Flexible Length

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.155154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.155154Z digest=sha256:c7c8b8e1b52c5ef1774c232493b2149b9d5772d9d8f53356a45aaf6d4837524d

Observation 7282f635-559e-47c4-891b-d17a25385c20 · outbound

This paper cites Variable-rate discrete representation learning.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Variable-rate discrete representation learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.158348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.158348Z digest=sha256:43366f1400df7bb7d3ed6c76797dcb06a55b69df2cd2a71e875b20769c556b5b

Pith citing papers

No inbound Pith citation observations are available.