Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.494016Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 258 outbound references and 47 inbound Pith citation observations for arXiv:2411.13577.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.494016Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:07.865325Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
100 of 258 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 9c26b626-7630-461b-a1f5-3fe009b28cd0 · outbound
WavChat: A Survey of Spoken Dialogue Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c2d8f8-9a71-4eab-b4b4-1096737dc39c · outbound
WavChat: A Survey of Spoken Dialogue Models MusicLM: Generating Music From Text
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ecaa38d-9edb-4a0e-9628-3526b7d01beb · outbound
WavChat: A Survey of Spoken Dialogue Models HILCodec: High-Fidelity and Lightweight Neural Audio Codec
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 002ca7ce-610f-4862-a914-196ffd609c03 · outbound
WavChat: A Survey of Spoken Dialogue Models APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9d8ba8-8bc2-43ed-b261-2854fb0377fc · outbound
WavChat: A Survey of Spoken Dialogue Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d0a9442-fe6d-4354-83fb-95ad0a456654 · outbound
WavChat: A Survey of Spoken Dialogue Models PaLM 2 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b9cc61-d310-4e55-9554-f66189e4215b · outbound
WavChat: A Survey of Spoken Dialogue Models SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956ed686-9594-4924-b517-422a317da9f2 · outbound
WavChat: A Survey of Spoken Dialogue Models Common Voice: A Massively-Multilingual Speech Corpus
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a393ab7b-61a5-490b-98a7-7beda0ca40bf · outbound
WavChat: A Survey of Spoken Dialogue Models XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e28eb88-9e4c-4b9f-9016-4a6e2e8443a8 · outbound
WavChat: A Survey of Spoken Dialogue Models wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315dc7ec-e5eb-4f92-b4a9-501bac386e64 · outbound
WavChat: A Survey of Spoken Dialogue Models Qwen Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a12570-d991-47bd-bf60-b71f05c3238b · outbound
WavChat: A Survey of Spoken Dialogue Models An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76dcd302-7a21-4582-a493-5dc93222cddd · outbound
WavChat: A Survey of Spoken Dialogue Models Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6800971-c408-481c-8671-963e5dd4fe91 · outbound
WavChat: A Survey of Spoken Dialogue Models Seamless: Multilingual Expressive and Streaming Speech Translation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1daf4044-5c06-4330-8d8f-c482137b6063 · outbound
WavChat: A Survey of Spoken Dialogue Models Curriculum learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab89427-e657-4506-9ccc-2cebbf9860b7 · outbound
WavChat: A Survey of Spoken Dialogue Models Medleydb: A multitrack dataset for annotation-intensive mir research
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2af126a-f211-4433-8a83-ed71a03ba5d4 · outbound
WavChat: A Survey of Spoken Dialogue Models A comparison of sound segregation techniques for predominant instrument recognition in musical audio signals
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4765c84f-026a-49d4-966a-021719ab9b7e · outbound
WavChat: A Survey of Spoken Dialogue Models Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258ccbfc-8bbf-4433-b616-92d0b2f50e44 · outbound
WavChat: A Survey of Spoken Dialogue Models Iemocap: Interactive emotional dyadic motion capture database
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2fb2d2e-336a-42d4-bf86-be0f562b8128 · outbound
WavChat: A Survey of Spoken Dialogue Models Msp-improv: An acted corpus of dyadic interactions to study emotion perception
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b44d82e-5b4b-41e4-b5dc-22689577b166 · outbound
WavChat: A Survey of Spoken Dialogue Models XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 018c40ae-3892-4bec-b1e2-e4c4d8558bbc · outbound
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceeb780d-66d9-4dca-a1e0-b4cfcddc395c · outbound
WavChat: A Survey of Spoken Dialogue Models Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db6cd1f-8b9a-4688-a0af-bcaaf51a1b75 · outbound
WavChat: A Survey of Spoken Dialogue Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c02b3be4-debd-4f67-8ee1-6d6aa4965463 · outbound
WavChat: A Survey of Spoken Dialogue Models EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7b9dba-156d-462e-9bf3-948ea0e899da · outbound
WavChat: A Survey of Spoken Dialogue Models Evaluating Large Language Models Trained on Code
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2a1efb3-59e4-437e-ae96-2947bae9c2b3 · outbound
WavChat: A Survey of Spoken Dialogue Models Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df371f9-bab8-4f0b-931c-14d381ef70e5 · outbound
WavChat: A Survey of Spoken Dialogue Models BEATs: Audio Pre-Training with Acoustic Tokenizers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b47d8fc-ad34-4938-87f1-403ed37ecbb8 · outbound
WavChat: A Survey of Spoken Dialogue Models VoiceBench: Benchmarking LLM-Based Voice Assistants
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c645ca3-7c0a-47bc-a393-ef41920af236 · outbound
WavChat: A Survey of Spoken Dialogue Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef5c41f2-bef0-43b9-8800-83c3b39d8a6a · outbound
WavChat: A Survey of Spoken Dialogue Models Audio albert: A lite bert for self-supervised learning of audio representation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac48763-3162-485f-909a-b2c4aed89fd4 · outbound
WavChat: A Survey of Spoken Dialogue Models Coding Speech through Vocal Tract Kinematics
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c593901c-4945-49fb-aee4-e89607ae5845 · outbound
WavChat: A Survey of Spoken Dialogue Models Qwen2-Audio Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf343a91-b7a9-4c22-ad3d-d1e7b00db920 · outbound
WavChat: A Survey of Spoken Dialogue Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be047d4b-f468-479f-84da-bd1ead5288e8 · outbound
WavChat: A Survey of Spoken Dialogue Models Vector-Quantized Autoregressive Predictive Coding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed8bf43c-20bd-4ecf-a844-feff12ec830e · outbound
WavChat: A Survey of Spoken Dialogue Models MusicRL: Aligning Music Generation to Human Preferences
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce81594f-53ac-49bc-acf0-c7e9a3e76d4c · outbound
WavChat: A Survey of Spoken Dialogue Models The fisher corpus: A resource for the next generations of speech-to-text
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1da52ec-8c73-41a9-868d-68f46c881fcc · outbound
WavChat: A Survey of Spoken Dialogue Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28f2a5d-7bd0-4402-ac24-8f356ba2676b · outbound
WavChat: A Survey of Spoken Dialogue Models Training Verifiers to Solve Math Word Problems
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b36e97-f652-4c13-bb1d-f662964f8706 · outbound
WavChat: A Survey of Spoken Dialogue Models Simple and controllable music generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d1dd1b5-327b-4a08-8f2b-1ff04dd74303 · outbound
WavChat: A Survey of Spoken Dialogue Models SpeechVerse: A Large-scale Generalizable Audio Language Model
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ec07a2-e710-4c85-902a-ca44a7fbba2b · outbound
WavChat: A Survey of Spoken Dialogue Models FMA: A dataset for music analysis
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a93b1fac-b5e3-48d3-8260-ea26665cacb8 · outbound
WavChat: A Survey of Spoken Dialogue Models High Fidelity Neural Audio Compression
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29fb114e-c123-4060-8e32-bf2cf5c3210b · outbound
WavChat: A Survey of Spoken Dialogue Models Moshi: a speech-text foundation model for real-time dialogue
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cefef90-71e3-40e3-a52f-0be6618637f6 · outbound
WavChat: A Survey of Spoken Dialogue Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a5a61f-24bb-4d91-9549-a875aab4479e · outbound
WavChat: A Survey of Spoken Dialogue Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b4d91b7-1791-4e0c-8b0f-a307743f6348 · outbound
WavChat: A Survey of Spoken Dialogue Models Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ba73b8-d666-4736-81fb-0eb8a5cd1376 · outbound
WavChat: A Survey of Spoken Dialogue Models AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ce1a66-22c5-4ec6-b4da-662fd9ea569f · outbound
WavChat: A Survey of Spoken Dialogue Models CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a316dd0-6089-4e12-b772-fe85c4600f4e · outbound
WavChat: A Survey of Spoken Dialogue Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f229422-b4d8-4554-8113-4f578c56de24 · outbound
WavChat: A Survey of Spoken Dialogue Models Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb83fbb4-cbb0-442d-94fa-42c7379c0d8b · outbound
WavChat: A Survey of Spoken Dialogue Models The Llama 3 Herd of Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40b767ac-f322-4d9e-983e-73d4f7a3b8d4 · outbound
WavChat: A Survey of Spoken Dialogue Models Some signals and rules for taking speaking turns in conversations
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcc02c13-c6d9-4d37-969e-037da153b3b7 · outbound
WavChat: A Survey of Spoken Dialogue Models On signalling that it’s your turn to speak
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b49f78fa-0bd8-428e-9379-b744193cd15c · outbound
WavChat: A Survey of Spoken Dialogue Models TurnGPT: a Transformer-based Language Model for Predicting Turn-taking in Spoken Dialog
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4670fd8a-6965-4992-afbc-6b8c907bc4e8 · outbound
WavChat: A Survey of Spoken Dialogue Models Neural audio synthesis of musical notes with wavenet autoencoders
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7753991-651d-4217-b129-458db8c2962a · outbound
WavChat: A Survey of Spoken Dialogue Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7443e8b-4bfa-4c66-a79e-5940a9a44655 · outbound
WavChat: A Survey of Spoken Dialogue Models MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c892ff8d-7257-4dde-9b61-a0165ba24c30 · outbound
WavChat: A Survey of Spoken Dialogue Models Meisd: A mul- timodal multi-label emotion, intensity and sentiment dialogue dataset for emotion recognition and sentiment analysis in conversations
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fbd5429-1d82-4448-8d25-882be6851d98 · outbound
WavChat: A Survey of Spoken Dialogue Models Fsd50k: an open dataset of human-labeled sound events
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d0aa0f-6dfe-4ed6-b2a5-2b9f75f34491 · outbound
WavChat: A Survey of Spoken Dialogue Models VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320fcf89-fc93-43c6-9d0a-23b137322614 · outbound
WavChat: A Survey of Spoken Dialogue Models A new algorithm for data compression
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 424a669b-2016-4d62-a083-84d59e362c72 · outbound
WavChat: A Survey of Spoken Dialogue Models The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2eef2e6-8a45-4db4-ac5d-13225d60b5d2 · outbound
WavChat: A Survey of Spoken Dialogue Models Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef5af30d-ea6d-4a93-8541-893a3a1f877d · outbound
WavChat: A Survey of Spoken Dialogue Models Audio set: An ontology and human-labeled dataset for audio events
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5e90a7-14b3-4374-b708-ab3f1a2d3636 · outbound
WavChat: A Survey of Spoken Dialogue Models Audio Dialogues: Dialogues dataset for audio and music understanding
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f44c55d5-6058-46b7-be64-741411121f5b · outbound
WavChat: A Survey of Spoken Dialogue Models Joint audio and speech understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bc3aeb-6d83-4f5a-bdd0-4ea1250cd62b · outbound
WavChat: A Survey of Spoken Dialogue Models Listen, Think, and Understand
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e5d15ca-1908-4d76-80bd-966b6e98ba46 · outbound
WavChat: A Survey of Spoken Dialogue Models Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cf570b-5c58-4d67-86c8-06a44039bacd · outbound
WavChat: A Survey of Spoken Dialogue Models SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9d32c8f-cb84-44a1-9072-da299d4e3e45 · outbound
WavChat: A Survey of Spoken Dialogue Models Prompttts: Controllable text-to-speech with text descriptions
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7b6ff2-74d1-46e2-a9fa-00471ee69a4f · outbound
WavChat: A Survey of Spoken Dialogue Models Prediction of turn-taking using multitask learning with prediction of backchannels and fillers
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9737d212-b214-4e07-ad82-87b2e567048c · outbound
WavChat: A Survey of Spoken Dialogue Models Turn-taking prediction based on detection of transition relevance place
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2d1c409-67bf-4a72-a12b-104db40e5015 · outbound
WavChat: A Survey of Spoken Dialogue Models Textually pretrained speech language models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d7304d9-8e79-476a-8f6e-2b93f738b65a · outbound
WavChat: A Survey of Spoken Dialogue Models Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f45b7f13-d2c4-477a-bccc-97b8b805163f · outbound
WavChat: A Survey of Spoken Dialogue Models Measuring Massive Multitask Language Understanding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c9351f-eaa2-49e5-8f47-6030e0f2cbca · outbound
WavChat: A Survey of Spoken Dialogue Models Cnn architectures for large-scale audio classification
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcebdad7-c22c-42f3-b06d-3063928ef62e · outbound
WavChat: A Survey of Spoken Dialogue Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4983beac-99aa-4b8f-be82-d5ce90e852b9 · outbound
WavChat: A Survey of Spoken Dialogue Models Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c156cbbf-a0d3-4c34-8ea9-10c02d575d6b · outbound
WavChat: A Survey of Spoken Dialogue Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800bb696-9382-4988-94b7-9b2b9178d651 · outbound
WavChat: A Survey of Spoken Dialogue Models WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68359b0-08d7-421c-ac43-232da744a15f · outbound
WavChat: A Survey of Spoken Dialogue Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78790e59-b9bc-4cfd-9456-50750a7d4d98 · outbound
WavChat: A Survey of Spoken Dialogue Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74190d8c-0307-4a0a-bc1f-7a279eae948e · outbound
WavChat: A Survey of Spoken Dialogue Models Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9636719a-5540-4c0f-b20a-eea6761db9b7 · outbound
WavChat: A Survey of Spoken Dialogue Models Audiogpt: Understanding and generating speech, music, sound, and talking head
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fac6b86-0a5d-4cc2-a969-cf23b331ccaa · outbound
WavChat: A Survey of Spoken Dialogue Models SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5802cc6-2358-49c6-9ae1-f78fe999229d · outbound
WavChat: A Survey of Spoken Dialogue Models RepCodec: A Speech Representation Codec for Speech Tokenization
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a72e59f-fcf6-4b85-be7d-a80f72e08bc0 · outbound
WavChat: A Survey of Spoken Dialogue Models Residual Quantization with Implicit Neural Codebooks
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297bc52d-5e5f-46a9-9a58-0ef872718c9a · outbound
WavChat: A Survey of Spoken Dialogue Models The lj speech dataset
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22c33a22-114c-4876-b8a3-46a984a0ee61 · outbound
WavChat: A Survey of Spoken Dialogue Models Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606ba9b5-f3b4-4a90-bb14-e3bb82890f64 · outbound
WavChat: A Survey of Spoken Dialogue Models MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd7dced9-e31f-430f-9d86-819999d292a9 · outbound
WavChat: A Survey of Spoken Dialogue Models WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162d83f8-3c4e-496a-9284-758657d4296e · outbound
WavChat: A Survey of Spoken Dialogue Models Textrolspeech: A text style control speech corpus with codec language text-to-speech models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7bd55b-9218-4dcc-9454-8da0c777e3c1 · outbound
WavChat: A Survey of Spoken Dialogue Models ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4339d445-0079-4168-abe7-9884393aefd6 · outbound
WavChat: A Survey of Spoken Dialogue Models CVSS Corpus and Massively Multilingual Speech-to-Speech Translation
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbdc066-4c28-4f04-949c-950e41ec9983 · outbound
WavChat: A Survey of Spoken Dialogue Models Mixtral of Experts
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d82b99ff-d4cf-4386-8c52-5efd5a7850d2 · outbound
WavChat: A Survey of Spoken Dialogue Models Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f123214f-8997-456f-9770-df9e795a12df · outbound
WavChat: A Survey of Spoken Dialogue Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63a201b-1a26-444b-89ed-9f5f2bd1ded3 · outbound
WavChat: A Survey of Spoken Dialogue Models Duplex conversation in outbound agent system
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad135ec-e442-419b-929a-63b5ea94ab09 · outbound
WavChat: A Survey of Spoken Dialogue Models Efficient multimodal large language models: A survey
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03dbb64-8ad8-47ff-8b52-f7946be435e7 · inbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners WavChat: A Survey of Spoken Dialogue Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85f77c63-dc55-4ddb-942b-4ab0570238d1 · inbound
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training WavChat: A Survey of Spoken Dialogue Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d119e6-23bf-4caf-a8a9-8ff3c6c41e7c · inbound
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models WavChat: A Survey of Spoken Dialogue Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac489d20-9269-4d47-9b1f-015ee9aef61b · inbound
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios WavChat: A Survey of Spoken Dialogue Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cae7078-9d07-49f5-b004-9cd8c21c982b · inbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction WavChat: A Survey of Spoken Dialogue Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5207cc1-6394-4a24-acb1-e3f238b866c6 · inbound
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey WavChat: A Survey of Spoken Dialogue Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ec01bb-3d89-4e32-b02f-f0bc7006d77e · inbound
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her WavChat: A Survey of Spoken Dialogue Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d76f053a-4f51-44c6-9c0d-6b8f7b739e43 · inbound
On The Landscape of Spoken Language Models: A Comprehensive Survey WavChat: A Survey of Spoken Dialogue Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ffc30f93-d563-4ea0-9351-857f2670ad54 · inbound
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis WavChat: A Survey of Spoken Dialogue Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd7ba0be-36c5-4455-8098-6c4bf2e123b0 · inbound
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs WavChat: A Survey of Spoken Dialogue Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24297cbb-faa3-4c2a-a25a-9cf84fe8d9b4 · inbound
Speechless: Speech Instruction Training Without Speech for Low Resource Languages WavChat: A Survey of Spoken Dialogue Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c3dccc7-2e75-4285-ad77-48527129a280 · inbound
BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM WavChat: A Survey of Spoken Dialogue Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c742b7-aa8e-4ea6-b906-76cf35b79278 · inbound
Breaking the Barriers of Text-Hungry and Audio-Deficient AI WavChat: A Survey of Spoken Dialogue Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44080555-c607-49e4-a760-a70df34c9de0 · inbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · inbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65c5f37-2a5f-4cad-b67c-2cbb5e689769 · inbound
OpusLM: A Family of Open Unified Speech Language Models WavChat: A Survey of Spoken Dialogue Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0594c1-2ab7-40ed-84a5-211ae59b4860 · inbound
UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding WavChat: A Survey of Spoken Dialogue Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8c503e55-8ba0-4943-9f41-c006f5503240 · inbound
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding WavChat: A Survey of Spoken Dialogue Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f140ba98-4637-4a27-b434-b70b112dc254 · inbound
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling WavChat: A Survey of Spoken Dialogue Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74403b0a-948f-405d-9850-ed7df164adbb · inbound
Dual Information Speech Language Models for Emotional Conversations WavChat: A Survey of Spoken Dialogue Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5288e3c4-2220-4f74-bd8a-ccf2e787a69d · inbound
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs WavChat: A Survey of Spoken Dialogue Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f744dbb-9fdf-44a8-9b65-209172234df5 · inbound
ChipChat: Low-Latency Cascaded Conversational Agent in MLX WavChat: A Survey of Spoken Dialogue Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab9084d7-ed4c-4928-add3-05e6c2c8fd8c · inbound
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models WavChat: A Survey of Spoken Dialogue Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b19cbf1b-6668-4905-b09c-0e80e2297a68 · inbound
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM WavChat: A Survey of Spoken Dialogue Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 271fc78e-d208-4431-9e80-c67457bd918e · inbound
TiCo: Time-Controllable Spoken Dialogue Model WavChat: A Survey of Spoken Dialogue Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 535b2d08-65a5-4aa4-ba7a-7e720b089a69 · inbound
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff WavChat: A Survey of Spoken Dialogue Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd73e9c-b565-485d-acf0-d87c0ab92ffc · inbound
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection WavChat: A Survey of Spoken Dialogue Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8ea7f99f-86af-441f-a610-0d9a707b88a3 · inbound
Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs WavChat: A Survey of Spoken Dialogue Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e8014a72-e25e-48c7-9938-bfde1e7c5e27 · inbound
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models WavChat: A Survey of Spoken Dialogue Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6de18e82-e85e-4747-9da7-4f588149be8f · inbound
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models WavChat: A Survey of Spoken Dialogue Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 36ee66e8-0806-429d-94e4-0367b3d78a91 · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing WavChat: A Survey of Spoken Dialogue Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7a2fed10-bd5c-417e-bed0-684bfb935262 · inbound
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades WavChat: A Survey of Spoken Dialogue Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6d257d6d-1135-4f57-aec4-5ffc20237160 · inbound
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades WavChat: A Survey of Spoken Dialogue Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e1a5b93-42aa-42d3-b65a-0f686d7bb436 · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook WavChat: A Survey of Spoken Dialogue Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 292c42a6-2bcc-4fdc-b4d3-5472e442d172 · inbound
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects WavChat: A Survey of Spoken Dialogue Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 56dd88e1-fd28-4878-aafe-63dad0995384 · inbound
UniVocal: Unified Speech-Singing Code-Switching Synthesis WavChat: A Survey of Spoken Dialogue Models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9e561681-2cfb-4359-be47-55f7f9d8f1e8 · inbound
Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models WavChat: A Survey of Spoken Dialogue Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5615766a-21f1-4d73-8a28-d4a93a13076e · inbound
Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering WavChat: A Survey of Spoken Dialogue Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fc803355-0644-4879-9499-af55bdc620cc · inbound
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0a6e3651-30ce-49cd-a0a5-a1d5c29ac73a · inbound
Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems WavChat: A Survey of Spoken Dialogue Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5bff76c7-2b3a-478d-bc3e-aaf2a390e528 · inbound
Audio Editing in the Era of Foundation Models: A Survey WavChat: A Survey of Spoken Dialogue Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c068f586-cd16-4bd7-9cdb-a91e167203ce · inbound
RedVox: Safety and Fairness Gaps in Speech Models Across Languages WavChat: A Survey of Spoken Dialogue Models
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 16ed510f-853e-45a6-b971-8f702480c31a · inbound
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1aacc01f-e37b-4fce-9c64-46d6f0f3761f · inbound
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65231ce-691f-49ab-bddc-1b16309664a0 · inbound
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs WavChat: A Survey of Spoken Dialogue Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26ce0f7a-9521-40c9-a8ca-99e9b2e5484c · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment WavChat: A Survey of Spoken Dialogue Models
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 773619b2-f9bc-4a5f-aa97-df3bf5dc5ab0 · inbound
VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference WavChat: A Survey of Spoken Dialogue Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.