Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:42.951304Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2506.00975.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:42.951304Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-12T08:53:13.408944Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T21:28:58.307813Z
85 of 85 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3c6e1d7f-e66e-450c-b2e6-694dcebd8e65 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction URL https://chattts.com/
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29e318d-d351-4b0b-b673-5f2dafab736c · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction L., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c068f11-9a6e-4a59-b3d1-ebd2798c923b · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895298cc-c608-419f-998e-e946641413b7 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1636c54-3f43-496b-9d53-5b339f20af2f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Common Voice: A Massively-Multilingual Speech Corpus
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5438fe-1fe1-4d58-925c-5c8a7a0fdcdd · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be36cb6e-4e8f-4c7f-a43f-6d65697acb32 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Wav2vec 2.0: Learning the structure of speech from raw audio
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ea4246-ec6f-4fc2-9b03-b0f163153d80 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiolm: A language modeling approach to audio generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bb04f5b-c50f-4ee3-982c-7d70c3d03ece · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction D., J \' u nior, A
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9e0bea6d-ba96-4772-ad3b-b985ffe392f6 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3105a6d-23ac-44bc-adf1-d7dbaa5fd501 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78a49792-6902-455d-b0a8-d4bde5038fe8 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Qwen2-Audio Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a904c0-bc0b-4e51-b1b9-2c30286229e6 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Fisher english training speech part 1 transcripts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22a12f79-7fe4-49e1-8e7f-61dfd7ec94ef · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Simple and controllable music generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 312dfe16-cff2-4271-8325-6e2db28426ef · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Y., Ermon, S., Rudra, A., and R \' e , C
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d474f38-9b34-4734-a226-eb1c4ca373f0 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Moshi: a speech-text foundation model for real-time dialogue
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d72e998-71f2-475c-9171-1452dc0e2651 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Pengi: An audio language model for audio tasks
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8035faf3-7c08-4760-b83f-b6f966129199 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2143a5-894e-4cd2-8ea6-84f5f8324f3d · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., and Wang, H
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92319e28-3802-475c-9737-cb73be4d4fd7 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Taming transformers for high-resolution image synthesis
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe33887-8a6b-4137-8a82-948411c9230e · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiochatllama: Towards general-purpose speech abilities for llms
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3827e764-6204-4997-b37d-b787117fdf61 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Vita: Towards open-source interactive omni multimodal llm, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3af4459a-7331-4898-9375-ffa2d8b34895 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Funasr: A fundamental end-to-end speech recognition toolkit
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac321634-5ca9-484a-8c90-13f474f18b1f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Gat, I., Conneau, A., Kreuk, F., Copet, J., D \' e fossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 551d551f-dee4-47fd-8ee6-60bd884b59d8 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Kenlm: Faster and smaller language model queries
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8dc67ce7-02a3-4c47-882b-a35b76ec12f8 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70df1984-f0c7-4d8c-8fcf-ba11ed3134da · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mistral 7B
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3edb39cc-aa46-4f93-b2bc-0f6f5ec4b43f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7481c0-d0d8-401b-a7cb-ccaa79b20de2 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e6a44f7e-c36a-485a-8fcc-6b0313bede3d · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4acb6830-e1a4-4ae2-9dd2-108b2a38de82 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Diffwave: A versatile diffusion model for audio synthesis
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8aef1af1-6c15-46ed-aec2-25318c405e47 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Audiogen: Textually guided audio generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99c49b36-5fe5-49e4-9401-1637cbae224d · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Voicebox: Text-guided multilingual universal speech generation at scale
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c18203e5-cdca-4886-98ff-5af9c24c8ace · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Autoregressive image generation using residual quantization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9cada8b0-4ef0-4bcf-917b-c8d5e96ebc34 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51370c96-703f-4c0e-a66f-2859258b346f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7cd156e9-0f03-46af-80cd-2e2aa83ec3e0 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Autoregressive Image Generation without Vector Quantization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b84a7b0-9973-4d37-b53d-67fb7fac0404 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Evolutionary-scale prediction of atomic level protein structure with a language model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577737c9-fc2d-49ef-a029-31c956a9fa82 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction P., Wang, W., and Plumbley, M
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fe9c039f-14ea-439d-a24e-e7beb456a9aa · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d6e50ae6-0bd2-4ffe-89d5-8c9e59471355 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Language Model Can Listen While Speaking
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb002973-bc68-4d7f-9ed1-0634d35b79a6 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction R., Subramanian, S., Mohr, B
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75dea8af-703c-4fbc-ad56-db0becf0e6c3 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f5d090-d7d9-4f53-aac1-56f46012cc2c · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction J., and Ramanovich, M
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b78aef10-d1ba-489a-aa7c-b115592f2056 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., and Dupoux, E
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc02d0d4-486c-4a1a-8ae7-bc2233c0cdcd · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Spirit LM: Interleaved Spoken and Written Language Model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d38a4a-ce27-44c7-8e0b-09a04e20b094 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction GPT-4 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df8a2e81-3e6c-49fc-9f93-86634a4510a2 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27299d89-e9bc-4255-8f4d-c4c55c3362ed · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction S., Constant, N., Raffel, C., and Callison - Burch, C
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1cd5d8c-5ea6-4c03-856d-6bec1acbf740 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Efficiently scaling transformer inference
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 42664b0f-f8b7-4927-9566-7a989af13eb5 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Efficiently scaling transformer inference
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1d2de32-f4f6-489b-80c9-03afcfa5fba9 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5bca6c3d-f9c0-4063-a333-af37c71766f6 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0a44f52-f6db-4094-9d7c-9c56bafaca96 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction The candor corpus: Insights from a large multimodal dataset of naturalistic conversation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f46aba9-72bc-4c48-9fe1-a15b3e8476f3 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 251546c9-78fa-4bdc-8f80-5870134b192c · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction A., Bekas, C., and Lee, A
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4be0a642-827c-40a8-89a8-0ade6f0c6960 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4458dc50-c08b-4f9c-9dd6-67226ad76cde · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Graphaf: a flow-based autoregressive model for molecular graph generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 76212d70-838c-4c6e-88c9-e8652cd920d8 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a36bdfc-2d8a-4785-9013-a3753838569f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f64fc8b-29c9-40b4-b397-55cc1f6fb24d · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction PandaGPT: One Model To Instruction-Follow Them All
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0eecd77-e1a7-4f59-aac7-d6522552d8cb · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 53eefead-cdb1-4489-aac4-d01aed5c75f5 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction SALMONN: towards generic hearing abilities for large language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef1f5c20-0cd7-4bb2-a8a8-2266bcff8bb3 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction C., Ture, F., and Lin, J
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ead31476-df3c-467a-b001-a59e2152fb08 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3628434-6402-4c47-b395-c74b301bb76f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Gemma: Open Models Based on Gemini Research and Technology
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b29c5b4-6da7-4c8d-952b-b502da87403a · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6014061-9c6e-4f7c-8312-53d06888a81b · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Neural discrete representation learning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66c2f283-ffc8-4f9b-9b40-b25d6adb6bc8 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4870119f-084f-4438-ab2b-8d42b446748a · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b2b3e3-7e84-42ff-a520-9b7525c41262 · outbound
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation feb87d25-b143-43c8-b5d0-e1124abada8c · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction SpeechGen: Unlocking the Generative Power of Speech Language Models with Prompts
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7218aea1-0087-49b5-831f-f71b3e529ccc · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Next-gpt: Any-to-any multimodal LLM
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 682e39be-d364-4755-8a56-0b5f53469a1f · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction J., Wang, W., Lin, K
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b034d105-1558-44c9-8ac2-b0fa4e255f43 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c7c0083-2de3-4e05-9bd1-c91a2c79b57a · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Diffsound: Discrete diffusion model for text-to-sound generation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e3e54503-2c5e-4ae0-946b-8f85d0a409e8 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Instructtts: Modelling expressive TTS in discrete latent space with natural language style prompt
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3596fa4e-db4f-4031-b596-9e8ae20c528a · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction L., and Leskovec, J
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4e621224-e9df-44e0-abb1-1c2d5ea4649a · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Soundstream: An end-to-end neural audio codec
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0ce57509-82bb-4136-938e-900155ee3eae · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1ef71db-56d8-4cd7-bc49-777c03c8b47a · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a439b41-1ec4-43d5-8231-6bb96448f359 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model, 2024
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 997ba1d2-2a75-48b7-84e9-318ae13e46b4 · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Mmspeech: Multi-modal multi-task encoder-decoder pre-training for speech recognition
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 69a63b19-1c08-4ca1-928d-38e2304b0bab · outbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction write newline
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6d85ff-3033-4893-b0eb-51c374cca2bf · inbound
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05a346ee-4bdb-427e-b5f7-2ae849d7bb49 · inbound
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.