Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:30.804516Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2505.24496.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:27:30.804516Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T21:31:09.399885Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T20:16:11.072472Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c198cb4f-2277-46e4-b818-dc7780d4147a · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edaab79b-3a62-4c0c-91b8-ba89f5b2d5e5 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bfee4669-8a90-4c16-9f93-88822094d260 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation SyllableLM: Learning Coarse Semantic Units for Speech Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cca677-1b93-4b49-a515-98d54bc050b1 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bebc49f9-7056-4d7b-a1a7-053c58262520 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation SoundStorm: Efficient Parallel Audio Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cf99313-2e9f-452c-ad98-1832a9f59aad · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2aea3f9-24c4-499c-ab07-bcca87572d2a · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf7acdd-960c-4c94-9766-173297f31e72 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a8f3746-4c35-4145-85b4-f581da39e7d2 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff74aab7-912a-4854-8803-81fc31d1ff0c · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and Diarization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8039c6-d6e6-4c4d-a396-592356a8a347 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Sylber: Syllabic Embedding Representation of Speech from Raw Audio
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678a8410-e5f5-47eb-b1e5-fb9d8b1148b3 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754937e3-48ee-4aac-bf30-f7d5d76257a1 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation High Fidelity Neural Audio Compression
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ee3594-44f4-489c-90d2-8113b1b7325f · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Moshi: a speech-text foundation model for real-time dialogue
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9b3575-3c63-406e-9df8-319dba0bc389 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a9b888a-0be7-4c57-a0d2-f5ba9de84b66 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Variable-rate discrete representation learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b36240-57df-4b54-8520-42b0fefa9669 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80d3b2b-0f52-4f09-91ce-94141c23eb23 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 474b2f4d-d387-4824-ac27-39948ab44abc · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 804f4053-fb7c-47a2-8cda-497b332e9879 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24e8798-a39f-40c1-b133-d801cb3473be · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac5391e-e554-4825-98e6-738d8d2b430b · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8a22c26-ed85-4d7a-8176-72047c6f98a2 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64b67d22-e821-4b10-83e8-f4dad82a4fd7 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4cc8cad-7112-4835-b95d-48f150541425 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13c7ab8f-b3e7-4ad4-b1fd-78eb7cf3f024 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc8600bc-c246-4104-8726-e3f149b77625 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af9e4769-dff0-4dbc-8696-882fbfa1d5fe · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cba51742-9535-4d59-9282-9d7c22eee1e2 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fd6b9d05-f1fe-4556-badc-32f6e149c0b5 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 409174c7-b487-4315-9c12-d8f74aa4f8fc · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e468d2-c148-40b4-af74-958742a806e8 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f4a5b07-c22b-41c6-bf22-5d42a07dcd83 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ec0cdd-833f-4b56-a1f9-cc5307e78352 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9c938da-1af9-4a58-b535-e9c9e7662bac · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e970c14b-eae9-4458-9ce8-70ac03908f86 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MLS: A Large-Scale Multilingual Dataset for Speech Research
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf9bf34-8d6b-4cff-8e05-7c6e4a1bed9f · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6d42c4b-40f1-49e6-98a3-8d5b3f9d1870 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d61baa9-d03a-428a-a061-6c8cb9922bea · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43042dcb-3e18-4ce3-9eeb-ba2d7105f563 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb30411-f999-45e3-831b-7e0cb6e92f21 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 38305736-0f56-4965-b26b-c9792c6a6ce0 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88b84e91-c007-4a5d-9186-1ade72b1f353 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b079e618-c75c-4a31-acc4-37d02cfc53ea · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e81728e-ab1f-4880-a6ab-9b88524bc37a · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation LLaMA: Open and Efficient Foundation Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4072c288-ec98-4003-8ff6-bdcc3fc82fba · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee274426-26c7-41a7-83ea-6e304359fe61 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1681cb1f-31c1-41c1-a254-49e330700f19 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05bcfd94-028f-404c-86a9-ce88db9cb736 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e80bb3a2-ddd5-4cec-b903-be5f1ad0f61d · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 68424da3-44ef-427b-8bc8-191d0020f98c · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4cc1ab-fa79-4a54-947f-2a87b497e724 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2fdcfe29-5aad-4dd9-8383-9aa86becdca2 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740be111-7553-461f-b72b-20d9ddaa4162 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc5ac70-5974-4e85-8076-9615b6a0b18e · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f0d1cd-218f-4da7-ac90-59d6970cf277 · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6e1986-3774-4bbe-82ef-ebc2b9637baf · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Frontiers in Communication 3 (2018), 25
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e6daf83-adb9-4827-9bec-cf163542bb2c · outbound
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation Nature Communications 16, 1 (2025), 803
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bb66ea14-8f2a-41c6-be0f-073aabe5ae2a · inbound
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a6d9c15d-b53a-43b1-9aba-e28e92808ccf · inbound
Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.