Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 69 inbound Pith citation observations for arXiv:2312.15821.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:43:41.177732Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
5
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation c1ec0f3c-064d-4904-b008-082d18969987 · inbound
Movie Gen: A Cast of Media Foundation Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 82917358-2cb2-4ecd-a980-b3ad05c3c46f · inbound
Flow Matching Guide and Code Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b145e811-63a4-4038-8a01-4e156adc5cd8 · inbound
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea093d5d-1bc7-4f91-b41f-9bbc2cd20518 · inbound
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5402d15-6c29-4779-b6de-572cb59c2fc0 · inbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16951b98-ed79-4e0b-884f-4bed32300366 · inbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f37537-7038-4ae3-bf5a-78615cee647e · inbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72c87f90-3104-4ca3-aaaf-cd187fc35c2a · inbound
ETTA: Elucidating the Design Space of Text-to-Audio Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae97aa0-ae99-4f7c-87d9-bcd2b1685718 · inbound
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d9f7e2-56b7-46f9-80fe-ff1f47f62436 · inbound
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 425947f0-4c8c-414e-9f71-c685b3c9fc17 · inbound
FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e0fe9cf-f618-4751-95ca-173c00fee50e · inbound
Overview of the Amphion Toolkit (v0.2) Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f6b83bd-d021-4081-b844-7266387cb578 · inbound
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · inbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2415d845-f9a1-4ae3-bb0b-742007a90c25 · inbound
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163393c2-c867-4746-b125-fe627ac4467d · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1de5ba83-5a11-4c40-9eda-c9d19a19f0fa · inbound
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · inbound
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab885eb-2fbd-44b9-a22a-87a76227beed · inbound
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 053015d6-0d7b-49f1-b8b8-ec50576d9198 · inbound
LoRP-TTS: Low-Rank Personalized Text-To-Speech Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5236f2fc-4bde-4a3f-a90b-7afba7cc9d0f · inbound
RenderBox: Expressive Performance Rendering with Text Control Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efdf9d49-2f0a-44b7-b1d1-1595bad1553c · inbound
OmniAudio: Generating Spatial Audio from 360-Degree Video Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82070b1-dc75-47cd-8e42-a328bcf9458e · inbound
Learning to Highlight Audio by Watching Movies Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c81ab182-aedb-4645-a4b4-e43539110abd · inbound
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6df0607-d94d-45b3-8f6f-d6798b6b6e8c · inbound
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae32be0-7d26-4b9d-b5f7-32d78822003b · inbound
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec86a7c-90b7-4ae1-af69-b106424ed52e · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cad407e-8458-4f77-ac9a-8ffe4e44edda · inbound
In-the-wild Audio Spatialization with Flexible Text-guided Localization Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a35ce5e-51c7-41c7-b247-0f98169daa3c · inbound
InfiniteAudio: Infinite-Length Audio Generation with Consistency Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c4f62d-2778-41f1-879b-66735e0cfc0f · inbound
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · inbound
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77c8120-21a5-4283-bef7-4f137d735d91 · inbound
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca6c30e-5db6-455d-9495-cd915c13d794 · inbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3fd9f59-84fe-4f00-8876-9f085173922a · inbound
DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1be8c0-385a-42bf-8d7f-0378490c6428 · inbound
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0c609199-248d-450a-8ba7-1316a92b2f4f · inbound
Testing chatbots on the creation of encoders for audio conditioned image generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c5340e-9405-4a0c-a097-db088a4fb847 · inbound
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe87faf-64e1-4cfe-9174-d11e04825029 · inbound
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54156058-05fa-4072-93ba-39b32cde10fd · inbound
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d2d6a9-4041-44da-9dfb-4a344fb57ac6 · inbound
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ed3b4b-f6c3-4c91-ba67-bc68f8dad714 · inbound
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b384b49b-ef6e-4168-93ae-0a5b55b3b17d · inbound
Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ffc0fb2c-350d-44f9-86f3-c1c906de81fc · inbound
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0746874f-c294-4504-8dd5-241db0182cff · inbound
A unified perspective on fine-tuning and sampling with diffusion and flow models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1d819b93-7627-48bd-b92b-619c12b3a0ee · inbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 66c3dcb2-33d5-4afa-a404-39d9a448298f · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fae6d979-9042-4207-ac15-07057f9144df · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation adaf4418-7ee9-45c4-a77b-e9c57146a96f · inbound
Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c1d55db-e97a-4f06-b3ca-42f193b15c9e · inbound
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 444d884a-4575-45da-9dac-ac58abdf689f · inbound
Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 688eb36f-523f-4e77-a7fc-6e0b0fb11064 · inbound
ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 087fe4f2-13f7-432f-bd41-19685fa2e3c3 · inbound
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 435d69f2-c82f-4b7e-b121-cff147ec5199 · inbound
EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 590a8127-654d-4387-90f6-df1426c6ce91 · inbound
VoxCPM2 Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2c590f6f-ae5c-47a6-bbc3-0c9cbd2afdb7 · inbound
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2278720b-7402-4a96-ade9-baf10dabe26a · inbound
AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 980e0768-d242-4eb0-a098-4155afd207f7 · inbound
Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a3d23cec-77d5-403f-9493-0bf03a9ec394 · inbound
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4231bb2-8c63-4c46-82fd-17b2e69f233a · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7a7202ec-e2cb-4f8f-b476-2fe65d39ecad · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d288361-f52f-4898-a045-1092ca4af2de · inbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e525aefd-e97b-4c32-93ce-dc1f3fa08ff8 · inbound
Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c4d6f4-71c1-480f-bd87-420838acd450 · inbound
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820fd6b8-e181-498f-af51-69118670a257 · inbound
CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3026f6-d04c-4633-8cde-7ce50f2bfcb9 · inbound
VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdbad638-fd23-48e1-bf05-11a5343a5750 · inbound
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57e5ad33-fe50-4da2-b06f-b3c4ed53a099 · inbound
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d7220c-6f7b-4fc1-92d5-a9c9ae5d514e · inbound
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.