Pith. sign in

Paper Citation Record · LEDGER

Audiobox: Unified Audio Generation with Natural Language Prompts

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 69 inbound Pith citation observations for arXiv:2312.15821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.15821 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 69 of 69 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:43:41.177732Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c1ec0f3c-064d-4904-b008-082d18969987 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.065148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:ee75f40cfedcad9a228f865a3b4d98a3576bddbfad31fd7d9c3d957f4a45b602

Observation 82917358-2cb2-4ecd-a980-b3ad05c3c46f · inbound

Flow Matching Guide and Code cites this paper.

Flow Matching Guide and Code Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:28:14.138973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T10:28:14.014706Z digest=sha256:9b515c5c8287b3a6fbce40e0d41a0ca977f04efb1a3459eea2220c043d64d830

Observation b145e811-63a4-4038-8a01-4e156adc5cd8 · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.434184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.434184Z digest=sha256:e713d8c057c2102aa806bce4627159d0f0fffd99ba0fbb624ce79fd684a00c15

Observation ea093d5d-1bc7-4f91-b41f-9bbc2cd20518 · inbound

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor cites this paper.

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:52:03.920679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:52:03.920679Z digest=sha256:5d7d716003928dc7134081ffd69bc0bc5f86ea7d0c33d0f000100a816b03d35e

Observation c5402d15-6c29-4779-b6de-572cb59c2fc0 · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.619152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.619152Z digest=sha256:773a11b45fa9496213dec7beade0f9d6ae5786adbc293515a6e604fe0ad1db14

Observation 16951b98-ed79-4e0b-884f-4bed32300366 · inbound

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text cites this paper.

SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T23:05:27.439825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:05:27.439825Z digest=sha256:0b529d302a216be6a4c63634136f3cba444637e189e880c287eb63bec3361a59

Observation 61f37537-7038-4ae3-bf5a-78615cee647e · inbound

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis cites this paper.

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:50:14.370607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:50:14.370607Z digest=sha256:09522e317f237b291150848b7ae3c8576d1d5e92bd05d0ffe6340553928d66bb

Observation 72c87f90-3104-4ca3-aaaf-cd187fc35c2a · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.900293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.900293Z digest=sha256:5c4897579ca1a1fbc51c093afaa87d4b7fe35d0337ff05bd7d1fe2e654581198

Observation 9ae97aa0-ae99-4f7c-87d9-bcd2b1685718 · inbound

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization cites this paper.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.706605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.706605Z digest=sha256:507e63598c7f3f9ccee0221f9e5c337c1ab4af96a2afadf937caf1f8cf06106d

Observation 87d9f7e2-56b7-46f9-80fe-ff1f47f62436 · inbound

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts cites this paper.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.100232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.100232Z digest=sha256:9f878d02dd7af4c14ec43bd48daaf46d7e955abcc8b930466d4836abc0bf98a8

Observation 425947f0-4c8c-414e-9f71-c685b3c9fc17 · inbound

FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching cites this paper.

FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:31:29.191010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:31:29.191010Z digest=sha256:71ce32c0f06c47b1a1476ca9a1f2a9fd32a94f27543bbb59575d7480acf63d79

Observation 3e0fe9cf-f618-4751-95ca-173c00fee50e · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.979989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.979989Z digest=sha256:5b1c05d4fc934b45eae465ac1b4e4bee959c3857b3324776f9b216cb8b3fde1e

Observation 1f6b83bd-d021-4081-b844-7266387cb578 · inbound

CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions cites this paper.

CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T10:58:29.045502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:58:29.045502Z digest=sha256:3ca236b2438a56390b80203b0b164a10b15ed5b1dff6120052bff1b924d48a92

Observation ba086c40-95b3-425b-8253-8c5412fe4d0d · inbound

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video cites this paper.

VisualSpeech: Enhancing Prosody Modeling in TTS Using Video Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T20:50:17.331087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:50:17.331087Z digest=sha256:150061ea8d8a39d65d4bb2dbcc2a22a01bd8362131260f5474ea9ed26f14b48e

Observation 2415d845-f9a1-4ae3-bb0b-742007a90c25 · inbound

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation cites this paper.

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-09T18:51:12.656812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:51:12.656812Z digest=sha256:7eacb090520439c6ed2064b8684d184bc8485f573ef5bfe6f92a79e8cf320a8a

Observation 163393c2-c867-4746-b125-fe627ac4467d · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.761884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.761884Z digest=sha256:0b6e447119cabcca02e91dfe492d7ecd97908e65fe9a1c0f819e53005b5a2164

Observation 1de5ba83-5a11-4c40-9eda-c9d19a19f0fa · inbound

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound cites this paper.

Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:30:51.367209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T00:30:51.265062Z digest=sha256:8b5e9868d2fcd33c3c4ae59c9ef354dda4a7f42ec8c0f0d0a48833b673888c42

Observation f9b274d1-daa9-4b5a-85f6-5411061b25d3 · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.205829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.205829Z digest=sha256:b736d522db0cb785dea3fbfebccde2efedc9f5fbaac73b00f5e55068c42e22a9

Observation 2ab885eb-2fbd-44b9-a22a-87a76227beed · inbound

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction cites this paper.

Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:06:32.937055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:06:32.937055Z digest=sha256:1eb153cacc7aaa1906b2a28d3ddb694fa36fa6df14e4f0fb12f011919a399e83

Observation 053015d6-0d7b-49f1-b8b8-ec50576d9198 · inbound

LoRP-TTS: Low-Rank Personalized Text-To-Speech cites this paper.

LoRP-TTS: Low-Rank Personalized Text-To-Speech Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:21:44.261324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T12:21:44.261324Z digest=sha256:f0aefef482e0234b78a9cedcaf723f2b98c5f4ad79f5fa08b915932dec73a9f5

Observation 5236f2fc-4bde-4a3f-a90b-7afba7cc9d0f · inbound

RenderBox: Expressive Performance Rendering with Text Control cites this paper.

RenderBox: Expressive Performance Rendering with Text Control Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T11:51:21.819764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:51:21.819764Z digest=sha256:b6ed2a2d5b96911805dd34faebfe331782910f8e479a1ebf286f3a943c542806

Observation efdf9d49-2f0a-44b7-b1d1-1595bad1553c · inbound

OmniAudio: Generating Spatial Audio from 360-Degree Video cites this paper.

OmniAudio: Generating Spatial Audio from 360-Degree Video Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:43:41.177732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:43:41.177732Z digest=sha256:c263e33d57f372958d1fb804a9bc252c3e9861532bd6fe65164b63ec04ef921f

Observation e82070b1-dc75-47cd-8e42-a328bcf9458e · inbound

Learning to Highlight Audio by Watching Movies cites this paper.

Learning to Highlight Audio by Watching Movies Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:56.733736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:56.733736Z digest=sha256:b452ac9d62640e4ab6cd851d96771337a02ab3b5b49c624097426450de6ec622

Observation c81ab182-aedb-4645-a4b4-e43539110abd · inbound

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching cites this paper.

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:30:51.011593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:30:51.011593Z digest=sha256:0edddbf896343fec9d97860844acab69998f30d848c1de8e4ef53c6fd1ff06d0

Observation d6df0607-d94d-45b3-8f6f-d6798b6b6e8c · inbound

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers cites this paper.

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:11.530233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:11.530233Z digest=sha256:aa2621142bf366d59322c10d9832ecaf0e20a5d9413ed1954a92ba014c780367

Observation 9ae32be0-7d26-4b9d-b5f7-32d78822003b · inbound

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations cites this paper.

RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:42.653144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:42.653144Z digest=sha256:cd1d8240e8438f2b651e0f9f49bd17302c4ae50c851e047e6269110ae43f1f32

Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.842801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.842801Z digest=sha256:2c4e935c987e70ce7fd66f1ea16e7c376b0a26459045cc8492c069a0850d34ff

Observation 7ec86a7c-90b7-4ae1-af69-b106424ed52e · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:04.049255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:04.049255Z digest=sha256:67736553fc2f6ae9b915c08a84b54e00606387d9e4921fe8b2f6aab80cef7bc9

Observation 0cad407e-8458-4f77-ac9a-8ffe4e44edda · inbound

In-the-wild Audio Spatialization with Flexible Text-guided Localization cites this paper.

In-the-wild Audio Spatialization with Flexible Text-guided Localization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:49.503496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:49.503496Z digest=sha256:7927b7461382bf74ac0b76c094cbb13f61b59923d4c61271c03f5d89a3236671

Observation 8a35ce5e-51c7-41c7-b247-0f98169daa3c · inbound

InfiniteAudio: Infinite-Length Audio Generation with Consistency cites this paper.

InfiniteAudio: Infinite-Length Audio Generation with Consistency Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:29.298666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:29.298666Z digest=sha256:d0aa010001007c78aa39b9714d3c24f40109495d5e4503cc5009b4039bbcd087

Observation 25c4f62d-2778-41f1-879b-66735e0cfc0f · inbound

Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation cites this paper.

Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:24.856067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:24.856067Z digest=sha256:b551b27ceb9b7439a6a14c8c0bd32acc6f61864b7758ad90e23ee4331529071e

Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · inbound

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching cites this paper.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.234441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.234441Z digest=sha256:7808f3c32b5a102c5362c71978463dee84e45c12010270f627fd13f432ba9395

Observation d77c8120-21a5-4283-bef7-4f137d735d91 · inbound

Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation cites this paper.

Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:17:14.641009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:17:14.641009Z digest=sha256:e387b037ee229ef55dc53521536c186b1f23563894f842a07f85e845e1e97a87

Observation cca6c30e-5db6-455d-9495-cd915c13d794 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.451574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.451574Z digest=sha256:e82ad96cd6b506eabbbbee1d1cb4738272490c4116a0ef5418ffc716358af059

Observation f3fd9f59-84fe-4f00-8876-9f085173922a · inbound

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization cites this paper.

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:49.645421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:40:49.645421Z digest=sha256:a92a9b70a3190ae7ac043f33a00a6276d8f8b513cf563234cf3ad6acc8e13473

Observation 5e1be8c0-385a-42bf-8d7f-0378490c6428 · inbound

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models cites this paper.

DreamAudio: Customized Text-to-Audio Generation with Diffusion Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.672162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T18:26:51.583145Z digest=sha256:96638444cc58de4cedd440cd85ae4b37e3ca84fd5160a34921f715458a955889

Observation 0c609199-248d-450a-8ba7-1316a92b2f4f · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.303876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.303876Z digest=sha256:442b51e355a2fd2a7156bc893a770c55f1011c7676661bdab2468010e4383448

Observation 85c5340e-9405-4a0c-a097-db088a4fb847 · inbound

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement cites this paper.

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:29:17.809858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:29:17.809858Z digest=sha256:e08f6fdec91dd9131dbe8a5c73351eeee755902886f686982b753cdefb5bc003

Observation abe87faf-64e1-4cfe-9174-d11e04825029 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:08.922024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:08.922024Z digest=sha256:e58b5a8f16b5950bd68d136ec5ff570327b5b49bcae16b98cb3fb7cc8e7dc4db

Observation 54156058-05fa-4072-93ba-39b32cde10fd · inbound

FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation cites this paper.

FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:14.257145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:14.257145Z digest=sha256:f53fd6c62cb14ab944c4dd849f97819ef8a253ae0d10b561c3633890e41d785f

Observation 78d2d6a9-4041-44da-9dfb-4a344fb57ac6 · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:33.115934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:33.115934Z digest=sha256:0589b74b6d52b905b0047dd83e612807bd76238d5cb2209320714a3366b753dd

Observation 42ed3b4b-f6c3-4c91-ba67-bc68f8dad714 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:51.864433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:97d418242f1c6063c61d6e27603631f51aad7d5ee67b20ca84fd0ff52828b35d

Observation b384b49b-ef6e-4168-93ae-0a5b55b3b17d · inbound

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control cites this paper.

Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:28:04.692081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T22:25:53.037164Z digest=sha256:d2d9d212d2279896feead6d7135ea2c70b22494bdacd240bb93830013fb0f402

Observation ffc0fb2c-350d-44f9-86f3-c1c906de81fc · inbound

PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing cites this paper.

PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:56:04.361972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:21:43.379525Z digest=sha256:ced0f94d08cfac63fd73d863ec2363fe63f48fd2c5d26fa658a787c6dc63e790

Observation 0746874f-c294-4504-8dd5-241db0182cff · inbound

A unified perspective on fine-tuning and sampling with diffusion and flow models cites this paper.

A unified perspective on fine-tuning and sampling with diffusion and flow models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:21.370267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T19:43:42.331642Z digest=sha256:5c69c186e34f754be6e9d90029bbc888d2330f4bae6448a52bd1087e3d69933b

Observation 1d819b93-7627-48bd-b92b-619c12b3a0ee · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:41.933376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:1763d399ad90fcaa05cfb39993ecfa5156b28785bc6d0e9af86264c43d3ab711

Observation 66c3dcb2-33d5-4afa-a404-39d9a448298f · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:36:45.292578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:dc46d3a82d4d3665146ae80006db21391c89232100e3c169ecf3ffbb74cce58b

Observation fae6d979-9042-4207-ac15-07057f9144df · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:35:07.877386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:9b282a5d079e863ee529376aa6733c7bb0a1111f570e1d90181c5e1876d83a78

Observation adaf4418-7ee9-45c4-a77b-e9c57146a96f · inbound

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation cites this paper.

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:13:21.342726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T14:08:30.802619Z digest=sha256:fc3296390eb33fe6514e0a0c7d31a908928dbb6a3a10eb249fe4d78dbe144cb1

Observation 6c1d55db-e97a-4f06-b3ca-42f193b15c9e · inbound

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models cites this paper.

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:44:41.261751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T13:35:01.226818Z digest=sha256:719195d2f4ebb3485d5525d46380b2e562880d1df069a4e16906d843f1297b35

Observation 444d884a-4575-45da-9dac-ac58abdf689f · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.930200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:918dd4ff2cb09ace9c9902e301bba19fd08a75a19963a97b28b72316299b9e59

Observation 688eb36f-523f-4e77-a7fc-6e0b0fb11064 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.626271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:1792095841fe5e28d8b5becb7d5e1311bd0ee68a68658b9034e91f9cc9041e29

Observation 087fe4f2-13f7-432f-bd41-19685fa2e3c3 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.917213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:af892a140a8e437fea7a19eeb7e631a2baf0f6548f8bf1ff755da20b7d143b02

Observation 435d69f2-c82f-4b7e-b121-cff147ec5199 · inbound

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement cites this paper.

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T12:42:08.983227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T12:34:06.024192Z digest=sha256:44a2e96849251046a2462afcce051148ce48a4f6e2e0864db6168527ad19b7e1

Observation 590a8127-654d-4387-90f6-df1426c6ce91 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.711855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:e3a6b0f03da6965f7f996dd915c5a5ee5e9ad46b4849f725c2e86a0bb89eeb6a

Observation 2c590f6f-ae5c-47a6-bbc3-0c9cbd2afdb7 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.662075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:65fae6fcfcd506676123d4c6411aa3f13acc1280f89da347645e2c0de5eb59b0

Observation 2278720b-7402-4a96-ade9-baf10dabe26a · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.917841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:a418f6a4f1d532db6173eb0cfbcab9ad6a1069aec3c805acd3a1d9e0fa9205a5

Observation 980e0768-d242-4eb0-a098-4155afd207f7 · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.956747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:2730d8d0ef987f1721c5af02f20ad6657c12d3611785d6c5c9a1a1363e0cd51a

Observation a3d23cec-77d5-403f-9493-0bf03a9ec394 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:df3c670fe2b5e418a197c29495b30b8bb5e1b3e6978c9efa580a8cb11cb7bffb

Observation a4231bb2-8c63-4c46-82fd-17b2e69f233a · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.480374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:77a58190e35009cb206b4b12b4cb4730e7e30a4496197812f087ae56852445cb

Observation 7a7202ec-e2cb-4f8f-b476-2fe65d39ecad · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:b113736109a092f8afaf51cbf61ce4679999d62044753eae6632de0ea5cda705

Observation 4d288361-f52f-4898-a045-1092ca4af2de · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-30T14:08:14.459751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:08:14.459751Z digest=sha256:25c88dd9ef87bc610de3ef88cd5f158f3590a375a97b6f79dae9a3535e40851b

Observation e525aefd-e97b-4c32-93ce-dc1f3fa08ff8 · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T10:17:26.573299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:17:26.573299Z digest=sha256:2ab45bdf1aeae40f9aae935f31326934b1048ee2ec48650d6d51e54cffcd471a

Observation 81c4d6f4-71c1-480f-bd87-420838acd450 · inbound

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation cites this paper.

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:42:59.876083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:42:59.876083Z digest=sha256:50cbbb46e45219388db36ce1a4d42fb71a5cfee443505ccc93ba9012c1aa1e5c

Observation 820fd6b8-e181-498f-af51-69118670a257 · inbound

CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control cites this paper.

CustomDance: Customized 3D Dance Generation with Coarse-to-Fine Human-Centered Interactive Control Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-10T21:56:38.468327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:56:38.468327Z digest=sha256:ae99fe649281c49a40a1ff467291e1831b8a365b0dfa9753d98ff10ceacae8a0

Observation ee3026f6-d04c-4633-8cde-7ce50f2bfcb9 · inbound

VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics cites this paper.

VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T00:44:06.171201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:44:06.171201Z digest=sha256:173f35d1cef5361b6f740d06e8cbfc14f1206e93fc1299c52c8d521a93a9668a

Observation cdbad638-fd23-48e1-bf05-11a5343a5750 · inbound

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation cites this paper.

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:54:55.057843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:54:55.057843Z digest=sha256:9331916040b55bc776f7b37183cf507b587913c3333b94b25ab128c7e004ea40

Observation 57e5ad33-fe50-4da2-b06f-b3c4ed53a099 · inbound

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation cites this paper.

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:54:55.062732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:54:55.062732Z digest=sha256:65f3fef3541521c8122aaf2343e2fcce5b66b826c8d926bf7248c6e1aa2f6766

Observation d5d7220c-6f7b-4fc1-92d5-a9c9ae5d514e · inbound

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching cites this paper.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.785493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.785493Z digest=sha256:7b08bf553e174f8abb1a63e91c5bf845e299c647cf3816c056f908aeb726203a