Pith. sign in

Paper Citation Record · LEDGER

UniAudio: An Audio Foundation Model Toward Universal Audio Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2310.00704.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.00704 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 46 of 46 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.810390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

14
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e4360b2b-7a9d-4c88-ab5f-5dff3bc5fc37 · inbound

Moshi: a speech-text foundation model for real-time dialogue cites this paper.

Moshi: a speech-text foundation model for real-time dialogue UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:13:22.565669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T08:13:21.962488Z digest=sha256:527c29b901c217e5c475c22247f8dacaa1c023aa94c01d7d2663af028c1607c1

Observation a78a21a8-abbf-4ac3-9259-c77b9d0d1bf3 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.810390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.810390Z digest=sha256:83a19431f7feca69f9e9e373815d7e7b35261e0c678c81ab500a73a6bb2bb8b4

Observation e1ad2248-1f26-4077-b858-4f3daf0ffb90 · inbound

High-Fidelity Simultaneous Speech-To-Speech Translation cites this paper.

High-Fidelity Simultaneous Speech-To-Speech Translation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.194777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.194777Z digest=sha256:8548ae08f7dc29f4eb42b84e4a7018a4d95497e17912b9b9eafd6a8f4380ef9c

Observation 9cc133a0-5e46-45ce-98ad-a3c78c6985f1 · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:22.509705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:22.509705Z digest=sha256:815d801956b5c8cdd062aef09b3f92dce73ac2bd7d3eb0ec55454ec434000f0a

Observation fe997328-c058-4a0c-b19e-22884066b795 · inbound

Do we really have to filter out random noise in pre-training data for language models? cites this paper.

Do we really have to filter out random noise in pre-training data for language models? UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.446787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.446787Z digest=sha256:dda0ef35330e9651b3a18e677ac3aa1119823385e48eade02442883497d9a896

Observation 0c1be9db-808d-43b2-8d74-453dc411fda3 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.112262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:76d6fbd087dd8866907b43f11a05523732f77e191e065c95f94fbb4f0db50044

Observation 6c5afbac-e9d2-4862-8f5f-cab96a22e33d · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:16.151902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:16.151902Z digest=sha256:219912eee8c7c80470fadc44a9632ce10cd79481be8a93777443112b449a1268

Observation 48e7124c-05e4-43c9-bb0d-10ed4f8b5f88 · inbound

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion cites this paper.

AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:08.829562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:08.829562Z digest=sha256:bd1985a7de8638139e6e4e423da87afee6788bf06107afddf2ad56bdc6a5c20e

Observation c87ec37f-a114-40b6-acff-57d356b935c6 · inbound

In-the-wild Audio Spatialization with Flexible Text-guided Localization cites this paper.

In-the-wild Audio Spatialization with Flexible Text-guided Localization UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:49.805442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:49.805442Z digest=sha256:683a8b268a6d0aaf147c07a9dc38354b38b12d837bdc31262ae7df906d85d3c9

Observation 4eb02c9c-0049-4412-b20a-1d0f0b385f28 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:58.032031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:58.032031Z digest=sha256:7be7d40c9a6d17a2229bc9f87f42fc48c66e455bda8f151a990ef8951b87cc74

Observation 93b38794-9dc5-4c3f-b4e9-219efef4dac9 · inbound

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion cites this paper.

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:24.284181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:24.284181Z digest=sha256:b70af4421e848152df4f03ef5cc14c939a5b435435820641979f812cbc0f2bcb

Observation e9596670-4fda-4713-bc55-68aeb0e506e2 · inbound

Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration cites this paper.

Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:33:55.754374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:33:55.754374Z digest=sha256:770859a7b7f05801f9e758c0360ddff68cca8d12b64835b6ee37697e84367ccd

Observation 4572b879-d5a4-4980-9e0d-41350217fb08 · inbound

Scaling Laws of Motion Forecasting and Planning -- Technical Report cites this paper.

Scaling Laws of Motion Forecasting and Planning -- Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:23:32.781421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:23:32.781421Z digest=sha256:93157fd693186c11e20cbb6ce4aece81bb0531e8973fced86e7cc0ccdf096e94

Observation 981006d2-d1a6-4647-8332-b80cd2f36eda · inbound

Exploring State-Space-Model based Language Model in Music Generation cites this paper.

Exploring State-Space-Model based Language Model in Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.387244Z digest=sha256:b90e5f1cc94482da7905231e9d8bbc019f9ced0a3ebd93c0e6899048c9bc5c28

Observation 72949c76-f247-4187-949e-bf301ca95cc7 · inbound

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment cites this paper.

SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:07.244165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:14:07.244165Z digest=sha256:7d3d92e300a92c9f502d519db89bf7c62878a54cbd2b417de5c5b0b9c1931650

Observation d682d2a4-1638-47e1-b3b7-c04993354b10 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.524187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.524187Z digest=sha256:c39750bf925a144e70a438f580e54161f410fd452aa85666606e16658b5010bd

Observation c6e79082-fc53-44c9-8f1e-6ec750618d11 · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.500277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.500277Z digest=sha256:48207959a5c486383219d6a1b8d609e6f971cd6f93f69fd35d3d490e2aa9f528

Observation 118c0019-ded0-417d-97e5-48b2b48b530c · inbound

Small Data Explainer -- The impact of small data methods in everyday life cites this paper.

Small Data Explainer -- The impact of small data methods in everyday life UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:28.943831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:05:28.943831Z digest=sha256:c95fe8fb34a7a2d88d4c932ad731a70483c5a388811a7f3fefe7a4d9884ff26f

Observation 506e6954-1a24-43c5-aaa2-a2c3c992b00e · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.479314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.479314Z digest=sha256:53e4c1df67b871910d60d9df482893495c18031a7d6ad9e489e8b04f2d8853b8

Observation 8ff6b37b-4359-4374-8562-1d9f47b94cca · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.199707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.199707Z digest=sha256:85636ce8ce78bb94faef30f3152c7a58848b86917852755347047c05acd237ba

Observation ecd473e6-b5a8-42f0-9157-10e01aae665a · inbound

Your Spending Needs Attention: Modeling Financial Habits with Transformers cites this paper.

Your Spending Needs Attention: Modeling Financial Habits with Transformers UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:59:05.807002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:59:05.807002Z digest=sha256:73193e59632f634992b85511d9f0941bc5fc218f7b02cca3113c73c49219558e

Observation 404acb60-0c1c-40f7-93ad-3c6b2d3cc334 · inbound

Ego-centric Predictive Model Conditioned on Hand Trajectories cites this paper.

Ego-centric Predictive Model Conditioned on Hand Trajectories UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:29:25.573605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:29:25.573605Z digest=sha256:a3dad4ae0c3edad7238bbe1a1e651db20a110dcba6cfc02c8032b89d1c9ab289

Observation e6197b20-9dbe-4590-945c-9a1418149def · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.543438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:6912e9620128f76270356a71b5cbbb7d16e1ca0beca0f0517b39c83102839942

Observation a6ccd059-323c-4831-b26c-401f61f95b9c · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.870189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.870189Z digest=sha256:fc73dcb10af066f8de1a38287b6095495dffbc75a9bda1f46647bdcb6ca7ab0b

Observation 7c4c84d1-c061-4e10-a641-3c466918b1cd · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.644203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:1e0b2891fb230309d7e8f6c894d6709be0dd07f79d7ad0edd7bcf7010edcbee2

Observation cefc0374-dafb-439b-9a9f-a548839f2bdf · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.106539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:fa10ff6628ff669c661ef9db2e9882827380aadc9e8bafc316a928ae47f2e351

Observation b07c37f9-d40f-46f2-91cb-41d3251b0bbc · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:42.653482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:42.653482Z digest=sha256:4603f58d02c03ad4f8bb2019e8a511400b2ba074e4d0f04336c92858aa41056c

Observation 3fb4263e-3399-4e4f-b0ee-b933c3288000 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.473180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.473180Z digest=sha256:df176f5c890943daf7072833ad161f4b142ae7233c07ca883633c8ebadbd4e3d

Observation c6f8bd97-d23e-4c84-ae04-9a6fff638289 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.178157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:9330b326a834fc4fde53b37a0e6e2a2977a9723d394aa3a4ea4b69d0fd83da58

Observation bbcbef27-93d7-41b0-986d-1c53182bed16 · inbound

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards cites this paper.

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:03:33.272230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:03:33.272230Z digest=sha256:07a3b15c46c2166384a9268fb1811ca2970412db4abe71ff68d78306f8b416e1

Observation 71031306-d505-4e03-98b8-542f9d4370e0 · inbound

Simultaneous Speech-to-Speech Translation Without Aligned Data cites this paper.

Simultaneous Speech-to-Speech Translation Without Aligned Data UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:57:01.237575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:57:01.237575Z digest=sha256:6a74e504571d471efe3466064e664d1fd862f7090b685a77335e1e96e32c282b

Observation 0f88f706-4546-4784-8303-9f521e2f6234 · inbound

Fine-grained Soundscape Control for Augmented Hearing cites this paper.

Fine-grained Soundscape Control for Augmented Hearing UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T20:00:06.136667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:00:06.136667Z digest=sha256:24cc609b4852ee00ade1222c86d769ce0aa867601cfbeb7d4c84ed0ec956bb4b

Observation 1f131168-718f-4581-ae3f-3eab37b2d5eb · inbound

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cites this paper.

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.928071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:58:26.634355Z digest=sha256:57f9d4c2b6317ce7c961c507f73cdb9a47c11d9dff276abd5813be09e476c4ca

Observation c0fb6750-12f8-47aa-b72c-659f602c332e · inbound

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation cites this paper.

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:47:43.034952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T18:47:32.003545Z digest=sha256:78a2312e69143c5157441ac84ae12de69b63c2dc53ee48ffb18d6f6ab59d972c

Observation e2101993-a06f-4051-818d-136ce8d283bc · inbound

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts cites this paper.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.934369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:8e7bff4bd6f026cdd3a536479b953196bd6ba3abc753ba19423c43d3cbb73180

Observation 323b2215-b214-4cd6-a90a-ff2a2dc5a7c7 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.586510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:53fc619d30d8f62b8c944d68474da3adeea2a7306b938628b74b38344210909b

Observation 3f22033f-a2cb-4491-8224-66d7c5fdcebc · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.522604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:1a62848f563a0535e1e3015157d4e547ef1739e87e4408b6a4fbe1a73b047621

Observation 7746aa85-7f8c-453c-b45b-f0bc2f686713 · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.873935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:31f8028bb73235c5a13755de3798e3fc3eed4547f9b9f7266654bfc61419757d

Observation e05d13f6-79a7-4bf1-a267-e5cfad63baa1 · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:59:50.989373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:6104fabdac20ac6a304250dec42e837f141ee98247009c9bf78fae8681384ed1

Observation 22c60912-0aec-4dea-9cd9-2e2be7da006e · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 49

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.397390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:e7aa86798aa56250f563fe5f0416d27308286201d9a63908f15f508484b222f3

Observation 59d0817a-b1a7-4eff-8df7-4707f059cc02 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:9ef857bd42b6d37a56e4ef189cef9e84906b5a0117c60e237d7f63c3337834ec

Observation 2aa98018-0654-4444-959f-a554c30ce503 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:b2711f798720597da960e1fe7579b46f1ccd4df8bc84bf6339739b99f9c468dc

Observation ba1e27fc-e26b-4c2d-8126-fa61f865c25f · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-30T14:08:14.574038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:08:14.574038Z digest=sha256:3ba2140b58c8ca98d7e68dc7b3ff10ce710ba12fd02834493bf61e426dfda07d

Observation e8563ad7-2fc0-47f6-abc6-efa67a33c4ee · inbound

Qwen-Audio-3.0-Gen-Preview Technical Report cites this paper.

Qwen-Audio-3.0-Gen-Preview Technical Report UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T10:17:27.671246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:17:27.671246Z digest=sha256:983a0f9af993b4d30f17d23af00656d0cfacd6ea6f12143e8477d76387a2d526

Observation f6a686b1-0ac0-4803-9929-bcd8db3a2396 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:30.088106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:30.088106Z digest=sha256:7978691ae48bfc60d820992675ae76a561f56e0dcde2a29629ae4158486f5f8d

Observation bcd2f08e-ee7f-4ed1-b3c2-e482f20237cf · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:49.016529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:49.016529Z digest=sha256:36526dd4febf8ab6c2f3e8555ec8fefba00094a2535b0033e73d2634246c8445