Pith. sign in

Paper Citation Record · LEDGER

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

As of 11 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2508.00733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00733 v4

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.971825Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:25:58.872482Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.612458Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b310c81-4c27-40cb-8c4f-8af5af29d927 · outbound

This paper cites VoxSim: A perceptual voice similarity dataset.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation VoxSim: A perceptual voice similarity dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.281776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:28.939172Z digest=sha256:d49bdfb61c7b6d335e43566a3bdceba9ade9d13d08480055b21dab58dc6d1d79

Observation 5f9e9fc9-d300-46c9-b166-d826846c8561 · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.181087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.181087Z digest=sha256:fadb0a592dfc5fb9bf624a8a9049738c1dbdbd06b4a89161a425dae697cce7cd

Observation 7e15cb2d-cf4c-47ea-9263-17ef14bee244 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.426627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.266404Z digest=sha256:cc68cd183e9f44bf85185e31080bd5bd9626efd459c3da350c90eef2667e9570

Observation 866d9cf6-06df-4edb-b9e8-78ed43b07394 · outbound

This paper cites Clotho: An audio captioning dataset.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Clotho: An audio captioning dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.415330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.760649Z digest=sha256:6aebd90b77885892ce49434c94c00984e626caae47e95bbc579a6270ad405dd7

Observation 437d8bde-4f5d-4bd7-941d-3210bf9447a2 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.852239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.852239Z digest=sha256:7046cbddbac67f5de91066d7626f314ccf133ece9dd0dcf1fadf2e076f07ebb2

Observation 01a1f4a8-eec3-40a9-97e4-402710127371 · outbound

This paper cites Vid2speech: speech reconstruction from silent video.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Vid2speech: speech reconstruction from silent video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.403698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.872025Z digest=sha256:03814ecbf80a4d26944c4799e08c5dca32e40c17ca44845ba9f3266d75d422ea

Observation ac1f3415-c111-4879-a6d0-e4066fe59202 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.876089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.876089Z digest=sha256:7bfb3ef49669656d1cd0d6a888631745c47b80c0f361d7f5bbcb236e1ff17354

Observation c7bb091b-edf2-421f-808d-1cf2109f21c1 · outbound

This paper cites ACE-Step: A Step Towards Music Generation Foundation Model.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.880406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.880406Z digest=sha256:a7a90f31871cef797e5c9462a935e4a0702aa0533afa15a1295888cdb426b58a

Observation b5e1a81a-9666-4ad7-80a5-dabf59f440e6 · outbound

This paper cites Effect of clustering on the me- chanical properties of sic particulate-reinforced aluminum alloy 2024 metal matrix composites.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Effect of clustering on the me- chanical properties of sic particulate-reinforced aluminum alloy 2024 metal matrix composites

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.393012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.884269Z digest=sha256:285dfe85c1d9c909a645ce3acc0b7fcd1c2fdad5f7459cc768116bdcae9a3934

Observation edaf99ee-7cd4-4aae-bc7d-8bb2d0200b9d · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Audiocaps: Generating captions for audios in the wild

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.372452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.891922Z digest=sha256:95caae859c89ce06186a1e749017b308bcfb65722748784cc883406161ea4631

Observation a8231362-d76c-4f73-a1dc-76c8d930a7af · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Efficient Training of Audio Transformers with Patchout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.895970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.895970Z digest=sha256:a103f4f8dfbb3f2435e5da3397da1585db28a7e4b1e0985e9e4ab45eb053ae0d

Observation fff1aa02-8cbb-415d-b915-9fc0222fc5f8 · outbound

This paper cites Ntire 2023 challenge on efficient super-resolution: Methods and results.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Ntire 2023 challenge on efficient super-resolution: Methods and results

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.360907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.906533Z digest=sha256:92c88bca77a6105078de2baec7a18f840d9e657f6438c05a1eb3a7a7b7b69d1e

Observation 7f62678e-05c4-4db8-b0a7-0fd6f90625a6 · outbound

This paper cites Flow Matching for Generative Modeling.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Flow Matching for Generative Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.910372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.910372Z digest=sha256:f7997b6b51440835c393aeb7ed9519878659eb486d66c442970b72ebcd88b506

Observation 1d75771a-4415-4798-b175-2d463d71af6d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.913997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.913997Z digest=sha256:04a289bfc558c891d05ff1cb5bf279ad27c15326e72011e36ce9028fa52ea5c4

Observation a90c9f78-cc57-45cb-a67f-e79f1a80ff1e · outbound

This paper cites SVTS: Scalable Video-to-Speech Synthesis.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation SVTS: Scalable Video-to-Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.917486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.917486Z digest=sha256:727c164749da1e396424e7165ef3c2246c2f3032a7165c08d9129edd697f0300

Observation f586a7f7-b5e1-4e4f-99eb-96e5cb855338 · outbound

This paper cites Effect of large cold deformation on characteristics of age-strengthening of 2024 aluminum alloys.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Effect of large cold deformation on characteristics of age-strengthening of 2024 aluminum alloys

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.350721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.921132Z digest=sha256:aa8e153b4f2094976864c0ff9873caf2c3a5a226719fae457da7bfadc90ca7ea

Observation e55a50ed-15c8-4340-baba-00de4d55dffa · outbound

This paper cites Egosonics: Generating synchronized audio for silent egocentric videos.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Egosonics: Generating synchronized audio for silent egocentric videos

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.339603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.924581Z digest=sha256:02fe7ecd6f3395d55b96be0de28854c2149b5c9c130ef2460c38034bb5a736d4

Observation f01b5d7c-4379-4464-bb4f-9c504a24c72e · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.328331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.928332Z digest=sha256:18f0b6415772fedc4835c2e93004a299bc7d8321a874a5ac62780af11e592857

Observation 3d927eec-728b-4229-b022-3f745738ec2e · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.931990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.931990Z digest=sha256:6a4d862782fedb18eced6c289f5b1d529ea18e5ad1b0ab6c528da9b0df178f91

Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · outbound

This paper cites DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.083418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.936928Z digest=sha256:02ab78fc98113c9f58ea09e66341552c339b6c30f8ca85aaee270fc8f55f3943

Observation 07986a81-9e19-4ab0-b216-8451bc314589 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Generalized end-to-end loss for speaker verification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.305063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.945502Z digest=sha256:bbc083d8375c7f03829dd754471b0085f71b2b149c36a0dc21912a361de6ae4b

Observation 85d6e32a-0eb4-48bd-b14e-ec1ec0be758d · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.952467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.952467Z digest=sha256:4319cf1d381e0e2058eb39b58d5f9c8f776deb7dd3477fe76c328330d1192036

Observation 479b8915-b319-4b6c-9bc6-c948f19904f9 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Qwen2.5-Omni Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.956434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.956434Z digest=sha256:17ba890d739e4028504c61209154668a2d61749aa096d3458ccd8e2b85fc8c79

Observation 3d9ffcc9-19c2-4fb1-9bcb-88c4668eab94 · outbound

This paper cites Forecasting china’s regional energy demand by 2030: A bayesian approach.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Forecasting china’s regional energy demand by 2030: A bayesian approach

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.293381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.960482Z digest=sha256:5e63b5423d70f6cc5e4a28a6d535c955580c42e6e68331eb395f55d54c0233ee

Observation bc06c7c8-a161-4847-8105-a415e8d1c1ba · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.963674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.963674Z digest=sha256:f9dd6869f8669bb4aadc0b52878008e964f7451a2339185a2802cc2eb91cbbe9

Observation 7e2735b9-945d-4da9-bf83-a8d4ac21bf7a · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.971825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.971825Z digest=sha256:967f695f9520dcdf74314a4c903b8b987bceb13506ae4263b29bab8278773cd2

Observation 0d639cc2-a4a9-48bf-9a87-a1675c0146ce · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Synchformer: Efficient synchronization from sparse cues

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.382529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.887959Z digest=sha256:ab147843bad10c50de10c826cced907dd5023951ef3fb3419ecc5dea8d57a0dd

Observation 1dd81864-a54b-460f-8af5-3e6306a72611 · outbound

This paper cites Temporally aligned audio for video with autoregression.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Temporally aligned audio for video with autoregression

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.317525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.941955Z digest=sha256:0a0e393b79c8693189ea814eeb47f909d8c837d987df5801a057acfa65245ca9

Observation a8811ba4-f8e9-4233-bc25-30129c684a9b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.948925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.948925Z digest=sha256:d0d4ce2f2661d8609ce9f70cf337c1ff3e3be5038009fb3fd21c37d5e07f7d30

Observation cf95400b-5535-4a27-8787-6e8f5fe34cb5 · outbound

This paper cites DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.967160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.967160Z digest=sha256:300a3034c25480a3571c98a14bb99124147724120e76243494ec67fd9c36f3ee

Observation a9f25310-eeaf-42b5-a2f9-e18080661d87 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.398373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.398373Z digest=sha256:5d56efa4b2bb3d9ff1f1ef6cb98fac534ff9b764c054edcaf65c26bb9ec7a5c8

Observation 30c34803-a6cb-4e46-873f-06920348492f · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioGen: Textually Guided Audio Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.899392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.899392Z digest=sha256:192039ef78e7f794de5c22d3f0dc4e3c774e107ce1a1ffe54a92f49eebdff467

Observation 46465fa8-6932-426c-8b48-9bfda381a29b · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.902919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.902919Z digest=sha256:31777bb4e0303c334c0ed69981b406ede692f8b3d70a023b8982a329ee6ed5c7

Observation e37e05cc-f7ac-4ffc-84c7-b3678e1ae25c · outbound

This paper cites Jukebox: A Generative Model for Music.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Jukebox: A Generative Model for Music

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.619107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.619107Z digest=sha256:b58f27cf1ae8e55fd1b95f79223de29c7b78fc2ecc06933e80b4a5274f80d121

Observation c770a0db-f760-470e-a2bc-9da33e913861 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.065491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.065491Z digest=sha256:b49505dd65fc616f6beb958ee43e14142c1ff7b284e622daea02da9e0da3e81e

Observation cc18a5d3-7b85-4190-9dbd-4c6d356f4faa · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.235068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T06:04:29.497216Z digest=sha256:22cbfcc514ca5c0997002eefbba677c0b59b79beb0d68bfb7a78dcdd668a7e7e

Pith citing papers

Observation 5db16c02-bee4-4b43-8764-32f82764cc54 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.872482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.872482Z digest=sha256:0b22597c3773c7d1a381ed5f616e25294ecdfee31e8e31b02b10539da241c0fa

Observation 92cd37fc-20d2-409e-97b6-218e161ecbb2 · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.031207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:b4f732102f47423158979de27fbed1578cf510d4057f1c308f849b1317d2dbe2

Observation 3cdf2ce2-884a-4cdd-8752-e50636cef16e · inbound

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories cites this paper.

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.838588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:02:52.748220Z digest=sha256:5373e1ab49c825c6396c032059d121fe3938ff5632acb822f397b8a1cbba3f96

Observation d4725871-7853-4a8c-b72c-0e3b8675d5eb · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.060801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:3fbecaef2d0aa629020cc78fb7b64835dcfa303aac61755d859954024b7bf17a

Observation 3baa02ea-0a50-43c0-be32-347d30bc946a · inbound

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer cites this paper.

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.355982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T21:17:48.886421Z digest=sha256:032d852ceb21165986fbd43754721d2a26df95e4d0fa7fca5d3e07e39c044fc3

Observation 1ae5565b-83bd-4351-8c3a-fa0453af18b8 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.615037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:f90ee7ac89b0dc7923022e0171dd0adbe0ed6bb39c946000c02c87522456caef