Pith. sign in

Paper Citation Record · LEDGER

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2508.00733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00733 v4

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.971825Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:25:58.872482Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:35.612458Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b310c81-4c27-40cb-8c4f-8af5af29d927 · outbound

This paper cites VoxSim: A perceptual voice similarity dataset.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation VoxSim: A perceptual voice similarity dataset

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.281776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:28.939172Z digest=sha256:460d39295d8630f1efe969f0afc91bb5c1cdefb811351902c1da8bdc52c70e9d

Observation 5f9e9fc9-d300-46c9-b166-d826846c8561 · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.181087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.181087Z digest=sha256:dbd2f8c0eb20054b0e2d07d856704cc1d944baed49d8ff2c969942f9b65405f1

Observation 7e15cb2d-cf4c-47ea-9263-17ef14bee244 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.426627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.266404Z digest=sha256:025144a68ee2dab5121eac10a486ef3fded4a8ab65860af1dd0b8f99420fad21

Observation 866d9cf6-06df-4edb-b9e8-78ed43b07394 · outbound

This paper cites Clotho: An audio captioning dataset.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Clotho: An audio captioning dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.415330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.760649Z digest=sha256:9375229cd31ead4943669e4e8ee9f9788b031939b2ae39d9ada94c747816b2df

Observation 437d8bde-4f5d-4bd7-941d-3210bf9447a2 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.852239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.852239Z digest=sha256:a9f598d14d3e5d6f95330fe4ab1569425625a1ca9fff08d3176f043eafd8068d

Observation 01a1f4a8-eec3-40a9-97e4-402710127371 · outbound

This paper cites Vid2speech: speech reconstruction from silent video.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Vid2speech: speech reconstruction from silent video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.403698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.872025Z digest=sha256:9cb62f5b4a3ca1f84225b2d7c395aad9335af6b5032075a5f451f108a7f00d5b

Observation ac1f3415-c111-4879-a6d0-e4066fe59202 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.876089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.876089Z digest=sha256:4138a6e0caedef096076c5e8f1b8b2dd2dd1089403fa3f2829a1de1d592cd925

Observation c7bb091b-edf2-421f-808d-1cf2109f21c1 · outbound

This paper cites ACE-Step: A Step Towards Music Generation Foundation Model.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.880406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.880406Z digest=sha256:cbbb0188e36b99e5b593fddfa49a38e4816604ab28a13f80cb48a56455a65c42

Observation b5e1a81a-9666-4ad7-80a5-dabf59f440e6 · outbound

This paper cites Effect of clustering on the me- chanical properties of sic particulate-reinforced aluminum alloy 2024 metal matrix composites.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Effect of clustering on the me- chanical properties of sic particulate-reinforced aluminum alloy 2024 metal matrix composites

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.393012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.884269Z digest=sha256:620326bb44615d8ab873b4138045c72db985a362f91a921d609f961e02d62532

Observation edaf99ee-7cd4-4aae-bc7d-8bb2d0200b9d · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Audiocaps: Generating captions for audios in the wild

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.372452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.891922Z digest=sha256:67bc14ad06918df5bb89a64f7d98ebdb710f3675fef7dfb3c46029bec735a617

Observation a8231362-d76c-4f73-a1dc-76c8d930a7af · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Efficient Training of Audio Transformers with Patchout

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.895970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.895970Z digest=sha256:8e5b6bd7a337085552cabbfe2aa33a49b4eb3531c7d545f621db52f971f4445d

Observation fff1aa02-8cbb-415d-b915-9fc0222fc5f8 · outbound

This paper cites Ntire 2023 challenge on efficient super-resolution: Methods and results.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Ntire 2023 challenge on efficient super-resolution: Methods and results

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.360907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.906533Z digest=sha256:6fc01eb48275aef317d8fae57a11e63ba12c04024a1e51355ad432a8f08c2ae3

Observation 7f62678e-05c4-4db8-b0a7-0fd6f90625a6 · outbound

This paper cites Flow Matching for Generative Modeling.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Flow Matching for Generative Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.910372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.910372Z digest=sha256:e718c999540e52dd8f58ae662069938d633ad16e66270d7a03ccd87969bd5d50

Observation 1d75771a-4415-4798-b175-2d463d71af6d · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.913997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.913997Z digest=sha256:e2e3e66be2acabd013dbda3af807da486f96f20b83c05418c68aaad3a21621eb

Observation a90c9f78-cc57-45cb-a67f-e79f1a80ff1e · outbound

This paper cites SVTS: Scalable Video-to-Speech Synthesis.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation SVTS: Scalable Video-to-Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.917486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.917486Z digest=sha256:7509cbcb020790699f0e216a7d053a383bd7a0deb3431a015f1d8ca8cafef253

Observation f586a7f7-b5e1-4e4f-99eb-96e5cb855338 · outbound

This paper cites Effect of large cold deformation on characteristics of age-strengthening of 2024 aluminum alloys.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Effect of large cold deformation on characteristics of age-strengthening of 2024 aluminum alloys

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.350721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.921132Z digest=sha256:3fde4354dc724493f2ccd158633c1b8b68fdcf53d56d57d48e44ee8833d189ae

Observation e55a50ed-15c8-4340-baba-00de4d55dffa · outbound

This paper cites Egosonics: Generating synchronized audio for silent egocentric videos.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Egosonics: Generating synchronized audio for silent egocentric videos

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.339603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.924581Z digest=sha256:a807b242305442ada285e96b5143ed34ebb19b7f1b6700686f1e46f985296d38

Observation f01b5d7c-4379-4464-bb4f-9c504a24c72e · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.328331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.928332Z digest=sha256:bb90c943f5dcd54a3fa18c02b8f8c56cbdd7b0f88d858121f7e971f3b73972dc

Observation 3d927eec-728b-4229-b022-3f745738ec2e · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.931990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.931990Z digest=sha256:c53c136c8fdf134144039b2f069da671491147f922eecf3ad2e962691aaab5ef

Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · outbound

This paper cites DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.083418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.936928Z digest=sha256:3f3ead3a36e80c7ae2a5daaf28ece43336fff7b7cdc542bbad9e1031c40eee8f

Observation 07986a81-9e19-4ab0-b216-8451bc314589 · outbound

This paper cites Generalized end-to-end loss for speaker verification.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Generalized end-to-end loss for speaker verification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.305063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.945502Z digest=sha256:cd908fdf7136f6844f299c1e8755474ba17e62628ee426353e89312e1aa80889

Observation 85d6e32a-0eb4-48bd-b14e-ec1ec0be758d · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.952467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.952467Z digest=sha256:b36b126e6edc6e2793f9bfce6b68f5022a7436a35f8d1b8f78eb96b04bfced92

Observation 479b8915-b319-4b6c-9bc6-c948f19904f9 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Qwen2.5-Omni Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.956434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.956434Z digest=sha256:907ee3c8af1aa952a07883014e21be4fd486ae2e8b6f0fd4a071fb222598af69

Observation 3d9ffcc9-19c2-4fb1-9bcb-88c4668eab94 · outbound

This paper cites Forecasting china’s regional energy demand by 2030: A bayesian approach.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Forecasting china’s regional energy demand by 2030: A bayesian approach

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.293381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.960482Z digest=sha256:75266dc4f74fcc9add7866d1dc781ae51fedbd96443e65abf9247b1d4a7df89b

Observation bc06c7c8-a161-4847-8105-a415e8d1c1ba · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.963674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.963674Z digest=sha256:0bbbcd149057cc869bbd0c21895746cac20fff333019bfc57c9a01689ff62bad

Observation 7e2735b9-945d-4da9-bf83-a8d4ac21bf7a · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.971825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.971825Z digest=sha256:cd088d1ab88220b0f3bb360e8eb79e93a9bd4fd46baff461bcc4d2dea5686df0

Observation 0d639cc2-a4a9-48bf-9a87-a1675c0146ce · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Synchformer: Efficient synchronization from sparse cues

Reference 2003

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.382529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.887959Z digest=sha256:2c7396518dcf0a46926676c6866842704071400eedaa5f91e7310d75c53691e2

Observation 1dd81864-a54b-460f-8af5-3e6306a72611 · outbound

This paper cites Temporally aligned audio for video with autoregression.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Temporally aligned audio for video with autoregression

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T06:04:30.317525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.941955Z digest=sha256:226e846ddd41539b65b4e016416d00965be58cfa8e9af59a0d36172131f581f6

Observation a8811ba4-f8e9-4233-bc25-30129c684a9b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.948925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.948925Z digest=sha256:81a0b539c98221b91287ceccfbffc872860add9337728948558b8cca65559f39

Observation cf95400b-5535-4a27-8787-6e8f5fe34cb5 · outbound

This paper cites DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.967160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.967160Z digest=sha256:db59ddc69391d7316b6f3c4f45be57c6d0e6dfb6e78713730940b1d7b1fddfbd

Observation a9f25310-eeaf-42b5-a2f9-e18080661d87 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.398373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.398373Z digest=sha256:ad3b8a12305aee31cf60f639694d5f496760263a300a94fd052a79f32d01ac1c

Observation 30c34803-a6cb-4e46-873f-06920348492f · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioGen: Textually Guided Audio Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.899392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.899392Z digest=sha256:f4295819eb62c029a6089adcf804b9d2ef1de6f769bfc7dd21a1aa342c2a8aa3

Observation 46465fa8-6932-426c-8b48-9bfda381a29b · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.902919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.902919Z digest=sha256:6577d15fe259ee1d551b975a63faeff52f140b0ebb350c17fa600cc65a1e9278

Observation e37e05cc-f7ac-4ffc-84c7-b3678e1ae25c · outbound

This paper cites Jukebox: A Generative Model for Music.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Jukebox: A Generative Model for Music

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.619107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.619107Z digest=sha256:5a8fb1e29d5b38eb8e8449e297a691b127745d756b9f66dadea1eda2c8deadd7

Observation c770a0db-f760-470e-a2bc-9da33e913861 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.065491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.065491Z digest=sha256:0f08e077e53ae2c2f67f9ebe119cac00d760d0e3b69371748b9f1acbf8896d4b

Observation cc18a5d3-7b85-4190-9dbd-4c6d356f4faa · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Intelligible Lip-to-Speech Synthesis with Speech Units

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.235068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T06:04:29.497216Z digest=sha256:38c4e0903948a229c3c71ab4a17671a28fc19f66326a21e54e1e2659f478825b

Pith citing papers

Observation 5db16c02-bee4-4b43-8764-32f82764cc54 · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:58.872482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:58.872482Z digest=sha256:a1e5ae08afeb7fdc7b3f8a1635e94734398cd3d56c04f1c388c01e2722910668

Observation 92cd37fc-20d2-409e-97b6-218e161ecbb2 · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.031207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:3d44d6c6106798be7fd09f7d62a38469bb9ac7b4c87f91571fc0f733c31dae5f

Observation 3cdf2ce2-884a-4cdd-8752-e50636cef16e · inbound

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories cites this paper.

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.838588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:02:52.748220Z digest=sha256:14e0428f4535873fbd54fbfc7bad5d471ba1c751b65847d9f647166bee8f34db

Observation d4725871-7853-4a8c-b72c-0e3b8675d5eb · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.060801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:2ec319b6148a9091e7bcdbf1760b5573ec5497b6322f128625d975b1ae9e5667

Observation 3baa02ea-0a50-43c0-be32-347d30bc946a · inbound

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer cites this paper.

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:16:12.355982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T21:17:48.886421Z digest=sha256:6824c080e8197ec653115bc774de99292b69f40ef8ba86e4060d8a1ed9e0c926

Observation 1ae5565b-83bd-4351-8c3a-fa0453af18b8 · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.615037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:2089f7123ed697055fb292550ba1faa813cd5308b8b52af27b13743c77d997fe