Pith. sign in

Paper Citation Record · LEDGER

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

As of 3 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2604.04348.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04348 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T20:23:36.774359Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T23:45:08.232528Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact5
  • verified fuzzy42
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d5871b25-db4c-4d8f-bd79-cdf3399182e8 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text LRS3-TED: a large-scale dataset for visual speech recognition

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.511047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:050508493139c73799361ffed91ecd9a8e4c56d194fd68401a1a7ccc6f099a28

Observation 96e47384-3e35-472c-abac-a350b4d9b12b · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.369239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:2c82257af1dded6b68653cfda5fc4e2fd5b17dc636b1fa83bd1a148f7181722c

Observation b050bbc0-4049-4f9b-b088-79ca37476313 · outbound

This paper cites Com- mon voice: A massively-multilingual speech corpus.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Com- mon voice: A massively-multilingual speech corpus

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.466475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:f7e30860c0c7e03db4cdf5ba15a46c1bd9988d7c244335abfe5cbf082547132a

Observation 8aa19022-5f78-464a-a6bd-4c127539ca8a · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.453117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:577bc38d88a061f41f436b255d4127a6371e89118c104011e07745b2b70aa01e

Observation 6eca2bf2-3824-4c21-af5b-47977ae1bb15 · outbound

This paper cites Video-guided foley sound generation with multimodal con- trols.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Video-guided foley sound generation with multimodal con- trols

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.457265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:1bd49f3bb07551fc0a96d4f87e3c06f2e6d34a06c6ea5c337f5b11b7d3752e77

Observation d283032e-c15b-4255-b48c-e83ca5528e47 · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Mmaudio: Taming multimodal joint training for high-quality video-to- audio synthesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.461838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ee81f19bf4523c6efffab09bea4deec92c40b98826e4ad699f7a4ead544b2476

Observation 2cd59036-3345-4036-9612-ada8d7506fef · outbound

This paper cites Scaling rec- tified flow transformers for high-resolution image synthesis.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Scaling rec- tified flow transformers for high-resolution image synthesis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.470968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:265c25ee3160f05f985886adf6506a8cac710add986e7a5b3844df1139d796a3

Observation 1cfe1090-5b0c-4c39-9d47-c81fadba63e7 · outbound

This paper cites Text-to-audio generation using instruc- tion guided latent diffusion model.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Text-to-audio generation using instruc- tion guided latent diffusion model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.475168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:06924961e79d2db93320f36d8457e03f9c90ac848dcf71795a296e1988132d69

Observation e4793e25-34b8-4a4d-aca4-f520c8c9e2a9 · outbound

This paper cites Classifier-Free Diffusion Guidance.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Classifier-Free Diffusion Guidance

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:00:47.536659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:1828d76c004e0d36fc9a133d45eca80c61937f731b7cc876ddb6ce551a17ede5

Observation 7088374f-c118-4150-a1cb-da42a3d21160 · outbound

This paper cites Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.426814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:adeb488ead97242252067e5e4ba69ce1e19956904fcd3343e258b02fe503d5f4

Observation 943b4866-f1cf-4a92-9882-25f341233f4b · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Imagen Video: High Definition Video Generation with Diffusion Models

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:31:08.340948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:613c87741d3f97ba7e383ac92ed5033f6517b170be3a83768896f981d76967b9

Observation 9c0a7983-c51b-44f6-a298-474dac8d5b04 · outbound

This paper cites Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Video dif- fusion models.Advances in neural information processing systems, 35:8633–8646

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.431154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:3697797aa885ef2f7ddcd446f14c45401341e4959fcda9c9ed979d5fa6d9d078

Observation d41a35c0-fe4c-4622-9a37-6521f67dab12 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.417646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:26b14755f36a091276f3e265056a683dfc6da9d7d87e35f6fe136403decf46ea

Observation 5b40324d-9df9-4461-a02c-04ef2c347ade · outbound

This paper cites Taming visually guided sound generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Taming visually guided sound generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.413435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:0b5648b56d0af61ece5edd706fcb321dac396299e6d1badc69beb15a708142cb

Observation 42c24fd8-13db-4852-8e03-fbc0d474231c · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Synchformer: Efficient synchronization from sparse cues

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.421687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:f63d1c329c1f3487f6e5d274c6093b469744ff6d7fe1ffc5e8c226816f7085b1

Observation 900a516b-0d04-4db6-aa16-b1618722252f · outbound

This paper cites V oicedit: Dual-condition diffusion transformer for environment-aware speech synthesis.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V oicedit: Dual-condition diffusion transformer for environment-aware speech synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.435228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:95ed5dcfd4abeb46ee572934d2ad391a180e171e99ac18ea8e46d7ad1a315678

Observation 4d5e1765-f75b-413f-ba56-231900503142 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.502823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:58e22b1e9b42abe9b347e0c74854b362a4d243c798a8fea05f7052f18b678aa6

Observation 8a79be29-c6d0-47ec-8fb7-d2ec2d161185 · outbound

This paper cites Guided- tts: A diffusion model for text-to-speech via classifier guid- ance.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Guided- tts: A diffusion model for text-to-speech via classifier guid- ance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.439435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:e34dd3956994fcd997a010b601c3647c9d35ef3bee55a9d6bfed6020de4a35ea

Observation b5c1bd87-780d-4943-92b7-260f673eccf8 · outbound

This paper cites Kingma and Max Welling.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Kingma and Max Welling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.409325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:2301696efbe8c47ad975673549ff37b9b496bd2109fc2e9ce6b7e6e44700997a

Observation 71d2c3fd-9785-49e7-9466-e673f878638a · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.Advances in neural information pro- cessing systems, 33:17022–17033.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.Advances in neural information pro- cessing systems, 33:17022–17033

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.444272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:c5f80ba893503ce5a3b68e3de460dc34608831939f6e24ca0113e3de132d945f

Observation 1a8ed80d-aa17-4abe-abab-26924a3a0a05 · outbound

This paper cites Audiogen: Textually guided audio gen- eration.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Audiogen: Textually guided audio gen- eration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.448850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:30bfcfb50793e9bca5be30362a4a3e1cae279452028a75ae194b18f031d1b40a

Observation 8eb1b796-b397-44d3-99f6-7e577d08ff61 · outbound

This paper cites Vintage: Joint video and text conditioning for holistic audio generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Vintage: Joint video and text conditioning for holistic audio generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.479262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:a5538d41590039c9869bb244334d96fcd4cd09519a9d92f52ab2025eecbd4b4a

Observation 49933331-aeec-4329-8fe2-24fabd908b14 · outbound

This paper cites an unresolved cited work.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-16T01:37:06.396482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:20c00fd55407315cf5259a2ad4529229e4536f5246cfd311719fa0f547a3509f

Observation 0401f434-9b66-4ffe-b986-a327a8d41182 · outbound

This paper cites V oiceldm: Text-to-speech with environmental con- text.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V oiceldm: Text-to-speech with environmental con- text

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.387377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:3729f1e94a68881a001fe057835e5769cc210ccca80067422bfdcd305ef0935a

Observation 2bc6ba92-01c8-47d1-8afd-b53d37b3e81b · outbound

This paper cites Tri-ergon: Fine-grained video-to-audio generation with multi-modal conditions and lufs control.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Tri-ergon: Fine-grained video-to-audio generation with multi-modal conditions and lufs control

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.400734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:4f7732e87f9c62277f946e7af38bffc394040d4c92693f6ca013563376cd40b6

Observation b6349fbb-f2bc-4e58-baa9-000c3196dfe9 · outbound

This paper cites an unresolved cited work.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-16T01:37:06.340924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:8bda64a36cef9cf5b5c9594e570a656d9d94cc6152468e8bc4db65c3e9810097

Observation 6bebdc63-228a-448e-8733-65480917d072 · outbound

This paper cites Au- dioldm: Text-to-audio generation with latent diffusion mod- els.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Au- dioldm: Text-to-audio generation with latent diffusion mod- els

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.345125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:9b2cf544d8a5b6976c65b1c8f58859217695203d3c11f4675627162e3e660132

Observation 4c3e5207-7ce4-4266-9df5-bc8887d6c3f2 · outbound

This paper cites Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Audioldm 2: Learning holistic audio gen- eration with self-supervised pretraining.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, 32: 2871–2883

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.373949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:b981a1a16c022184c96e336e6ea338070894a3bec2255c7831b320aa925ea0e9

Observation 85394fce-e6b3-4583-8371-1cd84c11e6bd · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.325402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:cdc6b935e9144902b039a847cd1d23d025028ee1dc6c084502aab404dafe7249

Observation 17f8027b-7823-4084-8e0e-4cbfee6274ed · outbound

This paper cites Decoupled weight de- cay regularization.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Decoupled weight de- cay regularization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.333215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:10caaec934076f73d29b53af34b8d4085cf31b593c2fd8ef474b5614284f2450

Observation b388d655-bc4e-4863-a8fe-90016b242897 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models.Advances in Neural Information Pro- cessing Systems, 36:48855–48876

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.321582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:f531e041bffc9af55cc444159e189a8dca37bbde74b89f150685f4ae031bad7b

Observation 5cf8fa71-ef8f-40a9-ab8a-45d733574964 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.329158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:e72a957c55b4af7bf753ea3f0b9522ed1e7019e4c0320255f42dd7146781dcc9

Observation 67070e8e-2bf4-4f95-aa9e-d6c6a8ecda85 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.317122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:a8f97658a285e41371d673dc03ae79865e494220e8bccb831f3e149031215856

Observation d9a3efc5-8c9e-4cd0-88fc-331eda484d60 · outbound

This paper cites Scalable diffusion models with transformers.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Scalable diffusion models with transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.312768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:c0d2ff32e15c75a94cdfa1ccff6c4042d13f49676921ecfec4fd7d68143bda94

Observation db27bd8a-e00f-4aae-b0d5-e6578a1b8a81 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Learning transferable visual models from natural language supervision

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.299666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:ecbed1cd971635e1025033389c0e287853224ca7f24aab7b0e1cffe988d1ec6f

Observation b8a5b66b-b1ce-4a39-95ee-c8ddca8467af · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Robust speech recognition via large-scale weak supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.303958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:b59c8abf3e24951f01cc6cf74e8765f0d3ee7ce5a867ca79781f96dac2d7abce

Observation e8695b06-01ef-4ff9-89ae-83c5f70cc6d6 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.308183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:fba16960ee7b2cda470d102427f5f1682eb03d1a8510b180ebbe2cff8f4a7587

Observation 0b3f25cf-3ae0-4bfb-8a04-bf26a5b63826 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text High-resolution image synthesis with latent diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.337089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:3d227f3a7bb82f00297d3ac08152be086463abde7e0bc766c781490c14f2cfbe

Observation 11a95843-eead-4cc0-a314-db33cee3c359 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.496754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:6da98f977669bd394ca660300b611ed16e64b8daf86c3390f678818720bff827

Observation e3f2adc0-e94b-406a-b0aa-db1d6ff79c72 · outbound

This paper cites I hear your true colors: Im- age guided audio generation.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text I hear your true colors: Im- age guided audio generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.349434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:281106b4a9b6bf1bd5946de4bc9bf785f0293aeb93ebd5cc585934c830cbbee6

Observation dbe4e294-7924-4f73-b6e8-a00a45ce0ea5 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Make-a-video: Text-to-video generation without text-video data

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.483737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:9c8a8a9e573cdb12fff88b5406db74449d8511455db7da1141ea6abd30b6acf0

Observation a4d74983-a2cb-4444-aa8c-d3c32bececf2 · outbound

This paper cites Score-based generative modeling through stochastic differential equa- tions.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Score-based generative modeling through stochastic differential equa- tions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.378333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:c9cb47b5d8f4b08142f29a4700a1843f94040f75457924a8b38a693d24595142

Observation c946163e-c7ca-45fa-8784-d5988bcf4ad9 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.383158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:fb51a61da94799e043c678be08774ba203ed7f78cb172a992e71b6520337c618

Observation 2f52d7cf-c424-4d50-b040-5e629ace856a · outbound

This paper cites Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6):4234–4245.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Naturalspeech: End-to-end text-to-speech synthesis with human-level quality.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(6):4234–4245

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.392337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:11f908b36acfcdced81df637b13527d17241454a5221054ad39eccd8854f1e7b

Observation ab2a469d-4513-4a1d-b1ce-c12323d55708 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.405129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:7f325276cf0ad4bfe2450a093782eb4208725565d5c204b896091436a3930064

Observation 175a2779-c67e-4365-b733-3e4900bcdfd9 · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in neural information pro- cessing systems, 37:128118–128138.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Frieren: Efficient video-to-audio generation network with rectified flow matching.Advances in neural information pro- cessing systems, 37:128118–128138

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.353573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:12cfb78b0570af0bfb4c5937dc9380ac2a8500514485ce982f8dd6d3e158da97

Observation 9ff5c10b-551f-4cf5-9ba3-4911237e2de8 · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Wav2clip: Learning robust audio repre- sentations from clip

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.357445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:31529d05d3fb156e835d8bb7c31c09ee89c002058c82e48a89023abe0d8fa754

Observation 5b64560b-b850-4f77-a50f-801584deec55 · outbound

This paper cites Son- icvisionlm: Playing sound with vision language models.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text Son- icvisionlm: Playing sound with vision language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.361354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:81568ac19e9ed30348ee752bcb77654927063628c4050a1ce60928efd74be5ab

Observation 652fc853-fccd-41c3-a7cc-f63e44889002 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:47.523966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:e03739ec28865214e5bbba5ac8ae26728dd7dc45ee04ed8f3af11635192c9790

Observation b6f7e65f-b241-4309-b285-fd1a76bfcdfd · outbound

This paper cites condition–unconditional.

OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text condition–unconditional

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T01:37:06.365145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T20:23:36.774359Z digest=sha256:cfab6c2529389d7f037a5abf9eac0e8b69915eba058a2c94b02744e6bd19ee7c

Pith citing papers

Observation f703365e-a69d-494b-bbf3-5a12f52a8deb · inbound

Do Joint Audio-Video Generation Models Understand Physics? cites this paper.

Do Joint Audio-Video Generation Models Understand Physics? OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.233900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f2f4f003755c9ccbcb27099251bd3c1fdae5e248cf2033edb4776dab3a70b548