Pith. sign in

Paper Citation Record · LEDGER

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

As of 20 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 1 inbound Pith citation observation for arXiv:2509.06389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06389 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.585912Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T23:46:32.435054Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T23:46:32.951229Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3600fb9b-cce2-4bb8-a054-c5e367bbf70a · outbound

This paper cites MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T23:46:32.956166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.435054Z digest=sha256:367ab1ece5cf68c52a1ee00626af0c48017307a4f1c15455d6cd73593b73b851

Observation 03f5461a-ad39-404e-a58c-f6d106d9c3ad · outbound

This paper cites On top of this backbone, MeanFlow formulation is introduced which directly models the average velocity to enable native one-step generation.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation On top of this backbone, MeanFlow formulation is introduced which directly models the average velocity to enable native one-step generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.363058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.440419Z digest=sha256:614111f0da5d8656f20f84cfd16937f361524fa7ea0992ecc0f8f2458456a057

Observation 3fd227ec-169f-479e-8daf-41354811e4b9 · outbound

This paper cites Multimodal Dataset The proposed MF-MJT is trained on multimodal datasets comprising both audio-video-text and audio-text pairs.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Multimodal Dataset The proposed MF-MJT is trained on multimodal datasets comprising both audio-video-text and audio-text pairs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.348985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.444989Z digest=sha256:7f2692a2ccf43b5ff559e1020ed94f4a9e9dab248a4085093086c6f8ca48b55e

Observation 5f345c19-dbae-41e8-8956-f0fc65fd41f9 · outbound

This paper cites Comparison with Baselines Table 1 summarizes the performance of the proposed MF-MJT against representative VTA synthesis baselines on the VGGSound test set.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Comparison with Baselines Table 1 summarizes the performance of the proposed MF-MJT against representative VTA synthesis baselines on the VGGSound test set

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-04T23:46:32.934088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.449742Z digest=sha256:769d02bdfe7d1066bc1e930ac9236e9116b9bb6bfbf273194f0ad62b98bdb97a

Observation 5f5c05e2-f170-4adf-a37d-ab118d5c469f · outbound

This paper cites an unresolved cited work.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-04T23:46:33.334130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.455488Z digest=sha256:3db066a52e717febf1d79b3b9b270f8637ece30888001dad7ecc5f8447a4c66b

Observation a67df04d-e494-46b0-a727-aed4387b61e6 · outbound

This paper cites Frieren: Efficient video- to-audio generation network with rectified flow matching,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Frieren: Efficient video- to-audio generation network with rectified flow matching,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.319821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.459826Z digest=sha256:74f56752eeeeb681493a2b223c60fc9e3972203eb4b4f2bc9e4124d7f804219f

Observation b2172030-e743-4194-8873-e560323e78b9 · outbound

This paper cites LoV A: Long-form video- to-audio generation,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation LoV A: Long-form video- to-audio generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.304827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.464269Z digest=sha256:0e68e6e34c266373858790a0cf8ad894951db6398c7c216481e924fec7075ebf

Observation 042929a9-f966-4f2d-a360-4236041ad7a7 · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioLDM: Text-to-audio generation with latent diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.289701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.468413Z digest=sha256:dfea7218f0b953dafce81d15f6e57372707f1ef7ee63edca48bcfafb7d5842aa

Observation 47b97ae1-b526-4cfb-a4a5-55c4653ba99c · outbound

This paper cites ImageBind: One embedding space to bind them all,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation ImageBind: One embedding space to bind them all,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.274510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.472472Z digest=sha256:268d02ce01639a170983caa654e094216a9a76bf2209ba309deaa0697c79c3f2

Observation 84baceec-4072-4255-a76d-e439ee6b142b · outbound

This paper cites Seeing and Hearing: Open- domain visual-audio generation with diffusion latent aligners,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Seeing and Hearing: Open- domain visual-audio generation with diffusion latent aligners,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.257882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.476656Z digest=sha256:890e0afdc417e493ac20d253a52143d634f15200584810659b124ff0a4ece3ef

Observation 0dcbd26a-08fd-4884-adeb-b9529976d844 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.481088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.481088Z digest=sha256:c0e463e0f1b12bb09943ba34a8f6c681a2197b149a027c89e98536c652ea1012

Observation 6b767ada-39d8-4d89-adaf-751159b082f7 · outbound

This paper cites TA-V2A: Textually assisted video- to-audio generation,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation TA-V2A: Textually assisted video- to-audio generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.242030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.485945Z digest=sha256:a7668f1f63d05bed59a7a88e89cc529431b0aa12848614074efedc2e4720336f

Observation 41ae4ce0-e219-4ab1-9cbe-e6fbf2058de3 · outbound

This paper cites MMAudio: Tam- ing multimodal joint training for high-quality video-to-audio synthesis,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MMAudio: Tam- ing multimodal joint training for high-quality video-to-audio synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.227680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.489971Z digest=sha256:b2d95f757580941bcd7197554e77c3e0b96db499f216d470eefbb55bc7e51662

Observation 30f5277f-ba65-4eef-bdb6-289bf375842d · outbound

This paper cites Kling-Foley: Multimodal diffusion transformer for high-quality video-to-audio genera- tion,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Kling-Foley: Multimodal diffusion transformer for high-quality video-to-audio genera- tion,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.494196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.494196Z digest=sha256:c551be55b9d89d77ae645c24b5e44ca5f98f671c38b7890917add29e2e21c7b7

Observation 4e0faab3-cb83-47e5-9dbc-ff112b26a7bd · outbound

This paper cites Denoising diffusion probabilis- tic models,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Denoising diffusion probabilis- tic models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.212191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.498259Z digest=sha256:47dbea8331e89a2e655f62e6884d583c96665dd0ac22ed50ad3de485650fd327

Observation 2bb2793f-c683-49cd-bbe6-6032eff702ec · outbound

This paper cites Flow straight and fast: Learn- ing to generate and transfer data with rectified flow,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Flow straight and fast: Learn- ing to generate and transfer data with rectified flow,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.197545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.502378Z digest=sha256:683220dfa73cd7b03317b8310709b30e6d530b9eff4b52e6cde2377adae2dc32

Observation a9688b03-26c1-46ab-b37e-091350255898 · outbound

This paper cites Instaflow: One step is enough for high-quality diffusion-based text-to-image generation,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Instaflow: One step is enough for high-quality diffusion-based text-to-image generation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.183440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.506431Z digest=sha256:abc525a26170120adcf14f51e7510a9be84d8e433786f236a7451960221b3d5d

Observation 896d1435-b761-40ad-a5be-728234741047 · outbound

This paper cites Mean Flows for One-step Generative Modeling.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Mean Flows for One-step Generative Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.510295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.510295Z digest=sha256:c47ab9d5600714d1b8a364ce61bbb2289da97f9888e3bd061c81abbe65dd6e49

Observation 1c1debeb-1c45-4715-bfe5-e3a9505e361b · outbound

This paper cites Classifier-Free Diffusion Guidance.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.514602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.514602Z digest=sha256:2d098b56ac94465a4365758eff57c5063d907c3e0308e5f074eb10da404e75f8

Observation 992acc49-5a1e-412d-86bd-554aa35a69ab · outbound

This paper cites CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.518893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.518893Z digest=sha256:3a3285cdf60740d7872b7062481fc52724ef6ad08471f69df1546c2569746be9

Observation d78d1673-8707-46fe-967c-4e491df61e79 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Scaling rectified flow transformers for high-resolution image synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.169343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.523361Z digest=sha256:40e2778e8f867f71948ce45fefe14b8d39f962fb764f577988c71b6fac54055e

Observation 56f499bf-339a-48f6-a196-eb479603b0a3 · outbound

This paper cites Scalable diffusion models with trans- formers,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Scalable diffusion models with trans- formers,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.154757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.527401Z digest=sha256:3cb04e30ec248d5f8f6fbb3e1a1cbfa5f8f737fc08427cb0bbc6762de8b1802f

Observation a1303cec-94e9-40a7-8d4f-5d6781368649 · outbound

This paper cites Learning transfer- able visual models from natural language supervision,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Learning transfer- able visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.140434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.531552Z digest=sha256:2e1379a73d95d7113b803ffee947f187eaaaa9c76ed0cdb33b6232d8d09d8cf6

Observation 5079f05a-0c9a-4aeb-b6fc-e0f069be909e · outbound

This paper cites A versatile diffusion transformer with mixture of noise levels for audiovisual gen- eration,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation A versatile diffusion transformer with mixture of noise levels for audiovisual gen- eration,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.126195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.535534Z digest=sha256:1e4f5c8f76e3fb62cab991b695a27237131d266394a0efb93d6d4371fb6fe638

Observation 358179d5-ffae-4870-b6d1-f346144c9806 · outbound

This paper cites Synchformer: Efficient syn- chronization from sparse cues,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Synchformer: Efficient syn- chronization from sparse cues,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.112032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.542615Z digest=sha256:f48bd55299348bcae38dd0def62bccd050e9112868dd2fd60c76eb6431eb0992

Observation 72ba48a8-6683-4bfb-93da-3286d88a10e2 · outbound

This paper cites Score-based generative modeling through stochastic differential equations,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Score-based generative modeling through stochastic differential equations,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.097241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.546543Z digest=sha256:7edf4fb1ba26ab085fb0d770056dea2de42eea327c10e38e0de2561ffa00ecef

Observation f9c9b6c0-aecb-4e90-bd51-744f0d1fa999 · outbound

This paper cites VGGSound: A large-scale audio-visual dataset,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation VGGSound: A large-scale audio-visual dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.083256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.550612Z digest=sha256:6a9b3859fbb32bfad1c720ef8f03bf23d56fc335b648762f4b85a79befae6261

Observation c7ca1149-03fd-48b4-b82f-a105d0d9e8b4 · outbound

This paper cites AudioCaps: Generating cap- tions for audios in the wild,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioCaps: Generating cap- tions for audios in the wild,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.067953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.555203Z digest=sha256:25722f3d5ec587842bc8b2781611475e90b26083f1d7f74ff69494ae7345b894

Observation ba099949-7db9-43f2-9403-2a65f35fabf5 · outbound

This paper cites WavCaps: A ChatGPT- assisted weakly-labelled audio captioning dataset for audio- language multimodal research,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation WavCaps: A ChatGPT- assisted weakly-labelled audio captioning dataset for audio- language multimodal research,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.052749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.559318Z digest=sha256:9522d827be6d4342aa753b9cfed59f6a38cd400e9fdd324974cc9e8e5dbb2323

Observation 11b40170-8da9-4314-948a-d1aadc9f118e · outbound

This paper cites Decoupled Weight Decay Regularization.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:32.563516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:46:32.563516Z digest=sha256:ae85c882f757c12e769ce7df8f3a78f79c885522825442671ab48ce6633a3a5a

Observation d42521d5-2ffb-4f22-a5c0-9e33a9c06816 · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Audio Set: An ontology and human-labeled dataset for audio events,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.038213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.568945Z digest=sha256:54788e31f270201ccb5277e46c7417640f6a121d258f0524e1e843bf7f6f9482

Observation b4f63100-b62e-409a-9b3b-bb3b38fa46af · outbound

This paper cites PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.020868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.573072Z digest=sha256:34b19c5db1b1ca1b3ebf1b2ebb771cdc372f4dc359ce76f6ec599745b95a58eb

Observation b1ebeab0-251d-4093-bf90-e8ce74829df5 · outbound

This paper cites Efficient train- ing of audio transformers with patchout,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation Efficient train- ing of audio transformers with patchout,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:33.005535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.577230Z digest=sha256:a0799b286f668405b31993c368646b6e9d997c4a5974a1794d802d3435ca6adb

Observation d9a819d7-f750-4095-bbfc-7dc47566a11d · outbound

This paper cites CLAP: Learn- ing audio concepts from natural language supervision,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation CLAP: Learn- ing audio concepts from natural language supervision,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:32.989932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.581541Z digest=sha256:58480cbce514779049533ae9b404cc8137eb28f31469b481dfea041ce2d43f56

Observation a7287d51-c426-474d-a050-35f35fa93fdf · outbound

This paper cites AudioLCM: Efficient and high-quality text-to-audio generation with minimal inference steps,.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation AudioLCM: Efficient and high-quality text-to-audio generation with minimal inference steps,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T23:46:32.973426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.585912Z digest=sha256:10c9e363cc65498afc535d1d3bd26806fba308d66a326c0e137ab56ef4d19f1c

Pith citing papers

Observation 3600fb9b-cce2-4bb8-a054-c5e367bbf70a · inbound

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation cites this paper.

MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation MeanFlow-Accelerated Multimodal Video-to-Audio Synthesis via One-Step Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T23:46:32.956166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-04T23:46:32.435054Z digest=sha256:367ab1ece5cf68c52a1ee00626af0c48017307a4f1c15455d6cd73593b73b851