Pith. sign in

Paper Citation Record · LEDGER

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2412.15322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15322 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:27:03.800014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:07:14.501348Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5cecb5b1-87b6-46a2-932a-2cc195f06d3d · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:03.800014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:27:03.800014Z digest=sha256:cf72b193d389bb3705c577859ae28236d2361eb2f81ee91045e84cf726a5e584

Observation 21929990-4d8e-41c6-992e-2b69a4e41bde · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.641742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.641742Z digest=sha256:b668344a00d570e8ea6a634d071a72463eb4ae3582c9a4be9fd3f713721bfb8c

Observation 91795adc-1b3d-410e-8887-008a4e0d6bb9 · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:07:14.505181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:c6c6f71e2aed47a83312fb691b4ec7a3c6dd71e57ac24d62adb8d4b2ac62a3fb

Observation 4ba353cf-9ed8-4907-825a-6f6705aa76b4 · inbound

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet cites this paper.

SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.999597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.999597Z digest=sha256:694ec988d5671d4c73d94670c9f7b4ad7208c12184370c37433e7b84ced5b87c

Observation 2309e642-f146-464a-a083-36fb2b903a66 · inbound

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks cites this paper.

Towards Video to Piano Music Generation with Chain-of-Perform Support Benchmarks MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:09.943750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:09.943750Z digest=sha256:be8222834bff44f56d16653334da4943fa3a1aa5df1367f448ad65180ee91d9c

Observation 30f33d15-5961-462a-a3c1-2596da605e46 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:53.035618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:53.035618Z digest=sha256:3e66e29a0b2c7cd11227000f3234041726b7de2d4a5c46b96af8fdfb61feb2f6

Observation 14e806fc-87e0-48fa-9c31-79ec9b6b96c0 · inbound

Sounding that Object: Interactive Object-Aware Image to Audio Generation cites this paper.

Sounding that Object: Interactive Object-Aware Image to Audio Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:14.895706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:14.895706Z digest=sha256:83612d106d677f86312455f3ebf4173b0725e1f94b9932e092bb301f89a3d4a6

Observation 713fd661-c934-40de-b955-c4023f195efc · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.686519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.686519Z digest=sha256:193def6a219b51d10d5aa7ff04e638c05987996070122b77712d861d3797da06

Observation 7244bc48-d4f4-41f9-9ec9-caa3a12af7b1 · inbound

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips cites this paper.

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.764241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:57:21.434793Z digest=sha256:e2c605eb415adacc57045e9c540c70bb238b6b493f8f746aebb0d55afcfb2d0d

Observation 19f406e3-594d-419b-8900-c45e54a35f16 · inbound

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation cites this paper.

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.093694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T16:11:16.551910Z digest=sha256:1666411506bacb9e07835fc779126a08484e8afe6de96ca854d4fd6775386c24

Observation 85d98edd-c600-454c-a9dc-983ee4baa906 · inbound

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation cites this paper.

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T12:34:20.057072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T12:34:20.057072Z digest=sha256:868d62097ec31c4ba8145fff21e28c28d07b3b5086fc622fa3ff56424f14d85b

Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · inbound

KVAE: Family of Tokenizers for Multimodal Generative Models cites this paper.

KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.957682Z digest=sha256:68896aed9afb49bae60a69768fe249bbf0822a96dc0b52a1574c275b7b72b35c