Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:43:01.110656Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.10768.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:43:01.110656Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.339998Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T20:27:03.913459Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dd153951-d89a-4a29-895b-f5070adee5ff · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation The Foley grail: The art of perform- ing sound for film, games, and animation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc69f13e-4ca3-4858-a248-4a172f5a6604 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation In- structpix2pix: Learning to follow image editing instructions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 394a27a7-aabb-408b-ba07-37715ce35c5e · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Video generation models as world simulators
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0963105e-5194-4970-b34c-0f821db9a0f4 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Vggsound: A large-scale audio-visual dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f456e97-c8d9-4bed-a688-96db9796cdd3 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Gentron: Diffusion trans- formers for image and video generation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7504d809-a4a5-4610-a200-d4f02bb44391 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audio-vision: sound on screen
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 42357cef-9a0e-4976-b4d7-37d60620d868 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scaling instruction- finetuned language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70a28615-c939-412f-addc-4ce07851924a · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e1b2c8-5007-4e4c-a5a6-da577476013d · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9bd243d-5c75-4739-afa0-c72c3a3200c4 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a8ec89-425b-4bb6-824d-e9428d88aac6 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Imagebind: One embedding space to bind them all
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe58a50c-5079-4ba4-9d18-971491a11c80 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Determining op- tical flow
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81536570-f0b8-4950-ab39-2777285f32f3 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a79237d-0f2d-4eb2-b0a3-fc343e6218b1 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Captivating sound
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43da5ea3-8cb9-4f97-b8a6-e16b0aee59ba · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Taming visually guided sound generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7bbc4dc9-61b1-4ad3-812f-df8ad7362d2c · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Mixing audio: concepts, practices, and tools
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2432f03f-54a6-4d14-abc8-427a56b07f0c · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Read, watch and scream! sound generation from text and video
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f01f30cf-82ef-45fb-bb91-d7747fc80aba · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3133d9e4-9647-48bd-bcd2-f71ab0d354f0 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 75b0328f-621c-491d-a3ed-870cadf9102b · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation AudioGen: Textually Guided Audio Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d1f91e-193d-48cb-a0d1-73980d2d4881 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b8b58e80-68ad-478f-a8f0-d83a2c93cfa3 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V oice- box: Text-guided multilingual universal speech generation at scale, 2023
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 67ef12d0-af5a-4b55-9d23-416a97f5b652 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Matching for Generative Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bddc6b1e-a04d-4b05-a33e-ea5dabcbdd7d · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audi- oLDM: Text-to-audio generation with latent diffusion mod- els
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 78e9addd-a28d-476c-9917-4230865d43b7 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8892f56-768c-46bc-bfca-c8d4807d0c24 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Plumbley
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8fcadfb4-164f-4c07-b2e9-6d9c8ec496b8 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d0d2b7-7e5e-410d-935a-c9cac5453cd8 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Separate Anything You Describe
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d314d68-d31e-46f1-b28a-60ad9c7dd7e9 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 19bb9429-a598-4d77-8b51-38262cfc3520 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Albergo, Nicholas M
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3c02fe20-b9e2-4a7f-9b7a-4668f091d418 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d496671b-d7de-463a-87d5-87e22465082e · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b1c667-3311-4ac6-be88-5c8557f19490 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation Synchronized with Videos
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c193d648-be82-4cc0-9ed6-b7975fab2e0d · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Bal- ancing act: Distribution-guided debiasing in diffusion mod- els
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 060323ee-af08-4265-8a17-867538b01346 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scalable diffusion models with transformers, 2023
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 341338be-00a1-4f5a-9cce-1df26449df12 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Film: Visual reasoning with a general conditioning layer
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91a4fadc-1435-4907-8ed2-f3cc48242973 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning transferable visual models from natural language supervi- sion
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b585338-d1ed-4056-9f83-e666efe31037 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fe98423-9360-4269-92b1-00befebb6bbe · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation High-resolution image synthesis with latent diffusion models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757ba374-dd75-4de7-a979-d360d2eb4d65 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation I hear your true colors: Image guided audio generation, 2022
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 42c483e9-14f8-407e-bf7b-ef23dc79e883 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Auto- acd: A large-scale dataset for audio-language representation learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b5e2a8e2-082f-4ce4-b38f-11c8bb590245 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning from Between-class Examples for Deep Sound Recognition
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9159ca6c-bf3d-4acf-b2aa-ae3f5fb14cb0 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Attention is all you need
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff84b15-706b-4178-a31c-c2e78cff4f93 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audiobox: Unified audio generation with natural language prompts, 2023
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 44a95349-babf-4205-b7ce-e6d28e03dc31 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5a5265e9-f206-4211-991e-1848bed6af07 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 559861e4-be63-48f3-a037-448bc3afc367 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Wav2clip: Learning robust audio repre- sentations from clip
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8c49149-27bc-4322-8696-04682191ead8 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d7192d4f-08a4-4c10-b83c-894b0be99496 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Son- icvisionlm: Playing sound with vision language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 436724e7-41b5-4930-a16c-78d0a5a465e6 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b8802d4-dd19-427d-b44a-c1a00998902c · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diffsound: Discrete diffusion model for text-to-sound generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1403e556-2132-4faa-b35f-1318ecaba71c · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 079eb143-1244-4a61-9e47-b3b193f10348 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8bafe478-14c6-4828-83f8-99f37de51dc7 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Llava- next: A strong zero-shot video understanding model, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c84ec568-408e-435f-8846-805b45ad41f6 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 10fc6dca-0337-4a2c-b5dd-4542c2b61058 · outbound
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d52cb8-48a5-4be9-989e-07211047f4bf · inbound
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa0d4c0-e44a-4942-bc74-9a9d04cb9093 · inbound
Sound Scene Synthesis at the DCASE 2024 Challenge VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.