Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T00:23:08.279261Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 12 inbound Pith citation observations for arXiv:2502.03897.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T00:23:08.279261Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.548806Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T05:56:40.900641Z
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eb615f65-4ba6-4ee9-b023-b5e053e3c579 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Scaling instruction-finetuned language models,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f0b09ba4-a94f-4ce7-aa0e-878c7af1993b · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Enhanced visual instruction tuning with synthesized image-dialogue data,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2ab83cb4-5302-43a1-ba55-d45aaeb2241c · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation High-resolution image synthesis with latent diffusion models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ebf2eaa4-9386-4918-89b0-5817a85a3a01 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Label-guided generative adversarial network for realistic image synthesis,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dbde849d-a369-4623-8afc-e976d57b79a9 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8664b6c0-2d65-4958-a1d7-29f336028c52 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Continuous emotion-based image-to-music generation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3f50a790-a285-458a-a68b-6cf1940de7c4 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5910f92a-b6dc-46fa-ab13-7ea114396aa7 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Open-Sora: Democratizing Efficient Video Production for All
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feaac7ae-57ea-42c7-b5aa-f90d78bff2a2 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Make-a-video: Text-to-video generation without text-video data,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 002714fa-c740-45ed-9635-e8fb280aefd2 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Align your latents: High-resolution video synthesis with latent diffusion models,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce0abe09-7a46-43f0-b842-5c0468621d8f · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Lavie: High-quality video generation with cascaded latent diffusion models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 07880440-c68c-4ad9-a1d6-90223bad2a21 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc0aa074-385f-4126-b52c-27a985094677 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Mm-ldm: Multi-modal latent diffusion model for sounding video generation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1c12b223-f672-482f-8076-7c5d1280b56b · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Scalable diffusion models with trans- formers,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7140319c-c2c3-472b-9d74-28f2498a66e8 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Mmdisco: Multi-modal discriminator-guided cooperative diffusion for joint audio and video generation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 78bd0cc2-3f69-46da-b5b9-4234d5abef5e · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eaaf7fb-2fee-45e3-96dd-0f55327083a3 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9f13782-2114-4f28-9752-f3c33f17ac6d · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Foley Sound Synthesis at the DCASE 2023 Challenge
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b07248fb-7b9c-4807-be26-6b34a7feffc8 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Conditional sound generation using neural discrete time-frequency representation learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c7bdd247-464b-4263-b9e2-d78b0daa5d70 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Audioldm: Text-to-audio generation with latent diffusion models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64883c8d-8162-4ce4-91be-88c0f9a550ec · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Taming Visually Guided Sound Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85bdd7f3-6e09-4ac7-a8aa-0412476b1f80 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Con- ditional generation of audio from video via foley analogies,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f8bc7a9e-a1ad-4621-b474-6e489e3f1578 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Diff-foley: Synchro- nized video-to-audio synthesis with latent diffusion models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3650c39e-68d2-4981-9145-f8c25141abf9 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7bd70ed-4be5-4cb0-9df4-5bf5ce4cc987 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Temporally aligned audio for video with autoregression,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd44a722-c613-4d0f-a766-13a72221c16e · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Tell What You Hear From What You See -- Video to Audio Generation Through Text
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50349149-657d-4dfc-b7a3-016504b7f5df · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Frieren: Efficient video-to-audio generation network with rectified flow matching,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 923602d6-db73-44ca-9b3f-e184c7770c71 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Sound-guided semantic video generation,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b4aefb80-8523-4756-b44a-0fd2be59938e · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation The power of sound (tpos): Audio reactive video gen- eration with stable diffusion,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9fe97c13-9c88-4e1c-bdbf-17a4e0228cd8 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Diverse and aligned audio-to-video generation via text-to-video model adaptation,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f49398eb-38ea-4e18-b229-589ac817dd69 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0d57e8-7353-4b1b-84fe-0b7ddc3d78e4 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Denoising diffusion probabilistic models,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ae537c-c41d-4ad7-a043-f2ce4b6e3f48 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Classifier-free diffusion guidance,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d4b9864b-da9b-49de-aaf9-2e154250529f · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Denoising diffusion implicit models,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d14f49b8-6357-4433-a39b-5cab2d52ce44 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Exploring the limits of transfer learning with a unified text-to-text transformer,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa2c650-9b39-4202-99b4-4168df74a94b · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 266eeb2a-d640-4b3e-8d00-02744a31c756 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Vggsound: A large-scale audio-visual dataset,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2fe81b8f-01fb-4973-bf7d-b1b073a6ae4d · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Ai choreog- rapher: Music conditioned 3d dance generation with aist++,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b4451862-8e35-4abb-b5f5-be8b6e6cecb7 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Audio set: An ontology and human-labeled dataset for audio events,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d4e6590-ae9c-439f-93b0-dd4cd9c76806 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation The benefit of temporally-strong labels in audio event classification,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf39378f-b5b5-42de-aaf6-65dea21aaef0 · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Aist dance video database: Multi-genre, multi-dancer, and multi- camera database for dance information processing
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5033d48-6d2a-4146-939b-5c3e6f9f99bb · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23d8b30-0744-4a48-9ac9-f57a8679632d · outbound
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65d3b03-373a-4102-baa0-0f2224156e3b · inbound
JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e3600b3-5b03-43bb-aa81-f95b258b9571 · inbound
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe916510-e314-4523-bf4a-49e286c8d953 · inbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a4669c-fe76-4651-b85e-aec31796905f · inbound
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 750fb1d0-85b6-4f75-8733-d431af5e8148 · inbound
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 195a3271-7eee-405d-a737-f4f90e7b23f6 · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 36151287-89fb-4b24-b8e5-a1965a16360e · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5472edbe-229b-4705-bada-76c3b06531ac · inbound
SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 432bf714-ac4c-41fc-8b02-6979188d94b2 · inbound
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 38385add-b7cf-4a4e-ad03-983f4a143131 · inbound
MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e276bec-38a2-4b47-87aa-98b4275b5842 · inbound
Inference-Time Scaling for Joint Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dd650d56-eda4-4d06-80ef-7bbfaa557f56 · inbound
Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.