Pith. sign in

Paper Citation Record · LEDGER

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 12 inbound Pith citation observations for arXiv:2502.03897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03897 v5

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:23:08.279261Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:16:52.548806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:40.900641Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb615f65-4ba6-4ee9-b023-b5e053e3c579 · outbound

This paper cites Scaling instruction-finetuned language models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Scaling instruction-finetuned language models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.939481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.073414Z digest=sha256:eb3ff0207e3c735f53f9da30768e8ebcf866bc4e0936fc9d99250e235cfe7057

Observation f0b09ba4-a94f-4ce7-aa0e-878c7af1993b · outbound

This paper cites Enhanced visual instruction tuning with synthesized image-dialogue data,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Enhanced visual instruction tuning with synthesized image-dialogue data,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.924486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.079309Z digest=sha256:1ac23ee8538f1c8a9731bfa661f837c8bebd0c1e65f6c361a6621a3359812b6d

Observation 2ab83cb4-5302-43a1-ba55-d45aaeb2241c · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation High-resolution image synthesis with latent diffusion models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.909334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.084674Z digest=sha256:05cfec1cb4837f71941bc2753fe70287f7892e2c65e4ead682a915d60db8c16d

Observation ebf2eaa4-9386-4918-89b0-5817a85a3a01 · outbound

This paper cites Label-guided generative adversarial network for realistic image synthesis,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Label-guided generative adversarial network for realistic image synthesis,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.894367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.090212Z digest=sha256:5010269cfbf51c3d6cb89eba2f21887d62321ee354465f5d859baac8a1f80242

Observation dbde849d-a369-4623-8afc-e976d57b79a9 · outbound

This paper cites Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Audioldm 2: Learn- ing holistic audio generation with self-supervised pretraining,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.095460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.095460Z digest=sha256:807b3ea95e6744a4668dab6a7772f53dfb2e21fd4ee8cbd8564da33ed4b9be02

Observation 8664b6c0-2d65-4958-a1d7-29f336028c52 · outbound

This paper cites Continuous emotion-based image-to-music generation,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Continuous emotion-based image-to-music generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.869392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.100701Z digest=sha256:67a26a82583abf27c56106a16bf63040244efad23558e68ac1a4799788b094c0

Observation 3f50a790-a285-458a-a68b-6cf1940de7c4 · outbound

This paper cites DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.106301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.106301Z digest=sha256:9fada8b9d26ea9a848c1fc13aff3a32f320ca0dffd412cfb079c4c617a0d8230

Observation 5910f92a-b6dc-46fa-ab13-7ea114396aa7 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Open-Sora: Democratizing Efficient Video Production for All

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.111482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.111482Z digest=sha256:7dfbc6bf528c723be7ee784bc4a9fc3435380544aad5a6847d29fd84efe45d53

Observation feaac7ae-57ea-42c7-b5aa-f90d78bff2a2 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Make-a-video: Text-to-video generation without text-video data,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.854523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.116640Z digest=sha256:9d69267f992e80b630d949719e6254d2f816a08354af5f21efd096d58b199212

Observation 002714fa-c740-45ed-9635-e8fb280aefd2 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Align your latents: High-resolution video synthesis with latent diffusion models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.839801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.121660Z digest=sha256:77208428042efccd91da36bb6aa43b61d4b6769f143109146a586aceaa01ad6c

Observation ce0abe09-7a46-43f0-b842-5c0468621d8f · outbound

This paper cites Lavie: High-quality video generation with cascaded latent diffusion models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Lavie: High-quality video generation with cascaded latent diffusion models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.825411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.126362Z digest=sha256:5973deab6807df74bd19a6072e7648e37f49a68a571bc49c836583f8cda22f85

Observation 07880440-c68c-4ad9-a1d6-90223bad2a21 · outbound

This paper cites Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.810670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.131422Z digest=sha256:531573e1f86bd97cc9e74a3d4b810a99445637ce9ca60504411144f813a19a98

Observation bc0aa074-385f-4126-b52c-27a985094677 · outbound

This paper cites Mm-ldm: Multi-modal latent diffusion model for sounding video generation,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Mm-ldm: Multi-modal latent diffusion model for sounding video generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.796083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.136257Z digest=sha256:b0bd4b11adb5e833383fe769face85d6f75b942fc1ba58c7e1d9091287c249a2

Observation 1c12b223-f672-482f-8076-7c5d1280b56b · outbound

This paper cites Scalable diffusion models with trans- formers,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Scalable diffusion models with trans- formers,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.780892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.140806Z digest=sha256:f7a712aa5df8649acd510cea793f4ef300a00442a8f3ceb1451867fea1de3701

Observation 7140319c-c2c3-472b-9d74-28f2498a66e8 · outbound

This paper cites Mmdisco: Multi-modal discriminator-guided cooperative diffusion for joint audio and video generation,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Mmdisco: Multi-modal discriminator-guided cooperative diffusion for joint audio and video generation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.765428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.145376Z digest=sha256:7d51ddfa6a51e4a40d30aa1fb24bfcb4b16992284082d29706950ba5fe842928

Observation 78bd0cc2-3f69-46da-b5b9-4234d5abef5e · outbound

This paper cites AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.150079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.150079Z digest=sha256:4ccab6b1ecaba62a2d9cfb298fbb0e2aacf6f235b9948444994c3a3121949a91

Observation 9eaaf7fb-2fee-45e3-96dd-0f55327083a3 · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.749722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.154852Z digest=sha256:afeaf2cba5e16cd5bc11238af6ca7ebfc6b0ce37a695fd274872887ddb0c0d52

Observation d9f13782-2114-4f28-9752-f3c33f17ac6d · outbound

This paper cites Foley Sound Synthesis at the DCASE 2023 Challenge.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Foley Sound Synthesis at the DCASE 2023 Challenge

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.159950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.159950Z digest=sha256:8f9d0d449b52b2ee97df35e6e5db35982db293eb3800e31dff65141276a7833c

Observation b07248fb-7b9c-4807-be26-6b34a7feffc8 · outbound

This paper cites Conditional sound generation using neural discrete time-frequency representation learning,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Conditional sound generation using neural discrete time-frequency representation learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.734547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.164988Z digest=sha256:e37432347b433fdf0bcf6d10f3e25e0ce533f66011d3b6fb7afee4dcb818cfa3

Observation c7bdd247-464b-4263-b9e2-d78b0daa5d70 · outbound

This paper cites Audioldm: Text-to-audio generation with latent diffusion models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Audioldm: Text-to-audio generation with latent diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.719456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.169478Z digest=sha256:edea02fa597fce47abf07c9a5245958d74e65b66f996fee8214b17e5a92495be

Observation 64883c8d-8162-4ce4-91be-88c0f9a550ec · outbound

This paper cites Taming Visually Guided Sound Generation.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Taming Visually Guided Sound Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.173861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.173861Z digest=sha256:d6296d3c51c5f8da6a86d6edccbf014a64f9c201f6c1fd094160e9f80059e238

Observation 85bdd7f3-6e09-4ac7-a8aa-0412476b1f80 · outbound

This paper cites Con- ditional generation of audio from video via foley analogies,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Con- ditional generation of audio from video via foley analogies,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.704609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.178960Z digest=sha256:3612531c665dc9c8819f3a0e5e52e32b09d5ac13a230ef1e650333363fd0de91

Observation f8bc7a9e-a1ad-4621-b474-6e489e3f1578 · outbound

This paper cites Diff-foley: Synchro- nized video-to-audio synthesis with latent diffusion models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Diff-foley: Synchro- nized video-to-audio synthesis with latent diffusion models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.688866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.183545Z digest=sha256:ce504ad317ae0e249debb01a814cbcba4d2dd043d9924d7693d78bed732fef95

Observation 3650c39e-68d2-4981-9145-f8c25141abf9 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.188006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.188006Z digest=sha256:ba47c23d92781b6f47a01d45500a7d1be67939fffce2ce5dfeb842100ba66038

Observation a7bd70ed-4be5-4cb0-9df4-5bf5ce4cc987 · outbound

This paper cites Temporally aligned audio for video with autoregression,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Temporally aligned audio for video with autoregression,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.672241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.193050Z digest=sha256:e2ae678400ccd59de4ef6a98980673953236350bc2865afbac1f3232acf50e39

Observation bd44a722-c613-4d0f-a766-13a72221c16e · outbound

This paper cites Tell What You Hear From What You See -- Video to Audio Generation Through Text.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Tell What You Hear From What You See -- Video to Audio Generation Through Text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.197960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.197960Z digest=sha256:bfbb74e8a21538018e7f7b3045de6b8486c2d3950c8e372cc201b5f13fa66008

Observation 50349149-657d-4dfc-b7a3-016504b7f5df · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Frieren: Efficient video-to-audio generation network with rectified flow matching,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.657337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.202869Z digest=sha256:ef80c0406da9808ed693aef7da6d4ea80582468b7d385e116d1ce66f708befe7

Observation 923602d6-db73-44ca-9b3f-e184c7770c71 · outbound

This paper cites Sound-guided semantic video generation,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Sound-guided semantic video generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.641726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.207696Z digest=sha256:e388150d8dfed4605bf8021f3c85646344cf66b773952d7b252c8bca4259b8fd

Observation b4aefb80-8523-4756-b44a-0fd2be59938e · outbound

This paper cites The power of sound (tpos): Audio reactive video gen- eration with stable diffusion,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation The power of sound (tpos): Audio reactive video gen- eration with stable diffusion,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.626702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.212365Z digest=sha256:7259fc54c7435024a8fc8ee9ab65e53a168adc16230b04dff118aeaa5987fc01

Observation 9fe97c13-9c88-4e1c-bdbf-17a4e0228cd8 · outbound

This paper cites Diverse and aligned audio-to-video generation via text-to-video model adaptation,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Diverse and aligned audio-to-video generation via text-to-video model adaptation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.611766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.217083Z digest=sha256:c7c95795c6ef910867431c7c7927ad1a8bf1cd172f4c57f7b3d9cb5164ad4ddf

Observation f49398eb-38ea-4e18-b229-589ac817dd69 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.221778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.221778Z digest=sha256:9e5d59bafae46648c8ed005714cd7c84f8a7463f13c9010ec94e5cd0b9763136

Observation fe0d57e8-7353-4b1b-84fe-0b7ddc3d78e4 · outbound

This paper cites Denoising diffusion probabilistic models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Denoising diffusion probabilistic models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.226795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.226795Z digest=sha256:9eff83f22e6a589f0673fff1f234768618b164110185d570d0c83e9a9c2a5e8a

Observation 80ae537c-c41d-4ad7-a043-f2ce4b6e3f48 · outbound

This paper cites Classifier-free diffusion guidance,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Classifier-free diffusion guidance,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.587790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.231336Z digest=sha256:ce44081bcce75cd9076d4d0605183419e587fe22424be59f42c718dc7ab60d32

Observation d4b9864b-da9b-49de-aaf9-2e154250529f · outbound

This paper cites Denoising diffusion implicit models,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Denoising diffusion implicit models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.572749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.236123Z digest=sha256:75e97854d499471a10de7f3fa685d2bf0da4ac2ddfc01ae543e9ab6d6c59251f

Observation d14f49b8-6357-4433-a39b-5cab2d52ce44 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.240853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.240853Z digest=sha256:925b92360a561c4692f7dc35cb2c78d721dc69e748c9818a09def8b33bdb2276

Observation caa2c650-9b39-4202-99b4-4168df74a94b · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.546974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.245669Z digest=sha256:32ac1510e281a1e94fd85a16adfac41ec244438ec4811a35a4d7518c3df5c6ff

Observation 266eeb2a-d640-4b3e-8d00-02744a31c756 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Vggsound: A large-scale audio-visual dataset,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.531559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.250345Z digest=sha256:0506d9beaf1a713a256b3fe8066500c9189896201f15989c809a1cd4b59b1c98

Observation 2fe81b8f-01fb-4973-bf7d-b1b073a6ae4d · outbound

This paper cites Ai choreog- rapher: Music conditioned 3d dance generation with aist++,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Ai choreog- rapher: Music conditioned 3d dance generation with aist++,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.516275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.255030Z digest=sha256:ce260b66e7b9e2a64cf51f400265cefd02614f05c9b421f57c3ae693530f2998

Observation b4451862-8e35-4abb-b5f5-be8b6e6cecb7 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Audio set: An ontology and human-labeled dataset for audio events,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.500522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.259716Z digest=sha256:f7632ff500880736209647a174d311254a8807d00a3cfdf81dee8d6a20d125bc

Observation 0d4e6590-ae9c-439f-93b0-dd4cd9c76806 · outbound

This paper cites The benefit of temporally-strong labels in audio event classification,.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation The benefit of temporally-strong labels in audio event classification,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T00:23:08.484410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T00:23:08.264537Z digest=sha256:7ee8897f8fcf5bc17b6aad887be6850d6ce4a9ce57a12497f44e4ad7bcf988cc

Observation cf39378f-b5b5-42de-aaf6-65dea21aaef0 · outbound

This paper cites Aist dance video database: Multi-genre, multi-dancer, and multi- camera database for dance information processing.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation Aist dance video database: Multi-genre, multi-dancer, and multi- camera database for dance information processing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.269384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.269384Z digest=sha256:a1945c49ef874918274875e2f7cccc3257ffe5e835fcaff06790dc1669fdbe2a

Observation e5033d48-6d2a-4146-939b-5c3e6f9f99bb · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.274133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.274133Z digest=sha256:54be7d8d3190f3138a2b8567428b3c61a985c651a0743ab12bc696d926f0289f

Observation e23d8b30-0744-4a48-9ac9-f57a8679632d · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T00:23:08.279261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:23:08.279261Z digest=sha256:3a4232d2f7244425ef2dbe4b6e7c59ef8b814f5243d7aaf2ee8acc5673cd8a2e

Pith citing papers

Observation a65d3b03-373a-4102-baa0-0f2224156e3b · inbound

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 cites this paper.

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:28.480478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:28.480478Z digest=sha256:77ddd2aa5a9a33b5238cbcb0baff82b8611295a4950565ba94da3da1c21ba6b5

Observation 4e3600b3-5b03-43bb-aa81-f95b258b9571 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.719488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.719488Z digest=sha256:27fede61100b7bf401b6fa5bb7a5b1f7aa960add6511669e59c7d3baccedca62

Observation fe916510-e314-4523-bf4a-49e286c8d953 · inbound

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction cites this paper.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.844041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.844041Z digest=sha256:1829b9b3e83467825128d953fb5fe27a6d98e8813b3c7ad697831669380efeed

Observation 86a4669c-fe76-4651-b85e-aec31796905f · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.911645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:6f3c323fc0f882b3d9a50d827077aefa0ac69f4781b48610fd8e288170137aef

Observation 750fb1d0-85b6-4f75-8733-d431af5e8148 · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T16:14:15.270070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:ef83b031f664522a00942e4c5b8eba2c74b6545745ff34fa3117463e674a9f6d

Observation 195a3271-7eee-405d-a737-f4f90e7b23f6 · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:44.178576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:28:14.734682Z digest=sha256:e9cf502228fb32b3de1b25ae09e073b11774e8d2cc2a0aaa495855f3c13ad808

Observation 36151287-89fb-4b24-b8e5-a1965a16360e · inbound

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation cites this paper.

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.816253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:26:46.077894Z digest=sha256:02234d308ae81f5bcb3405fc223d4f17a5eda4c75a00664897068ea752d524b5

Observation 5472edbe-229b-4705-bada-76c3b06531ac · inbound

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning cites this paper.

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.198346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T07:01:40.333448Z digest=sha256:7f73cdddc6d1de7f88160ea5ad64bcbef54483bef19cdb5b85a11a7d0cbcb6c4

Observation 432bf714-ac4c-41fc-8b02-6979188d94b2 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:18:03.069935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:17:11.690484Z digest=sha256:e9b1e89121bc6a5b424e9c8f38cbb94e6078866770520f8e630a6a24d9df7b48

Observation 38385add-b7cf-4a4e-ad03-983f4a143131 · inbound

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation cites this paper.

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.099423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:00:39.556424Z digest=sha256:19531fbf47e9d2b60e971f92f34df965b11c363751a3b04801412fa392cd88c6

Observation 8e276bec-38a2-4b47-87aa-98b4275b5842 · inbound

Inference-Time Scaling for Joint Audio-Video Generation cites this paper.

Inference-Time Scaling for Joint Audio-Video Generation UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:40.902116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:49:20.194990Z digest=sha256:32c6cd226d7c8c27cb4eccf2c3b919a6da44a7d0fceb8c8c92801a658bba2c74

Observation dd650d56-eda4-4d06-80ef-7bbfaa557f56 · inbound

Vorch-Omni: Multi-Task Orchestration of Sight and Sound cites this paper.

Vorch-Omni: Multi-Task Orchestration of Sight and Sound UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:16:52.548806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:16:52.548806Z digest=sha256:48d68b8656456f7427e7122c824d2e73e938198724716b05ef1acf9e84558332