Pith. sign in

Paper Citation Record · LEDGER

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

As of 12 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2412.10768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10768 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:43:01.110656Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:38:08.339998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T20:27:03.913459Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd153951-d89a-4a29-895b-f5070adee5ff · outbound

This paper cites The Foley grail: The art of perform- ing sound for film, games, and animation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation The Foley grail: The art of perform- ing sound for film, games, and animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:02.018261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.831775Z digest=sha256:f1178e2f644c7abd9993427a4b312e26af85802d7c869c3078fee10643841dac

Observation cc69f13e-4ca3-4858-a248-4a172f5a6604 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation In- structpix2pix: Learning to follow image editing instructions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:02.003279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.837425Z digest=sha256:56629e8c510998ce989e5bbfd850508ebbaaece4be605857c3e2fa4dcb7be9cd

Observation 394a27a7-aabb-408b-ba07-37715ce35c5e · outbound

This paper cites Video generation models as world simulators.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Video generation models as world simulators

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.842354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.842354Z digest=sha256:d2bfffb8e76850e9d3fba2b508549995d10f13a54b5c85222bad08144c3b4bb4

Observation 0963105e-5194-4970-b34c-0f821db9a0f4 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Vggsound: A large-scale audio-visual dataset

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.978661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.847396Z digest=sha256:e96d21c446a54a1825524e54925b7fb1842d6262fe427df400c897c200f7795d

Observation 1f456e97-c8d9-4bed-a688-96db9796cdd3 · outbound

This paper cites Gentron: Diffusion trans- formers for image and video generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Gentron: Diffusion trans- formers for image and video generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.963716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.852162Z digest=sha256:2bffc68b01ddee2aebd8126b7cdff5f5bfde74ae967c400159b4b5b35b35367a

Observation 7504d809-a4a5-4610-a200-d4f02bb44391 · outbound

This paper cites Audio-vision: sound on screen.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audio-vision: sound on screen

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.948849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.856979Z digest=sha256:a2d3f2110f207f2694c25335586e8bae1247176178e578c3de1ff6024180641f

Observation 42357cef-9a0e-4976-b4d7-37d60620d868 · outbound

This paper cites Scaling instruction- finetuned language models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scaling instruction- finetuned language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.933714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.861754Z digest=sha256:ce20a01dab17c91e15ea6537c91e20e5dab2937af2fa678696cfe2bb8bbcd8df

Observation 70a28615-c939-412f-addc-4ce07851924a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.867260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.867260Z digest=sha256:cdba7f73e34df825655fe6c187b4edc3304f74c2879015fb5d0c2c66eb97a306

Observation b5e1b2c8-5007-4e4c-a5a6-da577476013d · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.877098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.877098Z digest=sha256:af532ccb37b21cfbdb4fe1eccc0f51945d8a75475c5db8adfff3bb3aa9d786ca

Observation f9bd243d-5c75-4739-afa0-c72c3a3200c4 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.881954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.881954Z digest=sha256:bde10b4a58e39f5ad783c183bd7294ae084e0a4f8273cf0456cb3859119d5d56

Observation b0a8ec89-425b-4bb6-824d-e9428d88aac6 · outbound

This paper cites Imagebind: One embedding space to bind them all.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Imagebind: One embedding space to bind them all

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.917393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.887089Z digest=sha256:6ca59ecf0faf68d4eed21135ea9ff30fb4119eb0bc39b24aad45baaaa4e34e86

Observation fe58a50c-5079-4ba4-9d18-971491a11c80 · outbound

This paper cites Determining op- tical flow.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Determining op- tical flow

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.902604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.891837Z digest=sha256:65f85e10a33c52002e79893540f15393e2255dc9752792899de363b0ba9a76db

Observation 81536570-f0b8-4950-ab39-2777285f32f3 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.887430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.896690Z digest=sha256:5ee720dbcce913767239caa6899393bd87ca4d0789899d014dc59f577198fc6b

Observation 3a79237d-0f2d-4eb2-b0a3-fc343e6218b1 · outbound

This paper cites Captivating sound.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Captivating sound

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.871039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.901474Z digest=sha256:aa568039b239c2b85d4218140b90e22b1b4604d109fadd6ea8dbd7ff5d13000f

Observation 43da5ea3-8cb9-4f97-b8a6-e16b0aee59ba · outbound

This paper cites Taming visually guided sound generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Taming visually guided sound generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.855536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.906290Z digest=sha256:8a24d5bded9653e47d0c1603e4ffca90e2d57c8443b5112badd6d5740f1d0f90

Observation 7bbc4dc9-61b1-4ad3-812f-df8ad7362d2c · outbound

This paper cites Mixing audio: concepts, practices, and tools.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Mixing audio: concepts, practices, and tools

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.839857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.911140Z digest=sha256:67dfd81189a9ff689832063960b0fdd2853a98b5ba363d836ad7961bdcb47d0a

Observation 2432f03f-54a6-4d14-abc8-427a56b07f0c · outbound

This paper cites Read, watch and scream! sound generation from text and video.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Read, watch and scream! sound generation from text and video

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.825042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.915875Z digest=sha256:934b309b70ca679d8950b720e6d8a29c9cea587d074f06abe42560fa99f51829

Observation f01f30cf-82ef-45fb-bb91-d7747fc80aba · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.920633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.920633Z digest=sha256:8e38ee57afa0c4932dc316dcdc2dea04ba17db64c9ff75a35c9071765e4a7afa

Observation 3133d9e4-9647-48bd-bcd2-f71ab0d354f0 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Hifi-gan: Generative adversarial networks for efficient and high fi- delity speech synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.810087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.925539Z digest=sha256:71bcea49045387bb9f0389da95a4bc8a71068d66c774136417e7bbe460b6bf19

Observation 75b0328f-621c-491d-a3ed-870cadf9102b · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation AudioGen: Textually Guided Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.930103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.930103Z digest=sha256:fec86e206a27c5d9669d4847bfdf21b98f3f32c2223800412fc843983d66f80d

Observation 38d1f91e-193d-48cb-a0d1-73980d2d4881 · outbound

This paper cites Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:43:01.305284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.935561Z digest=sha256:0caf89d2cc96e4d67c01e6ee16d05f19b75ec03c921eeb4ec95a5463f1e1a719

Observation b8b58e80-68ad-478f-a8f0-d83a2c93cfa3 · outbound

This paper cites V oice- box: Text-guided multilingual universal speech generation at scale, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V oice- box: Text-guided multilingual universal speech generation at scale, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.793935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.940625Z digest=sha256:7f94f12c31edacca2a2886f62984f8e9e30f7959f5530a09532e2a82601af5a8

Observation 67ef12d0-af5a-4b55-9d23-416a97f5b652 · outbound

This paper cites Flow Matching for Generative Modeling.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Matching for Generative Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.945415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.945415Z digest=sha256:c7bdc7d026c10323ff910fe9f49454dfb6b2b2089cc85fff40a8f636e03c651e

Observation bddc6b1e-a04d-4b05-a33e-ea5dabcbdd7d · outbound

This paper cites Audi- oLDM: Text-to-audio generation with latent diffusion mod- els.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audi- oLDM: Text-to-audio generation with latent diffusion mod- els

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.776837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.950472Z digest=sha256:0f41b8348ec3ec30267b5a1a038daec049c0210987a6a1756329028f322ad4a0

Observation 78e9addd-a28d-476c-9917-4230865d43b7 · outbound

This paper cites FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.955420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.955420Z digest=sha256:fa54a77fc9e6bc1ae9120f01e94442779a007302266241d63c608a1d878ea60d

Observation a8892f56-768c-46bc-bfca-c8d4807d0c24 · outbound

This paper cites Plumbley.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Plumbley

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.759012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.960680Z digest=sha256:3edaf01b93acd7e9a15bf07312be8595c1e730335891c6750d794880badda1ad

Observation 8fcadfb4-164f-4c07-b2e9-6d9c8ec496b8 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.965632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.965632Z digest=sha256:109480406ee92e5de406751c5645cc68e0a2cfa141fc08fa65c996b1e65f6f3b

Observation 78d0d2b7-7e5e-410d-935a-c9cac5453cd8 · outbound

This paper cites Separate Anything You Describe.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Separate Anything You Describe

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.970707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.970707Z digest=sha256:8b5fe9e011f7612e1d1ed9c02c7d91d9b1730e3a1806c52bd7a987f19c5e5d82

Observation 8d314d68-d31e-46f1-b28a-60ad9c7dd7e9 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diff-foley: Synchronized video-to-audio synthesis with la- tent diffusion models, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.743876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.975676Z digest=sha256:0a6aee3fa057208e1834013d45c9fda3cf2a7f7bd2c343e0565ea70066c86cf0

Observation 19bb9429-a598-4d77-8b51-38262cfc3520 · outbound

This paper cites Albergo, Nicholas M.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Albergo, Nicholas M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.727746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.980418Z digest=sha256:99c6a810ba33c9ffab14aa97833b99b6c4d4f62374b01f4dc9843f7842454ead

Observation 3c02fe20-b9e2-4a7f-9b7a-4668f091d418 · outbound

This paper cites Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.710578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:00.985150Z digest=sha256:9f07593f2d954a238d9a5ac07e61f4d98f93e0e048dc503301f5d993d15a628b

Observation d496671b-d7de-463a-87d5-87e22465082e · outbound

This paper cites SampleRNN: An Unconditional End-to-End Neural Audio Generation Model.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation SampleRNN: An Unconditional End-to-End Neural Audio Generation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.989949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.989949Z digest=sha256:dcd69879948a9ee332f1c145fc70f0a2488b06be52fbd789cda0c3ab1f19869b

Observation b6b1c667-3311-4ac6-be88-5c8557f19490 · outbound

This paper cites Text-to-Audio Generation Synchronized with Videos.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Text-to-Audio Generation Synchronized with Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:00.994918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:00.994918Z digest=sha256:2fef07c91739005ce63fbb1ecd3d8a4c63e09723a89dbf72e54545562a774a2a

Observation c193d648-be82-4cc0-9ed6-b7975fab2e0d · outbound

This paper cites Bal- ancing act: Distribution-guided debiasing in diffusion mod- els.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Bal- ancing act: Distribution-guided debiasing in diffusion mod- els

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.695014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.000235Z digest=sha256:ceecfb20adaa6ed2a30d6d2ec4d9621477f938342e5b7085b15360d247fdec7f

Observation 060323ee-af08-4265-8a17-867538b01346 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Scalable diffusion models with transformers, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.679620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.004938Z digest=sha256:bc122c08319f54d3d7744e9c5e2597b24216764ccc3d13459852f2df24cd891e

Observation 341338be-00a1-4f5a-9cce-1df26449df12 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Film: Visual reasoning with a general conditioning layer

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.663487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.009584Z digest=sha256:f2d7d35fb3e51fa8cc69b74cdc8f0452df4da4470f4cfb2cc2297eba9fe19844

Observation 91a4fadc-1435-4907-8ed2-f3cc48242973 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.647834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.014105Z digest=sha256:9e9e5cc85516ee8f4b5038afcb0283a734204a83458ae1aa1f2112e58d13869e

Observation 3b585338-d1ed-4056-9f83-e666efe31037 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.018952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.018952Z digest=sha256:66f967138b742afee6321daa58d94da144bd8f3b6cb45b44360b989bdf473848

Observation 3fe98423-9360-4269-92b1-00befebb6bbe · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.024145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.024145Z digest=sha256:8c347572ab7909f892d2ea0008bc9dc0656bec4394ee3b402f739a8cb702d83a

Observation 757ba374-dd75-4de7-a979-d360d2eb4d65 · outbound

This paper cites I hear your true colors: Image guided audio generation, 2022.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation I hear your true colors: Image guided audio generation, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.608310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.029283Z digest=sha256:d917f5a6f74eaaefad94aa7c0f97a7878cc777e5c4cec5090291f8685afd835f

Observation 42c483e9-14f8-407e-bf7b-ef23dc79e883 · outbound

This paper cites Auto- acd: A large-scale dataset for audio-language representation learning.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Auto- acd: A large-scale dataset for audio-language representation learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.591925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.034089Z digest=sha256:591766349853dc7d278287a722d2a35e9ddf43d4a61b4ade65000fc0e9a1eb97

Observation b5e2a8e2-082f-4ce4-b38f-11c8bb590245 · outbound

This paper cites Learning from Between-class Examples for Deep Sound Recognition.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Learning from Between-class Examples for Deep Sound Recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.039003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.039003Z digest=sha256:365d4c048ee873a66b2be9c21f895c3e5db7a07935491414c02ae736538ccf22

Observation 9159ca6c-bf3d-4acf-b2aa-ae3f5fb14cb0 · outbound

This paper cites Attention is all you need.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Attention is all you need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.044268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.044268Z digest=sha256:6ba5f882e31e5810bab5a5125664aa3f29be7a9f5e0ac00bcbbc1cdfff52f9ad

Observation dff84b15-706b-4178-a31c-c2e78cff4f93 · outbound

This paper cites Audiobox: Unified audio generation with natural language prompts, 2023.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Audiobox: Unified audio generation with natural language prompts, 2023

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.562666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.049825Z digest=sha256:8c50c17e1551d803e38eab21dcdcda8a5de0e29c908a3b056bc008bacfc10506

Observation 44a95349-babf-4205-b7ce-e6d28e03dc31 · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foun- dation models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.545952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.054754Z digest=sha256:0618312acc8f3332e2cd7a04370ae5aae9ffd41c8e0276d07ed41c8b6557b0bd

Observation 5a5265e9-f206-4211-991e-1848bed6af07 · outbound

This paper cites ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.059533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.059533Z digest=sha256:7d009ef5f3ce37000204b15f0ea4b00be25bf34f22b0d78e0981a152e2c657ac

Observation 559861e4-be63-48f3-a037-448bc3afc367 · outbound

This paper cites Wav2clip: Learning robust audio repre- sentations from clip.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Wav2clip: Learning robust audio repre- sentations from clip

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.529275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.065138Z digest=sha256:f898b09c66b64f1f23184ae6924550aa98910a0281f9c17bb9066b410570d8ae

Observation e8c49149-27bc-4322-8696-04682191ead8 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.512058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.069941Z digest=sha256:46d9bb21715d66db25774c35fce89277d087efdb741c814c14d7d0de489ce5ea

Observation d7192d4f-08a4-4c10-b83c-894b0be99496 · outbound

This paper cites Son- icvisionlm: Playing sound with vision language models.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Son- icvisionlm: Playing sound with vision language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.494889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.074951Z digest=sha256:8ee575db4763b95afdc9c6811fdf87222fa15f6c43e1f5b5fbaa7613bc78c209

Observation 436724e7-41b5-4930-a16c-78d0a5a465e6 · outbound

This paper cites Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Seeing and hearing: Open-domain visual- audio generation with diffusion latent aligners

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.079975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.079975Z digest=sha256:37f95d906393fbd651f395403f3e31d262feb36d0d0d50ac9cfe4637c5c25dd8

Observation 6b8802d4-dd19-427d-b44a-c1a00998902c · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diffsound: Discrete diffusion model for text-to-sound generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.466318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.085034Z digest=sha256:b33aa6104b5d95253870ea80a0c24171e5d1a1ad4a9a1f035e37bc7ccdb266dc

Observation 1403e556-2132-4faa-b35f-1318ecaba71c · outbound

This paper cites Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Diverse and aligned audio-to-video genera- tion via text-to-video model adaptation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.450177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.090156Z digest=sha256:da5e35b378cd1769cb3d8cb89e2ff63caf31b42c5ddeb5bc6ba0162a648d17a2

Observation 079eb143-1244-4a61-9e47-b3b193f10348 · outbound

This paper cites Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Foley- crafter: Bring silent videos to life with lifelike and synchro- nized sounds, 2024

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.434140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.095147Z digest=sha256:4d976f223a06574e2a60a9690dad7ebb4b54e2e5a900eccb656e960be90b782b

Observation 8bafe478-14c6-4828-83f8-99f37de51dc7 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Llava- next: A strong zero-shot video understanding model, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:43:01.417921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.099877Z digest=sha256:8fbadf2c8ac3ae9c0bab4e4e63833d6b7f1f2a12aa85b7e7b9666ba6c0f5c1a1

Observation c84ec568-408e-435f-8846-805b45ad41f6 · outbound

This paper cites an unresolved cited work.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:43:01.402261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:43:01.104419Z digest=sha256:a9e7a2a82b06a5390dc3826621560101311afc250bc5b227974070dea675f325

Observation 10fc6dca-0337-4a2c-b5dd-4542c2b61058 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:01.110656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:01.110656Z digest=sha256:676c20bb8ab7557f57765d332d4463acf4ba17bdf721dd76c3245fd29ec183ec

Pith citing papers

Observation 52d52cb8-48a5-4be9-989e-07211047f4bf · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.339998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.339998Z digest=sha256:cc67ec4ee6c226862eb5dcce3ae161ac6f0667a98f5811e3a71b5c73a1f0bab7

Observation eaa0d4c0-e44a-4942-bc74-9a9d04cb9093 · inbound

Sound Scene Synthesis at the DCASE 2024 Challenge cites this paper.

Sound Scene Synthesis at the DCASE 2024 Challenge VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:27:03.921470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:27:03.773063Z digest=sha256:1107203882829bbdbf74668bbba2ae6d069370cabf776d2954baf81b3d409354