Pith. sign in

Paper Citation Record · LEDGER

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

As of 8 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 1 inbound Pith citation observation for arXiv:2506.12573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12573 v3

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:47.983763Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:40.786176Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:49:50.392438Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved41
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:74a449ca212181acadc8a75d6e6b8ce7ec967ce34d3df49297fc464c54cd9301

Observation ad5fa3ec-dacf-498f-bb97-38acd2d4cd7b · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.864936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.864936Z digest=sha256:a9a026e445db5ef41067f84b6e4f36959bf7a90ab5016332523d4c96a71e9ef9

Observation 5a892216-9ee5-4236-be83-f39bfade8dd6 · outbound

This paper cites We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.950159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.950159Z digest=sha256:2e2e4f027ab8f8fd971ad411a6af068960de4bb2fb9ff5aa852584ba9c8ac29b

Observation 96cf67e9-1b5d-4273-8464-6b780233dcf4 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.045157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.045157Z digest=sha256:f672094aeb31b77250cf4aa0b2653a53d0cc28c6fe4e65c7fbe265251bdeb34a

Observation 2749241b-9bb3-4ded-8c1c-40b2dd661c27 · outbound

This paper cites S” and “M.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections S” and “M

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:49:50.232025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.120789Z digest=sha256:adcd6f9a9334a4ddb5a854aa37ecb930434e06de2c620db527eb57971937d84d

Observation 2ccb242a-283f-46be-99c7-458adbebe840 · outbound

This paper cites Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.198985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.198985Z digest=sha256:97ddb22fc225ed4b3fdb999c72c5357c1d2ae08edca6525e06fad9881feeb990

Observation 330770bb-000c-4f02-bfd3-07d4e36c92cf · outbound

This paper cites To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.630279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.272353Z digest=sha256:d58d79d8b366b275f27359979b1b3a5861b3217c0d4c89e00887bc6746081f51

Observation aa0b77f5-740f-4a30-94a9-5bef220bc6b6 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:49:57.502642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.339642Z digest=sha256:fabc8d0219ebbcf240f7e07ec50af5aee515c9a188efd690840d22a7b3acdeb1

Observation a2b075f8-ad4d-41ad-9f10-d69886d45ea0 · outbound

This paper cites Soundtrack design: The impact of music on visual attention and affective responses,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Soundtrack design: The impact of music on visual attention and affective responses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.424060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.417326Z digest=sha256:be0c2bf1707ff7a8001e17159dc0cd81e6cc735a8fd35009a2409b670e2aa4a5

Observation 931c243d-a0be-450f-8fe9-b6abb48a5c18 · outbound

This paper cites Multimodal deep models for predicting affective responses evoked by movies.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal deep models for predicting affective responses evoked by movies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.347541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.494318Z digest=sha256:b848c58272cf9ab2523add93ac2a614758e8b148ec387cf17f238c1c82de3880

Observation 65756565-9c31-4f5b-acc1-9927be466874 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.863527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.584055Z digest=sha256:31d3f6f3a8caa5d397bb858d047f898144f4ae5d09259a9777c2c47b4968d984

Observation 2d14e174-ec1a-42a6-8b9b-228347da9a6f · outbound

This paper cites Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.192060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.662550Z digest=sha256:b1d9cf3a3ca0090143901d9f2594d2c60c32fe3765f7039c1781a2db165418bd

Observation 24dc64c5-acca-4af7-92cb-6137c395ebaa · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.551873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.739826Z digest=sha256:3a3b72d0097aab701aad14e62703f8b5c0da7385867ff9571e2228f47edd0319

Observation 4b747be7-3c98-4c30-ba78-447e0aa5feb9 · outbound

This paper cites Analysis of the roles of film soundtracks in films,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Analysis of the roles of film soundtracks in films,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.820319Z digest=sha256:73a12103e3a8b6dedf363a552d77c602cf4ec37d9a453a7607dd946bd8445f67

Observation 7f35f1ab-b34e-4616-b036-c2cfcc4bb411 · outbound

This paper cites Actions in context,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Actions in context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.884556Z digest=sha256:3a754b5796793668c50eeed8e6d86b6311af635bc7a82e430f5e16373dee2fa3

Observation 57e6cdd2-bd3f-4ded-a1b7-431a494383cb · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movieqa: Understanding stories in movies through question-answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.692466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:41.963832Z digest=sha256:590ad773f77886805c8494472318d16482cb8f1d87101bd1d18a8fe65455a4dd

Observation 178f64c9-084c-4961-8dd7-9428a98bcc0f · outbound

This paper cites Movienet: A holistic dataset for movie understand- ing,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movienet: A holistic dataset for movie understand- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.534391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.035597Z digest=sha256:4e71da4f8754897310ab36f576882a1a78f4245f050733ede50160917f95a918

Observation 8255ea0f-8c7d-4a2e-a083-578fb835f1e5 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mad: A scalable dataset for language grounding in videos from movie audio descriptions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.367972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.103527Z digest=sha256:de88ca0eeb775922b36ab9489a1cce7cc9853566fc2c9baa7e74c9defb5e2765

Observation 9f955c11-efdf-4d39-8a82-67e5f90dbc62 · outbound

This paper cites A dataset for movie description,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A dataset for movie description,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.278451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.197524Z digest=sha256:1ced6f57c5347b1162a5d3005e48b80e0271a4bbfdce6f128290f6bb27347937

Observation f61528fb-e050-4ddf-bbf7-9f2c67e5fc0b · outbound

This paper cites Moviegraphs: Towards understanding human-centric situations from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Moviegraphs: Towards understanding human-centric situations from videos,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.112885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.279348Z digest=sha256:70d60f17ab027b418b42ed5708970bcf97a0eb277054939f5d36788dbc63d061

Observation 9086561a-5ce8-448d-84b3-30760610658f · outbound

This paper cites Hlvu: A new challenge to test deep understanding of movies the way humans do,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Hlvu: A new challenge to test deep understanding of movies the way humans do,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.021400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.366444Z digest=sha256:8b88fffa1fa8406e9b94c12c39331764c3018b5d58170d822b0145fecc7ddd43

Observation 3237553e-2891-420a-abb0-136945d11bf6 · outbound

This paper cites Condensed movies: Story based retrieval with contex- tual embeddings,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Condensed movies: Story based retrieval with contex- tual embeddings,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.920877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.445904Z digest=sha256:22c0f04379fd45c447c5a71f62ea89ec2bdbf815c8fab3f53bdc41e8edd36027

Observation 4799cefd-2193-40fc-a7d3-0a70dd655e0d · outbound

This paper cites TeaserGen: Generating Teasers for Long Documentaries.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections TeaserGen: Generating Teasers for Long Documentaries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.511515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.511515Z digest=sha256:93f502e0e2420e85016f3e2a228d2b1fba4ec5eb6695ba0097d192b8b4dcc05c

Observation a9178205-829d-4963-b251-0bb2ad2877d2 · outbound

This paper cites Simple and control- lable music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Simple and control- lable music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.764831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.599104Z digest=sha256:792bdb6e71d1ba470ff093f89f35a9e47b09ca51cef4b81d469026e47f33d5de

Observation f1a5fa92-d749-4ce4-9143-a20f91670f51 · outbound

This paper cites Attention is all you need,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attention is all you need,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.675447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.675447Z digest=sha256:ba5fe8ea2b7929ee1d1f8ec741dc409e18d498b135b89598dbad390e0aa7e13f

Observation f47ca879-4b18-487b-9530-1e89f691e266 · outbound

This paper cites MusicLM: Generating Music From Text.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusicLM: Generating Music From Text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.750717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.750717Z digest=sha256:abe6de9b2e418634b7c545546fb0068f54a68496f639e0da8f837e74029d18d9

Observation 744cb494-a203-4312-a521-fafac7c164da · outbound

This paper cites MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.827934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.827934Z digest=sha256:5f5520cebe89161dc27730daf86c514a81543612b20d393f0c69ba06b5f1cb1e

Observation 1450699d-fda3-4db1-be6d-12e8ad47a34d · outbound

This paper cites VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.900461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.900461Z digest=sha256:e79da46f6979852be21a0665f9563fd074badf6eb0e07428caa48b7474d0a64c

Observation f33b4d9b-4689-465d-bb03-a5353cfe780b · outbound

This paper cites V2meow: Meowing to the visual beat via video-to- music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections V2meow: Meowing to the visual beat via video-to- music generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.665078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:42.977529Z digest=sha256:1997634a8d8cbce1ec08c9be2ab4018b4c2fc816d4f7c7a8efa53b757d6bae61

Observation ab03e73b-0eb7-424e-86c1-3955da2955bd · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.247373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:43.064927Z digest=sha256:7a8c4ced4ba40d209b039e21be9434c26552a50bcb9b2d7475eed5c3859016c0

Observation 7f8710a4-62bd-4f3e-ab57-14e14f6e0164 · outbound

This paper cites Riffusion-stable diffusion for real-time music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Riffusion-stable diffusion for real-time music generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.456481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:43.156770Z digest=sha256:8de8208c2f2d454c70b944e2b48bc4bd3082526175144bab2624499ed1746dad

Observation 92b97fe1-1c0a-456e-a812-f804e48e40f9 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.230631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.230631Z digest=sha256:1e82bdcdbec6893741f863442c30cd201ba3230ccfb69547585edb07c74905ef

Observation ff44a7fe-ccb8-437f-a06d-699d6dba7156 · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.294050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.294050Z digest=sha256:7154824862e5d27c1948e23ad86c92533b370d9db412b624811c8e3023b3bcf0

Observation 41e83f31-e1a9-4b62-9f2b-e0d367cf7fda · outbound

This paper cites Efficient neu- ral music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient neu- ral music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.286800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:43.347990Z digest=sha256:bc10e8424bf294b946a0c0c8b19470a6ce2659a2ed26263469021362778fba86

Observation 712b6267-1f37-4ea1-94e7-9cfb31d71bfd · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mustango: Toward Controllable Text-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.437966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.437966Z digest=sha256:0ebfd56090d3ed788ac71b606a803ebdb9e8065d2c8d128fa401e1fec94de738

Observation c168ce5b-df36-4fc9-a7bb-213c000ebc30 · outbound

This paper cites Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:48.971726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:43.507489Z digest=sha256:ac4be00d23d94620ce2076b7a9a0b661e44bf128d505e90fa59a1a5c5be398f1

Observation 7e3009df-3e6b-43a6-a441-32a41699f533 · outbound

This paper cites Stable Audio Open.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Stable Audio Open

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.573592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.573592Z digest=sha256:1c02f66dc846c65bce575e90a38be4c805fa1f55a70074f046934e1edbb557a7

Observation d9d5c618-8e40-4ceb-a172-1711575d2e26 · outbound

This paper cites VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.639290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.639290Z digest=sha256:ddf50c18feba99d5d3a4e14d8a246cdf38059921aeb63e44445a110c09caac9e

Observation 92882141-7d2d-423f-9e38-e178bfac7bf1 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.716277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.716277Z digest=sha256:729947d7c4015aba1b29b0f9b48d990db8c1e4328504004978e7f3cd6e7c509b

Observation 9b963403-04a1-475c-bfe0-7f4522db4f95 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.791478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.791478Z digest=sha256:0bff9b450e940689c8bc7116f5c76c1416dd028eb4dd1e9a3c8031e44b538c96

Observation 7611343f-ec97-4eca-803e-ef17b1d82238 · outbound

This paper cites Music controlnet: Multiple time-varying controls for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Music controlnet: Multiple time-varying controls for music generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.092002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:43.880967Z digest=sha256:4287aacb72e91aea72d4567a1c593d4b38ae4e778146ea1ccef10a1e62730350

Observation 2e6b4824-b7d1-43fb-a24f-ddec9c5092af · outbound

This paper cites DITTO: Diffusion inference-time t- optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO: Diffusion inference-time t- optimization for music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.987129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:43.945081Z digest=sha256:b8cb663b3897403205e7135cc77db25ee2467134b10c72ebcb639b8f7f7d9ad4

Observation 6d17aff5-b6c0-460a-bd6a-44a3e4baf800 · outbound

This paper cites DITTO-2: Distilled diffusion inference-time t-optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO-2: Distilled diffusion inference-time t-optimization for music generation,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.851023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:44.026351Z digest=sha256:d704184e583f3b52ea38715ae1ed91e600caff62bcd1340c2474f3800c1fe931

Observation 29b928b4-861d-4832-8757-7750994f7a18 · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.099128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.099128Z digest=sha256:9f556a69a8e7985afc80723c9a18d03957429e9fd12d74c20ff0070a46f5fd1f

Observation b0d5dc2d-5a3c-431a-b6cc-59798dfe8b86 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.175395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.175395Z digest=sha256:5e20b662881096f89201626ca641b75bcef108bef8dcdb5746e06ecc3f816612

Observation 6ff4f4fd-722b-4618-8a9d-efd043432fe7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Foley music: Learning to generate music from videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.708344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:44.227354Z digest=sha256:e076de1edb39b7ad5fd8a0f4a6eba4c5eb090dddf74ea25a457406d796703a1e

Observation 333a7a68-750d-4f41-8e71-6270ce04784c · outbound

This paper cites Video background music generation with controllable music transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video background music generation with controllable music transformer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.292165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.292165Z digest=sha256:0d3b8889ce03a759b89c80fe121b75879213888d5e5ee0019028f1b1e9ccca0c

Observation 34a5a28c-d19a-423c-8526-5ea63f78a210 · outbound

This paper cites Vivit: A video vision transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Vivit: A video vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.286387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:44.438782Z digest=sha256:b27c2e247279e921b6198065e9ebd778861a546dfaffeb4dc7a3160248f3ea54

Observation e1f7b1d7-f835-4570-92d1-e247baee4a2a · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Parameter-efficient transfer learning for nlp,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.522171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.522171Z digest=sha256:823f01f7a828e8219eb3697683b7a8630629bb0937dc272d48dcc7f8ff403ba9

Observation 7e28f1b3-58b5-4aa7-a2d5-331bf2c35972 · outbound

This paper cites AdapterFusion: Non-Destructive Task Composition for Transfer Learning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterFusion: Non-Destructive Task Composition for Transfer Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.596943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.596943Z digest=sha256:12a3fa76ff7f418cd516463b06f2f481053afcb7b68ad8101a6023e08f7dc12a

Observation 98494540-e731-4a61-96b1-884d338d1334 · outbound

This paper cites AdapterHub: A Framework for Adapting Transformers.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterHub: A Framework for Adapting Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.690125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.690125Z digest=sha256:b0e816f4adbcc0088b08566a455e1951639ee71f834e8aa5edab8cccfd10c5f0

Observation 32c0adc6-f711-46ee-acfb-e46541056b6c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.750532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.750532Z digest=sha256:8e7d955e89e73e9c28415777fbdc33ec450d894efd2e5d3250a5dfb1987080b5

Observation 72702529-6ad2-48f3-9b62-3d1570f1714b · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.107266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:44.833785Z digest=sha256:115c0db33a9810a100c209a85835010c83f35299a7bb5e922c63c93eb8ee63d6

Observation ce65a2d3-06f7-4f66-9924-d8574cc2a9e5 · outbound

This paper cites Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.938203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.938203Z digest=sha256:23eeaaeb973d25ca1ad835bf8dbc031b2f458da165c6f4bbce57523c58ac997d

Observation c0c004cf-0283-4c5d-9f20-94aaf665c911 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.008013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.008013Z digest=sha256:3b7b91f080c9523e232242651086d4379b67c186d381c5874c7148d0a4b98b0f

Observation 62e97fff-57f7-4bc8-8b5a-1aa173c5c3ae · outbound

This paper cites Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.121557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.121557Z digest=sha256:8ff821de2384118ef3e1f76330329a0007956d9a58d019a06a9659bf44c1818f

Observation cb5d3735-9a17-41c5-9b44-bd2b4875388b · outbound

This paper cites Quantized GAN for Complex Music Generation from Dance Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Quantized GAN for Complex Music Generation from Dance Videos

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.246773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.246773Z digest=sha256:ec12709a3217339e89f718e0c70a8b7e022fa564cda2b0fd8be7c3b94e1046b3

Observation 50e3b075-93b3-4d87-b3f2-18ffe2df324f · outbound

This paper cites AI Choreographer: Music Conditioned 3D Dance Generation with AIST++.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.341616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.341616Z digest=sha256:57a7cd8bc7c608c34898fb2d7a1196645a7eb535c1f09ba8d984cd40d36873e0

Observation e25abf34-5f26-4e9a-8f49-7743e2818e36 · outbound

This paper cites Video back- ground music generation: Dataset, method and evalu- ation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video back- ground music generation: Dataset, method and evalu- ation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.547185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:45.429434Z digest=sha256:cfe5c510db2e4e59ddb8b30ae1d0d8e3e59bf6493a9367d5d64006cfcb7dd555

Observation e3261b1c-afce-44ca-b1f9-9cf8749ed34e · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.721432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:45.523930Z digest=sha256:16047bf05ee385dc80cc891b0bbdb9688730c9cd02439c7a597b0b3f5e9f7800

Observation 5c29d675-b417-483e-b7db-5049e09cea7e · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The Kinetics Human Action Video Dataset

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.834553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.834553Z digest=sha256:6577c4f97ba21773f27214ada2f7584ebf2e8a93684a48ec11a3e9bbfe7d367d

Observation b536ca68-8cc5-4f46-b632-2d4f528d1b24 · outbound

This paper cites Diff- bgm: A diffusion model for video background music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff- bgm: A diffusion model for video background music generation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.453381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:45.792613Z digest=sha256:60ca3f7147794a565dd497efa3c504861e3b42164571354a5e002ebee413dc05

Observation ed757164-55b8-47f6-a5e6-548e964111ed · outbound

This paper cites The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.215919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:45.876102Z digest=sha256:21c2c0739d38e18e018dd3404ade9e3d9bb2a4d0ef5457522e67466476afcf3e

Observation 4a9c785d-d9ee-4467-81df-a3c8c7f3d811 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 65

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T00:49:48.478189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:45.996145Z digest=sha256:b1188918db220a9991abaedb4c953664b9115c495786dad5e756c43a80dee4cd

Observation e34b0e31-a537-43d3-a04c-e3180ff025b9 · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Benchmarks and leaderboards for sound demixing tasks,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.959182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:46.099559Z digest=sha256:48786691084e2516e72f648c6b15bc74f0179f6cb70476db277ec529e28f0097

Observation a7b7110e-f6db-48df-8c8f-0f298f8ce98e · outbound

This paper cites pyaudioanalysis: An open-source python library for audio signal analysis,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections pyaudioanalysis: An open-source python library for audio signal analysis,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.616912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:46.248153Z digest=sha256:b9814cbbfcffffbb2cbe5a8ebba2ac6de35fac6d4802003a71417b30dbce2a8b

Observation 9dfd5c2b-5a9e-4eef-a362-a5f81764df87 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.378501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.378501Z digest=sha256:873ab4450cf40615ff119205651487a3f7438612d4c9dea0094004a7d743eaf3

Observation 50cc1ea6-278d-4b79-96b4-68bbdbe10c97 · outbound

This paper cites A circumplex model of affect.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A circumplex model of affect

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.500264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.500264Z digest=sha256:0a9c209a0b546718720e14b0aa197634d4a7ea8fd838cf820ab400548c9825ea

Observation 44613096-827c-4252-953e-7ace2a26ea60 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.604276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.604276Z digest=sha256:34e0f09a5a4222da0dd9344a62502a62aad237de2888046e6c69d19f03179704

Observation 7fc4e7aa-0042-4bf6-bfa3-f9f85ef5d283 · outbound

This paper cites High Fidelity Neural Audio Compression.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections High Fidelity Neural Audio Compression

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.732431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.732431Z digest=sha256:a76d7c0e92dbe30036569e0fc2e3083871985960cbb190f02b2690537080729b

Observation 00b7d1d0-bfaa-4804-b2eb-2456a7ab899f · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.161900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:47.758112Z digest=sha256:864453d376bfe1f36a4c6bccccffc390c37d93e1e3fe5ddbbf909e73b7c2742f

Observation 4f9aa396-4221-48f1-89c8-704c1015dd39 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LoRA: Low-Rank Adaptation of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.967955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.967955Z digest=sha256:02c9262f257f1ed41df1c41f9977eb8efd066c2e1680eec857fc0de3e9971e1e

Observation 0ebc6fab-03ae-4375-a19c-4f2db7f4fc0e · outbound

This paper cites Language models are few-shot learners,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Language models are few-shot learners,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.186777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:47.082948Z digest=sha256:9d8408b2561e7d4d7059d21065ee2edbbf647c14fd2acec360d8d4f987c8252a

Observation 69f6d0cd-4b1a-421a-8448-b2f4c0123638 · outbound

This paper cites Design guidelines for prompt engineering text-to-image generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Design guidelines for prompt engineering text-to-image generative models,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.823139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:47.186083Z digest=sha256:acb97b3fd2ffd9f038fe730d065e8e7a95baae04277c348ad7cf1c4e4cc0b74b

Observation 6e8f562d-5af2-4bc7-a4b9-8d521f795dd6 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.311232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.311232Z digest=sha256:b7dc563e88599e98b10352c5bb2a02b76a11a67ec249f08365136461cd27f9da

Observation 18e95974-ec41-4849-a687-64572867160f · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.415707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.415707Z digest=sha256:d710ec04a0bbe673c5428dc186980125437abb8a5071f1fc64896cab9d73f873

Observation 81eda005-80ec-47dd-8f4c-298358cfd8da · outbound

This paper cites Decoupled Weight Decay Regularization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Decoupled Weight Decay Regularization

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.484231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.484231Z digest=sha256:a2ef6125d4e6188b3dc27d371537e95c1a4973663b79c79a73fdc32e61c52c89

Observation 5a4deb80-728e-4503-99cc-5da2dccca4ba · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.550534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.550534Z digest=sha256:fdad931cf790ebe89103debed21d4b123ea591bbff340ede9fd5aebedab5600f

Observation 50ab875a-1595-4b31-bd65-6e8a2f65f51d · outbound

This paper cites Presto! distilling steps and layers for accelerating music generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Presto! distilling steps and layers for accelerating music generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.477011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:47.631070Z digest=sha256:8aef8c3efd5cde1e2101fffef468f1d1c8fc623658e0360285a7561d925fc252

Observation 17565a16-a129-4653-ad4d-251f7255020c · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.697740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.697740Z digest=sha256:315facfec33cbd9ddeb96eca2aaa28c38d3fc678b1675a902c4c18d8ba1572bf

Observation 7bc0f0b4-575b-4092-9629-2b5e5b319de3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.832625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.832625Z digest=sha256:b44ac13cb91b3cbac64a9d785ac586633cae28c194a12c1bc0de6d3064297766

Observation 5851456f-0269-48f1-ae3c-d32dc18108bb · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Reliable fidelity and diversity metrics for generative models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:50.882373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:47.896765Z digest=sha256:77294b5256844b198243cb1eac88cd2cc698bd54b3c306e1a808ffa30a64d4e2

Observation cc2e518c-34dc-4078-9479-835127505348 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient Training of Audio Transformers with Patchout

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.983763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.983763Z digest=sha256:29d8b217605d76b3d773a727bd7841345ecac166ef7c9ba03000cdab830c9bc5

Observation 91732895-4c5d-4970-93e6-4f4a5c4fe0e1 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.673911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.673911Z digest=sha256:7c32960eef710a4022211f8a18688c5a9548784acc76f2acb4fc21a0641437c9

Pith citing papers

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:74a449ca212181acadc8a75d6e6b8ce7ec967ce34d3df49297fc464c54cd9301