Pith. sign in

Paper Citation Record · LEDGER

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

As of 17 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 2 inbound Pith citation observations for arXiv:2506.12573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12573 v3

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:49:47.983763Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.500746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:49:50.392438Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved41
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:26f43669896461dd9b32b78967915b223cff03053f0581ef32f63a554c051709

Observation ad5fa3ec-dacf-498f-bb97-38acd2d4cd7b · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.864936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.864936Z digest=sha256:99cf2f22c0dc6084d9fb6922eb74ac35e2a7f6fc3bdba3d785e354c252f4e64c

Observation 5a892216-9ee5-4236-be83-f39bfade8dd6 · outbound

This paper cites We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections We provide an overview of the comparison of video-music datasets in Table 1 and an illustration of our dataset con- struction methodology in Figure 1

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:40.950159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:40.950159Z digest=sha256:281c75dc29fa2fc4cc003c0de8b2e8b6ec53a0f0ccd83c4a712f08468b68e4dc

Observation 96cf67e9-1b5d-4273-8464-6b780233dcf4 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.045157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.045157Z digest=sha256:62336eb981ae5a2c9b5c8e75b43d7f814588a24a79c666be8a0a0f12bb94c487

Observation 2749241b-9bb3-4ded-8c1c-40b2dd661c27 · outbound

This paper cites S” and “M.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections S” and “M

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:49:50.232025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.120789Z digest=sha256:24d3d99a5d666ea9cdb645443d43a59e7c3ddf89e4270351fa8bea0ea69cb8e4

Observation 2ccb242a-283f-46be-99c7-458adbebe840 · outbound

This paper cites Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Distributional FidelityOur evaluation on OES-Com re- veals that fine-tuning on OSSL enhances distributional fi- delity

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:41.198985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:41.198985Z digest=sha256:e96532b23633b9a2d72ceb88dfbed1f1a632dbdf3a6bd0cc7a5adb625213be10

Observation 330770bb-000c-4f02-bfd3-07d4e36c92cf · outbound

This paper cites To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections To show the effectiveness of our dataset, we adapted a text-to-music generation model with video conditions and fine-tuned it on our dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.630279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.272353Z digest=sha256:1be5983b5cd1be4939be773579822c4d05d7ecb1691302c9e77e956a23ad34db

Observation aa0b77f5-740f-4a30-94a9-5bef220bc6b6 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:49:57.502642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.339642Z digest=sha256:37479f63b3340fe66e299c57773446bb38e7dec81244de97608f23449ce81c7d

Observation a2b075f8-ad4d-41ad-9f10-d69886d45ea0 · outbound

This paper cites Soundtrack design: The impact of music on visual attention and affective responses,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Soundtrack design: The impact of music on visual attention and affective responses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.424060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.417326Z digest=sha256:c5077ef64d1ed424beb118c04079fb5344b03baf74cdaaae30aa0ee4e9d91405

Observation 931c243d-a0be-450f-8fe9-b6abb48a5c18 · outbound

This paper cites Multimodal deep models for predicting affective responses evoked by movies.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal deep models for predicting affective responses evoked by movies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.347541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.494318Z digest=sha256:c8b5fddb6b0cadd0788ce80813494fbdacd4d6e1da9fd686bb824703d48c563b

Observation 65756565-9c31-4f5b-acc1-9927be466874 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.863527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.584055Z digest=sha256:46d4868f6990371fa2a169925b01eb73ef3c9cfc3652abf9276772783067a19b

Observation 2d14e174-ec1a-42a6-8b9b-228347da9a6f · outbound

This paper cites Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attendaffectnet–emotion prediction of movie viewers using multimodal fusion with self-attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.192060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.662550Z digest=sha256:890855eef01ac48d0e8f10d2464131407129b1ea970a2776c9d80d85887d0e86

Observation 24dc64c5-acca-4af7-92cb-6137c395ebaa · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.551873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.739826Z digest=sha256:b0c8db88a4ed47c1a9e16315d0339cd500894f2adff49d3d4da25086d47c3c32

Observation 4b747be7-3c98-4c30-ba78-447e0aa5feb9 · outbound

This paper cites Analysis of the roles of film soundtracks in films,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Analysis of the roles of film soundtracks in films,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:57.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.820319Z digest=sha256:63735dd8cbcc61eef28d9a99e5e9cdabd69d79f597eb593041f92eede3c9f16a

Observation 7f35f1ab-b34e-4616-b036-c2cfcc4bb411 · outbound

This paper cites Actions in context,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Actions in context,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.941092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.884556Z digest=sha256:ca83eb2935921fa1154bb5241ee158260e4468b2a9941c1193c79035d0f6b9be

Observation 57e6cdd2-bd3f-4ded-a1b7-431a494383cb · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movieqa: Understanding stories in movies through question-answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.692466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:41.963832Z digest=sha256:fa3c0ccd58b6cd4ae5bff9d75af807bedb9356412a0acdb0343815e3932aa079

Observation 178f64c9-084c-4961-8dd7-9428a98bcc0f · outbound

This paper cites Movienet: A holistic dataset for movie understand- ing,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Movienet: A holistic dataset for movie understand- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.534391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.035597Z digest=sha256:96b91090095125f8abc132d7e64f18e2d2559416a16c49dbb23ea42b3c6b8954

Observation 8255ea0f-8c7d-4a2e-a083-578fb835f1e5 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mad: A scalable dataset for language grounding in videos from movie audio descriptions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.367972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.103527Z digest=sha256:bbf5cc9f489236b4f38b404cfb73b183bd0b11bc7e3353d60c3b3c0542c5343d

Observation 9f955c11-efdf-4d39-8a82-67e5f90dbc62 · outbound

This paper cites A dataset for movie description,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A dataset for movie description,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.278451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.197524Z digest=sha256:827db1f2d7e81afcc0fa41aa1fd47b74c4dc203da71dab140e8162958ba6900b

Observation f61528fb-e050-4ddf-bbf7-9f2c67e5fc0b · outbound

This paper cites Moviegraphs: Towards understanding human-centric situations from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Moviegraphs: Towards understanding human-centric situations from videos,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.112885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.279348Z digest=sha256:12353ee97c64569ff3365072643e9360afbaa52d7a97ecbb89794280a626b78d

Observation 9086561a-5ce8-448d-84b3-30760610658f · outbound

This paper cites Hlvu: A new challenge to test deep understanding of movies the way humans do,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Hlvu: A new challenge to test deep understanding of movies the way humans do,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:56.021400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.366444Z digest=sha256:1ba896d58dc06bf0858fbaf83f309b87b0226afec77affbf2bed7d355726962d

Observation 3237553e-2891-420a-abb0-136945d11bf6 · outbound

This paper cites Condensed movies: Story based retrieval with contex- tual embeddings,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Condensed movies: Story based retrieval with contex- tual embeddings,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.920877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.445904Z digest=sha256:0d2f7cd9c8b50b7bdbebe21ff9c9595382b228ca2b446fb126d0c49421f19297

Observation 4799cefd-2193-40fc-a7d3-0a70dd655e0d · outbound

This paper cites TeaserGen: Generating Teasers for Long Documentaries.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections TeaserGen: Generating Teasers for Long Documentaries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.511515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.511515Z digest=sha256:d3975db37001458fabb8815ff39b307e9cbe3a84798e89f8697e8adfaa0b1e28

Observation a9178205-829d-4963-b251-0bb2ad2877d2 · outbound

This paper cites Simple and control- lable music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Simple and control- lable music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.764831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.599104Z digest=sha256:b818dc671c625bd32b00d5bdc6cf32d6b9e43eeee55023f7210f0a8514ac545c

Observation f1a5fa92-d749-4ce4-9143-a20f91670f51 · outbound

This paper cites Attention is all you need,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Attention is all you need,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.675447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.675447Z digest=sha256:224e86fb6062d8fbb642e8932a09195cedc122494a6dfd6a14bfa469e43c8972

Observation f47ca879-4b18-487b-9530-1e89f691e266 · outbound

This paper cites MusicLM: Generating Music From Text.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusicLM: Generating Music From Text

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.750717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.750717Z digest=sha256:bec0edbffcc898bb04a04c6c14ffb7a20cb920460b8c075f37a1ade56da83c6f

Observation 744cb494-a203-4312-a521-fafac7c164da · outbound

This paper cites MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.827934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.827934Z digest=sha256:83042abbe05d35eda89c06ec7fb6b2e11b82ee3cfcc1328734b89e798ca66861

Observation 1450699d-fda3-4db1-be6d-12e8ad47a34d · outbound

This paper cites VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:42.900461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:42.900461Z digest=sha256:2e44f794dcb7d77e1c9a14d8077ddef25854823cad1f46e41420e0c199b117f3

Observation f33b4d9b-4689-465d-bb03-a5353cfe780b · outbound

This paper cites V2meow: Meowing to the visual beat via video-to- music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections V2meow: Meowing to the visual beat via video-to- music generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.665078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:42.977529Z digest=sha256:1d680f024eaf29dc762515688104d8bb20972b978df16c571087a3d6ac5e15a1

Observation ab03e73b-0eb7-424e-86c1-3955da2955bd · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:49.247373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:43.064927Z digest=sha256:6262fc4c7227569d80bfe185bc719fed728fc84b667e48e4b1c5dc8480155950

Observation 7f8710a4-62bd-4f3e-ab57-14e14f6e0164 · outbound

This paper cites Riffusion-stable diffusion for real-time music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Riffusion-stable diffusion for real-time music generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.456481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:43.156770Z digest=sha256:856fca93200c21837c13bab32ee77bd33161029359ad756e243f6e282cb30329

Observation 92b97fe1-1c0a-456e-a812-f804e48e40f9 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.230631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.230631Z digest=sha256:39e1b22a377d43b9d089ff2db378ebe004a3829976dfe32b3a9c2c51a0f83dcd

Observation ff44a7fe-ccb8-437f-a06d-699d6dba7156 · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.294050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.294050Z digest=sha256:30d1e186ee66c7183d2f0bf521e86aae69fc2f514569f9437fa5010673b9e435

Observation 41e83f31-e1a9-4b62-9f2b-e0d367cf7fda · outbound

This paper cites Efficient neu- ral music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient neu- ral music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.286800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:43.347990Z digest=sha256:c5fd277e979e21c8c96d11499d34334e301ca00ea6f86a2ee5601cce063cee3b

Observation 712b6267-1f37-4ea1-94e7-9cfb31d71bfd · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Mustango: Toward Controllable Text-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.437966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.437966Z digest=sha256:a75cbdac1584f1271e63d1f38a1d9268c970db6242538561f6fc707e8f76bdc2

Observation c168ce5b-df36-4fc9-a7bb-213c000ebc30 · outbound

This paper cites Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:48.971726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:43.507489Z digest=sha256:c54e92010c6180476f344d981140f4f347a33e0284cc3bdd95ad361a5bab40da

Observation 7e3009df-3e6b-43a6-a441-32a41699f533 · outbound

This paper cites Stable Audio Open.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Stable Audio Open

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.573592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.573592Z digest=sha256:d6897b6b7b64444b2c5863ad899698f6b01e4c844a821707292e8ad3fabe20c4

Observation d9d5c618-8e40-4ceb-a172-1711575d2e26 · outbound

This paper cites VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.639290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.639290Z digest=sha256:1deb5e438474bce263e1a8f71c40e90541541e6a3cd0d92c17d2a4d8cbf66ac1

Observation 92882141-7d2d-423f-9e38-e178bfac7bf1 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.716277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.716277Z digest=sha256:7bdea18db7d6221111420c98eaefa6bd8dda15051501db584bb7055f76fe8611

Observation 9b963403-04a1-475c-bfe0-7f4522db4f95 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:43.791478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:43.791478Z digest=sha256:76baec3dcefc0c93842d7dfb0d58d8ad5d9c03293c45d6de14b263c160497b06

Observation 7611343f-ec97-4eca-803e-ef17b1d82238 · outbound

This paper cites Music controlnet: Multiple time-varying controls for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Music controlnet: Multiple time-varying controls for music generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:55.092002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:43.880967Z digest=sha256:8822c6404825dc7e8a00f491496382dbe0e871cbb4898df72daa6a321bf976e1

Observation 2e6b4824-b7d1-43fb-a24f-ddec9c5092af · outbound

This paper cites DITTO: Diffusion inference-time t- optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO: Diffusion inference-time t- optimization for music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.987129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:43.945081Z digest=sha256:0e69bb444b033713f428014b520b1f6a88793f90a2f3322af1aa7834b539ccc6

Observation 6d17aff5-b6c0-460a-bd6a-44a3e4baf800 · outbound

This paper cites DITTO-2: Distilled diffusion inference-time t-optimization for music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections DITTO-2: Distilled diffusion inference-time t-optimization for music generation,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.851023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:44.026351Z digest=sha256:6773788812d52ff424e2ce2646100ffe3d06be52c07cbe639f7605303ba307fa

Observation 29b928b4-861d-4832-8757-7750994f7a18 · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.099128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.099128Z digest=sha256:a6353873f7e6cdabc72d6bc23d2cee413471ad1a5a3be3f92b62f7cf680d04e6

Observation b0d5dc2d-5a3c-431a-b6cc-59798dfe8b86 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.175395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.175395Z digest=sha256:f962026d05d5319b5fba382ebb03600556761344e6175e6ded5ab11b67dbf897

Observation 6ff4f4fd-722b-4618-8a9d-efd043432fe7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Foley music: Learning to generate music from videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.708344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:44.227354Z digest=sha256:4b1534dbb5851e38fdfe07771ec9cfdbd0f122aa827ca880ea00c773f1e4fa88

Observation 333a7a68-750d-4f41-8e71-6270ce04784c · outbound

This paper cites Video background music generation with controllable music transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video background music generation with controllable music transformer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.292165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.292165Z digest=sha256:e3938e3eddc78ca59c05209e24e07a0f83e4e8bf78d3a6cb8a5d89c1e990fdda

Observation 34a5a28c-d19a-423c-8526-5ea63f78a210 · outbound

This paper cites Vivit: A video vision transformer,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Vivit: A video vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.286387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:44.438782Z digest=sha256:661d9d60dbdaa0e89c2cb4ef2c4f0029b748f6cec7af0c86864aefe364c6ae06

Observation e1f7b1d7-f835-4570-92d1-e247baee4a2a · outbound

This paper cites Parameter-efficient transfer learning for nlp,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Parameter-efficient transfer learning for nlp,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.522171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.522171Z digest=sha256:a714f8bfb9b9adb9e4cc6dc3d4a6d76d18e030a3d963735d42eadb919e060865

Observation 7e28f1b3-58b5-4aa7-a2d5-331bf2c35972 · outbound

This paper cites AdapterFusion: Non-Destructive Task Composition for Transfer Learning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterFusion: Non-Destructive Task Composition for Transfer Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.596943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.596943Z digest=sha256:0f1c2037fbab1a718ea52b5f8e92822f46e150ad24c5e35b739083eaed78082d

Observation 98494540-e731-4a61-96b1-884d338d1334 · outbound

This paper cites AdapterHub: A Framework for Adapting Transformers.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AdapterHub: A Framework for Adapting Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.690125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.690125Z digest=sha256:b16445708729ad8e9fc7b5f56db75ae2eb3eca5bfc78ae1a20d0d451f82782d5

Observation 32c0adc6-f711-46ee-acfb-e46541056b6c · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.750532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.750532Z digest=sha256:bc287b9906bd818f2587ac4d5ec8d6a5691fc8eff5fce208ebe0518c9589c4b8

Observation 72702529-6ad2-48f3-9b62-3d1570f1714b · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.107266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:44.833785Z digest=sha256:b71f27cebd5da1b3c9a273b83c08b8557441d7ed0f09374001164c5e527aa87b

Observation ce65a2d3-06f7-4f66-9924-d8574cc2a9e5 · outbound

This paper cites Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:44.938203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:44.938203Z digest=sha256:b02f839db9285798e21e705e5e903a3dcb468a54bdf357ea11aeff3809b5c4cf

Observation c0c004cf-0283-4c5d-9f20-94aaf665c911 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.008013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.008013Z digest=sha256:6424117214d87a15527665b37c617ca86bdd0fbb3c1d66b826efe04ed76ad2d8

Observation 62e97fff-57f7-4bc8-8b5a-1aa173c5c3ae · outbound

This paper cites Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.121557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.121557Z digest=sha256:ad4c023222e3478c505d6245ee1eae50b5312ba06c42b953d337e376f430cb5a

Observation cb5d3735-9a17-41c5-9b44-bd2b4875388b · outbound

This paper cites Quantized GAN for Complex Music Generation from Dance Videos.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Quantized GAN for Complex Music Generation from Dance Videos

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.246773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.246773Z digest=sha256:5448c87bb7fce96ba351cc13a6c776a7f58779de358f924a9afdc2d267407799

Observation 50e3b075-93b3-4d87-b3f2-18ffe2df324f · outbound

This paper cites AI Choreographer: Music Conditioned 3D Dance Generation with AIST++.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:45.341616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.341616Z digest=sha256:73c03c23ca2c72bbfb4cc9d7b8123c4d94d1907c95c59f1776bbee3f997a2555

Observation e25abf34-5f26-4e9a-8f49-7743e2818e36 · outbound

This paper cites Video back- ground music generation: Dataset, method and evalu- ation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video back- ground music generation: Dataset, method and evalu- ation,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:54.547185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:45.429434Z digest=sha256:8fc10ed48f6a8984791f2769f8e0a1f1923fe876e33cf27241154aca84bd6d78

Observation e3261b1c-afce-44ca-b1f9-9cf8749ed34e · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.721432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:45.523930Z digest=sha256:d5880bcd01d288c6b3096703b0aa8ced2e85dc15ee302d1d34242570848c8af2

Observation 5c29d675-b417-483e-b7db-5049e09cea7e · outbound

This paper cites The Kinetics Human Action Video Dataset.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The Kinetics Human Action Video Dataset

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.834553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.834553Z digest=sha256:8047db8b4e1a552e3b3067519c7983236e7cb2aa9c546ace7e5d1e834b8f54a2

Observation b536ca68-8cc5-4f46-b632-2d4f528d1b24 · outbound

This paper cites Diff- bgm: A diffusion model for video background music generation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff- bgm: A diffusion model for video background music generation,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.453381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:45.792613Z digest=sha256:833243753af6a87a428a404f707979ce0d757fe60f828318fcc50fd4d994b29f

Observation ed757164-55b8-47f6-a5e6-548e964111ed · outbound

This paper cites The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections The nes video-music database: A dataset of symbolic video game music paired with gameplay videos,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:53.215919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:45.876102Z digest=sha256:d1c345f3f9a9d9066c559f3336b347bf72bee83d937a88fa9b4fd9ca24aa930e

Observation 4a9c785d-d9ee-4467-81df-a3c8c7f3d811 · outbound

This paper cites an unresolved cited work.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Unresolved cited work

Reference 65

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T00:49:48.478189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:45.996145Z digest=sha256:ef43bf3abe24554f4f789014609d0b51e261ab573cf959d1743a9b6892077dc6

Observation e34b0e31-a537-43d3-a04c-e3180ff025b9 · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Benchmarks and leaderboards for sound demixing tasks,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.959182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:46.099559Z digest=sha256:adb14076ad1fa890d970ce8724e523b6f94123a10675b724b40684ad63f74c04

Observation a7b7110e-f6db-48df-8c8f-0f298f8ce98e · outbound

This paper cites pyaudioanalysis: An open-source python library for audio signal analysis,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections pyaudioanalysis: An open-source python library for audio signal analysis,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.616912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:46.248153Z digest=sha256:2dd70f92e3b01d632ccb114780c718bff908c2600d08504881c78427b36b3715

Observation 9dfd5c2b-5a9e-4eef-a362-a5f81764df87 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.378501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.378501Z digest=sha256:5f78ec4851a4292301eb6739bbf4e939d4734fe3eed58b6e43eaf1e177b059d8

Observation 50cc1ea6-278d-4b79-96b4-68bbdbe10c97 · outbound

This paper cites A circumplex model of affect.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A circumplex model of affect

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.500264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.500264Z digest=sha256:54fb8c7b61a5517c8c2accc2ca8f254206ec7b3fdd0f27b60d3d249eb1adfde6

Observation 44613096-827c-4252-953e-7ace2a26ea60 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.604276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.604276Z digest=sha256:d9e1e136a26fa40b86c32dc4579f43157eef83284514cafdb7396fc2211d4c7a

Observation 7fc4e7aa-0042-4bf6-bfa3-f9f85ef5d283 · outbound

This paper cites High Fidelity Neural Audio Compression.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections High Fidelity Neural Audio Compression

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.732431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.732431Z digest=sha256:327385a48d7e54e11204d54992735a8c97eb4691a261ce6c575e873de592141a

Observation 00b7d1d0-bfaa-4804-b2eb-2456a7ab899f · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.161900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:47.758112Z digest=sha256:8219eba47b18d25aa7891e722400298c50b1687aa9d857e58a99b4bf8223e399

Observation 4f9aa396-4221-48f1-89c8-704c1015dd39 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LoRA: Low-Rank Adaptation of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:46.967955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:46.967955Z digest=sha256:ed0a97cedc0ab2b041a8d4c77625e6cbcb0a79c69e0fefedf3b843f17f8af01a

Observation 0ebc6fab-03ae-4375-a19c-4f2db7f4fc0e · outbound

This paper cites Language models are few-shot learners,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Language models are few-shot learners,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:52.186777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:47.082948Z digest=sha256:d5da52df3dda74d8e3f4a1aab7eef4fbff446d4fbdc9c8e0c1c278e030915cba

Observation 69f6d0cd-4b1a-421a-8448-b2f4c0123638 · outbound

This paper cites Design guidelines for prompt engineering text-to-image generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Design guidelines for prompt engineering text-to-image generative models,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.823139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:47.186083Z digest=sha256:48a3934cae1f9f3b7d2d7fc8253d392382bc78c49a85fb1afd55d8ee9b7ad1a9

Observation 6e8f562d-5af2-4bc7-a4b9-8d521f795dd6 · outbound

This paper cites A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.311232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.311232Z digest=sha256:e7dc4d7132ea5cfee0f9fe3a7d86d98280c663e35467db005f7f3aaa6941a2dd

Observation 18e95974-ec41-4849-a687-64572867160f · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.415707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.415707Z digest=sha256:26532deb1342a2d2428a9c3e35fc6146f4e0c34f1cd542bd3227270c16cbd050

Observation 81eda005-80ec-47dd-8f4c-298358cfd8da · outbound

This paper cites Decoupled Weight Decay Regularization.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Decoupled Weight Decay Regularization

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.484231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.484231Z digest=sha256:b5a376ac0299afb57358869f9b11813cc081c51520e4a8faae3e799ea0c20f60

Observation 5a4deb80-728e-4503-99cc-5da2dccca4ba · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.550534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.550534Z digest=sha256:46a1eaa06d167121b0e3e3df28541eb02143bd32e5d2af9aaa52d680526fb48a

Observation 50ab875a-1595-4b31-bd65-6e8a2f65f51d · outbound

This paper cites Presto! distilling steps and layers for accelerating music generation.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Presto! distilling steps and layers for accelerating music generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:51.477011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:47.631070Z digest=sha256:10d9f078345d8b36a1fda8ac7ef57c4323cfef17d6d97cb6507832b5de76e449

Observation 17565a16-a129-4653-ad4d-251f7255020c · outbound

This paper cites Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Diff-A-Riff: Musical Accompaniment Co-creation via Latent Diffusion Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.697740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.697740Z digest=sha256:c284409093cfc21615ce12d3b284322e33f9b472149ddc3f69ea2107e1c43c45

Observation 7bc0f0b4-575b-4092-9629-2b5e5b319de3 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.832625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.832625Z digest=sha256:1522d5fa03fd7710017e237582cdaba22046563d21d5f5058cd89edc68b7483b

Observation 5851456f-0269-48f1-ae3c-d32dc18108bb · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Reliable fidelity and diversity metrics for generative models,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:49:50.882373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:47.896765Z digest=sha256:a3d69765d22da26d0488663029923e3fe20a69a127556b1795c6c9e473b61d98

Observation cc2e518c-34dc-4078-9479-835127505348 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Efficient Training of Audio Transformers with Patchout

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:47.983763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:47.983763Z digest=sha256:6750f535ce045c1278c47b0ff569801e8ba136d49a5eda55fc50eacb89bdbd52

Observation 91732895-4c5d-4970-93e6-4f4a5c4fe0e1 · outbound

This paper cites Available: http://dx.doi.org/10.1016/j.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Available: http://dx.doi.org/10.1016/j

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-07T00:49:45.673911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:49:45.673911Z digest=sha256:aa41ff998d4a6ac779d93393035aef6f89a137e4ed7afcdda4df1d89b543184d

Pith citing papers

Observation 5af4092a-926b-44ce-b64d-34ac5cdc75e5 · inbound

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections cites this paper.

Video-Guided Text-to-Music Generation Using Public Domain Movie Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:49:50.536973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T00:49:40.786176Z digest=sha256:26f43669896461dd9b32b78967915b223cff03053f0581ef32f63a554c051709

Observation 8147f160-bd41-4192-82f0-c30b73745ab6 · inbound

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections cites this paper.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.500746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.500746Z digest=sha256:718de42c84ef1d5b8f244b3def6445eac6eaab8077f9baeb6e1310ccc36a5999