Pith. sign in

Paper Citation Record · LEDGER

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2608.11576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11576 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.599014Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:36:42.391185Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T00:36:42.953731Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact6
  • verified fuzzy29
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9994dfab-e618-4b92-91d6-f403c9decf4b · outbound

This paper cites Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.957902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.391185Z digest=sha256:afba87e97f53e72e6e2daf43cf0f6a7094863a5a2f5c212fcc33ebe456474bb2

Observation 49f7049e-e796-46e3-9af6-b670310f557f · outbound

This paper cites Early work in modern video-to-music generation systems uses large-scale web music-video corpora and autoregressive modeling over semantic acoustic tokens to generate music [8].

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Early work in modern video-to-music generation systems uses large-scale web music-video corpora and autoregressive modeling over semantic acoustic tokens to generate music [8]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.330732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.395669Z digest=sha256:3fec7335008a0345f12c8851291445712cfb2f634559b8896ff78eca6327333f

Observation d1f0f9bd-1d9e-4177-aebe-5ce875711696 · outbound

This paper cites trance music.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections trance music

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.319454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.399453Z digest=sha256:90a493930ca075d73c80d729c9e0175ef8ef8ff485b19f2fde9f50c8bdca51bb

Observation f0743c4e-cba8-4a45-915f-35bb08426abd · outbound

This paper cites an unresolved cited work.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:36:43.308324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.403577Z digest=sha256:e7f246ca1371a52a9522528a10494d3dcea2cb1cca25f9557a64d0beb8dca50e

Observation 1369da0c-923d-4a28-b521-da386bbf7aa4 · outbound

This paper cites First, we bench- mark video-to-music generation tasks by comparing existing video-to-music generation models trained and evaluated on identical data, using OSSL-v2.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections First, we bench- mark video-to-music generation tasks by comparing existing video-to-music generation models trained and evaluated on identical data, using OSSL-v2

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.298362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.408016Z digest=sha256:3145bbf49031360e60d81d990a027861d7a89e67fbf26247af0d2da7ab0eddc2

Observation 442247b2-c215-4420-b9d4-a54b69ada218 · outbound

This paper cites +Dialogue.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections +Dialogue

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.287922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.412530Z digest=sha256:14072ea5f2659575a58568bf63149d5f394a03c2ee64e8f07d006fd72ae34834

Observation 4d97db7e-2296-4f1a-a2c9-6fff5d6fbc44 · outbound

This paper cites Be- cause the dataset is free from link rot and does not require separate web scraping, our dataset is suitable as a durable benchmark for the field.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Be- cause the dataset is free from link rot and does not require separate web scraping, our dataset is suitable as a durable benchmark for the field

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.275395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.416447Z digest=sha256:6aefcdd37c2ea4b7e70a305544f368ce04830ce807105b0198909a9ee127a46d

Observation df5e0c5b-906b-47e3-959a-cc4b271b19e6 · outbound

This paper cites Teaser Generation for Long Documentaries and Educational Videos.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Teaser Generation for Long Documentaries and Educational Videos

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.264105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.420783Z digest=sha256:853a07988a2ed000204a9cae1bd4202bd59a03c45154cf0ea69584d4c6851340

Observation 48d25803-3e24-4035-bc77-87021a8676c0 · outbound

This paper cites Attendaffectnet–emotion pre- diction of movie viewers using multimodal fusion with self-attention,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Attendaffectnet–emotion pre- diction of movie viewers using multimodal fusion with self-attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.252296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.424057Z digest=sha256:284dac4bd92a2e8bea787d0581251f8401a817c1f5091bd5db9b664906288436

Observation cefa76f8-82df-4414-a38f-80e911720855 · outbound

This paper cites Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.427623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.427623Z digest=sha256:a1b7883ed0a8b4f0949fb56a3d66a3f414d810868d762162558f4c00307130fe

Observation aefa4075-821b-4a94-9155-3c07c6c43813 · outbound

This paper cites The cognitive processing of film and musical soundtracks,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections The cognitive processing of film and musical soundtracks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.240784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.431355Z digest=sha256:1e8d26d0da3bd52e66e9c3fd70be56ab471a0014bf1363943929d75dbbd5539a

Observation 008a119a-ccf9-473b-b194-b9633b4b2f70 · outbound

This paper cites Multimodal deep models for predicting affec- tive responses evoked by movies.,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Multimodal deep models for predicting affec- tive responses evoked by movies.,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.229494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.434869Z digest=sha256:c14e5fb6bbd6cc227611b65f0175eecbef91965489c6d2e28021a2e51c6299de

Observation fa33de3d-8bb4-4105-b5cc-52ed53cdaad7 · outbound

This paper cites On music’s potential to convey meaning in film: A systematic review of empirical evi- dence,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections On music’s potential to convey meaning in film: A systematic review of empirical evi- dence,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.219280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.438501Z digest=sha256:4e6d026983a623d23b613fc679004ab4b089745b26681d511b146fc006514fd7

Observation e4bb1dcc-a488-4e3f-b982-43011b6020d4 · outbound

This paper cites Emotion Embedding Spaces for Matching Music to Stories.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Emotion Embedding Spaces for Matching Music to Stories

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.442268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.442268Z digest=sha256:eac27408657fac455cdf4eee29b32c0dd69ec70530885cd10a90d32cbe81e1b9

Observation b24fc0fe-dfd4-463c-96d2-b5d1293e74d7 · outbound

This paper cites Foley music: Learning to generate music from videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Foley music: Learning to generate music from videos,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.208400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.446296Z digest=sha256:7a4ad372dd61d9168862654a6d9af574f05de980fafe05ed60a8f225124fb7fa

Observation 4b2cefab-443c-4b9e-8172-064e5352d578 · outbound

This paper cites V2meow: Meow- ing to the visual beat via video-to-music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections V2meow: Meow- ing to the visual beat via video-to-music generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.197380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.450042Z digest=sha256:3c5835ba0a605cc5363442cdaf3cc5b1b6b183cbae9aca5fb5ec397f506fd3e0

Observation fd6a3ed6-191b-4b45-8089-1bf9bbd4218a · outbound

This paper cites Vidmuse: A simple video-to-music generation framework with long-short-term modeling,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vidmuse: A simple video-to-music generation framework with long-short-term modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.185861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.453909Z digest=sha256:0472f43cd93d980fd42f5bfd724226f096f20640d7355f8d75122e04c93ae89c

Observation 1ceb4568-70ff-40a4-9f39-ff64d417ccc9 · outbound

This paper cites Vmas: Video-to-music generation via semantic alignment in web music videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vmas: Video-to-music generation via semantic alignment in web music videos,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.174223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.457562Z digest=sha256:348e87908e2572ece402245c8c6f8485adda93e23273fdaff5bbd440addf3519

Observation 0a616f9c-5105-4098-9803-76aed2621f3b · outbound

This paper cites Sonique: Video background music generation using unpaired audio- visual data,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Sonique: Video background music generation using unpaired audio- visual data,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.163597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.461759Z digest=sha256:090c85f1c723f61c1aec005fbc9c66cff0fb124256a153c3410a3594b6d370ad

Observation 707f5136-b83a-41bc-bac4-8007b047d9ec · outbound

This paper cites GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.465952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.465952Z digest=sha256:64919e3958007a062712020fe2593ab3ce230a81f69306fa7c1cf713aaa51343

Observation 19f194b0-b420-4d16-8e89-42189f1c22a4 · outbound

This paper cites Vision-to-Music Generation: A Survey.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Vision-to-Music Generation: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.470428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.470428Z digest=sha256:a6c171a0339b958407a35f83f69327786eb025bd813a908f0c3d3cf62742324e

Observation a329da2a-9f37-45c7-a4d3-bad6ee9ea753 · outbound

This paper cites Ai- based chinese-style music generation from video con- tent: a study on cross-modal analysis and generation methods,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Ai- based chinese-style music generation from video con- tent: a study on cross-modal analysis and generation methods,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.151434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.474275Z digest=sha256:95cd1cc0d00d3f3170bfaeb43344201e8da10608befc959a58b45cc53967dc2f

Observation 81b7990b-37ce-403d-83d9-2d9c84b6d211 · outbound

This paper cites V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.893043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.478417Z digest=sha256:178475cc751813a2a682c0d4082aaa9e58a5568e1532f033f562f830b1b249fd

Observation aa944ece-843f-4cd3-9506-b9cc63eb9cc7 · outbound

This paper cites Diff- v2m: A hierarchical conditional diffusion model with explicit rhythmic modeling for video-to-music genera- tion,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Diff- v2m: A hierarchical conditional diffusion model with explicit rhythmic modeling for video-to-music genera- tion,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.139456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.482121Z digest=sha256:33eac161ad8770db381a15f5445c86f4b263ec69a8783a9a613da4dda6b404ed

Observation 10022ec1-ca19-4aa1-bf0f-aad20ccbb1fe · outbound

This paper cites Quan- tized gan for complex music generation from dance videos,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Quan- tized gan for complex music generation from dance videos,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.128702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.489105Z digest=sha256:d9ddbe6e658197f3bb7485f614672c6dfc90065ff3b8dfbade2dbe09a67fe8fd

Observation b2d7c4b9-30bd-4184-86b6-a939bf58ac3e · outbound

This paper cites Video background music genera- tion: Dataset, method and evaluation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video background music genera- tion: Dataset, method and evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.117905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.492941Z digest=sha256:9ee4f477da1684c47f33fd3ff72fa76b4b2577d11ebf3e4e8fbc46733b88323b

Observation 990607ec-3e49-4516-874e-35c7c1f52409 · outbound

This paper cites Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.496612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.496612Z digest=sha256:873756689855aef556c1c2cc6a0c596af7057824a30f109ab098d25d66ffed0c

Observation 8147f160-bd41-4192-82f0-c30b73745ab6 · outbound

This paper cites Video-Guided Text-to-Music Generation Using Public Domain Movie Collections.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Guided Text-to-Music Generation Using Public Domain Movie Collections

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.500746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.500746Z digest=sha256:718de42c84ef1d5b8f244b3def6445eac6eaab8077f9baeb6e1310ccc36a5999

Observation fa6e8aab-9e17-454e-b054-7b958e060045 · outbound

This paper cites Acoustic profiles in vocal emotion expression.,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Acoustic profiles in vocal emotion expression.,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.106258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.504700Z digest=sha256:b16281452b69b75b2e596b6d27d4c6d86c187749d2bba01283b89dc2dda64b87

Observation 6abfe8f2-f194-4e2c-84c1-41efdb5a99a1 · outbound

This paper cites Background ducking to produce esthetically pleasing audio for tv with clear speech,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Background ducking to produce esthetically pleasing audio for tv with clear speech,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.092776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.508491Z digest=sha256:507384bceb714cc519cdbaf26583ef33e7f6be95b820088665f97be140fddcec

Observation bee2f69a-b172-452f-ae3e-7047d87f4192 · outbound

This paper cites Improving dialogue intelligi- bility in streaming media,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Improving dialogue intelligi- bility in streaming media,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.080172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.512290Z digest=sha256:48d5e0e2b7727dc767612fee06b6d9a815d740b0947b08c023d2c36668ecde0c

Observation 120bc34a-93c5-4792-90e6-0256dc84d0c3 · outbound

This paper cites Creating a multitrack clas- sical music performance dataset for multimodal mu- sic analysis: Challenges, insights, and applications,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Creating a multitrack clas- sical music performance dataset for multimodal mu- sic analysis: Challenges, insights, and applications,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.068775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.517678Z digest=sha256:f5e1ac5cc1ada76cb8f3c777f63cbe6eb5492245e46bd58504f8019b3cf4cfcf

Observation ebc889d5-7b93-4f4d-9a5e-3aeb1039d244 · outbound

This paper cites Video2music: Suitable music generation from videos using an affective multimodal transformer model,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video2music: Suitable music generation from videos using an affective multimodal transformer model,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.058045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.521020Z digest=sha256:708f3c64c984d885b399ca58260341052759e1ea94048e3dec313c77bd6dd364

Observation e1e7e50f-29fd-4335-a102-3ddf5172aec5 · outbound

This paper cites Extending Visual Dynamics for Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Extending Visual Dynamics for Video-to-Music Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.524714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.524714Z digest=sha256:9257be0892a8fa28e7bc275d7c8e9f988f46d03b3f2f991561bb3471b33f0620

Observation bd94a98a-47b9-4228-ba40-0b7c2eac11fb · outbound

This paper cites VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.528088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.528088Z digest=sha256:e80c3851b1790e4068595a73b36d1c033d7c6a02787e8137c3ccead8e65b719f

Observation f06b666f-bace-4d58-a581-5d524fbabd79 · outbound

This paper cites MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.531770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.531770Z digest=sha256:2f4d8d5b687f7828517f6217fdd90fe73b2edc5d3946142ecb5535e66d1ddd0c

Observation 6e1b4654-799c-4717-9f79-8ce89d3db55f · outbound

This paper cites Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video echoed in music: Semantic, temporal, and rhythmic alignment for video-to-music generation,

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:36:42.818139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.535975Z digest=sha256:18896877fa9996819f7f89ab2cec155b3b9a38160e93af28452e1c2bfb913176

Observation 7c0c2411-6d68-43cc-8757-d2ece80bb540 · outbound

This paper cites Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.743314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.539678Z digest=sha256:23f3396efaa222c7aa98e9a689a6ff824ba91b93f266c553a85767035930e0ee

Observation 21d31ded-83a0-4005-be0e-91e719cd9342 · outbound

This paper cites Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.726835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.544084Z digest=sha256:c46cc7c69d88ca543227b498879ddb02e073c4ed8422b0c1aa338e708afd0539

Observation 1209e8e0-5c05-4638-a493-3f2d2ff08163 · outbound

This paper cites FilmComposer: LLM-Driven Music Production for Silent Film Clips.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections FilmComposer: LLM-Driven Music Production for Silent Film Clips

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.548207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.548207Z digest=sha256:44e29bdfed7cc3efcba400c24d9c2c406e6ab9ddfcf4750c591e8513ff543ed6

Observation 72562688-b4de-4467-b3ec-8b104b63b241 · outbound

This paper cites M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.551715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.551715Z digest=sha256:57fd22086f1e749344d8c675805dbf422d0af545540977215de1a60a5deab1fb

Observation 16bf29af-c87e-4b33-9971-4c33d6f8d77d · outbound

This paper cites MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.555094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.555094Z digest=sha256:6632bb9503c30dac55d3aa65b85e4b80f4c773cf1c23569c3edfafe550b2ae1e

Observation 8d8c3492-5c3a-4909-9fa0-35fc7b2ed464 · outbound

This paper cites Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.558498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.558498Z digest=sha256:78ccbbbbbd19b3c07f2b66b49e9549db925d2d7f3fa75447f80d78de52605986

Observation b074ce61-01be-4387-8b04-d49dfe2f24fd · outbound

This paper cites Benchmarks and leaderboards for sound demixing tasks.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Benchmarks and leaderboards for sound demixing tasks

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.661931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.562317Z digest=sha256:8138b8315f3db63d0691b6a614a62d1eb5b286ae54fc53e1e1d0ab4eb790b49b

Observation eb3d6160-8e0d-4e0f-a8db-8702c047d2af · outbound

This paper cites Panns: Large- scale pretrained audio neural networks for audio pat- tern recognition,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Panns: Large- scale pretrained audio neural networks for audio pat- tern recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.046961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.566111Z digest=sha256:0c6f3decb6c1b84ad38db23f005981d38752dc71cf446224bf56a16db6a4dfb1

Observation 653a13bf-c98b-4d2a-83cf-051647f1e180 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Film: Visual reasoning with a general conditioning layer,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.569655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.569655Z digest=sha256:63a1d454b909e7eaf1f86890d6b10e7b87e587c1c5271779d3fd2b6077b825fb

Observation da8bf902-2b3d-4bbe-9776-dc74b80f2ac7 · outbound

This paper cites Simple and controllable music generation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Simple and controllable music generation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.028037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.573729Z digest=sha256:3b7208f6f6168b671acbbda0d9fb4b5c9b2fefa1f7ce1cab34d6e5416215b1ae

Observation 3ce23509-f9dd-4bc1-aec2-972436a940c0 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Learning transferable visual models from natural lan- guage supervision,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.578021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.578021Z digest=sha256:2afff8dc130bbaa85673d21d44241de1cb9aef02a8feba2f26020edebe6cca1e

Observation 3eee001c-eab3-4d51-b092-e28182dd0b42 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:43.008983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.581737Z digest=sha256:5e89ffb8009a9f6f0656ab3b473ed21884314d167615029aa8ee105d4597b649

Observation 8dba42b6-6f13-4d4f-ac83-62e87cb04ec0 · outbound

This paper cites Stable audio open,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Stable audio open,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.585657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.585657Z digest=sha256:371dab7868ff23db3ec58aa35a55dc6ba6a0881c890f34161a3fa2f3d32e01fc

Observation 845fcd3d-3188-46eb-b21f-4d82c9bd4ff7 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.589124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.589124Z digest=sha256:4917632f2b6c79b1210ad417eae0bc629b9f361ec7cd52c8480dfef0a12b6b71

Observation 033875e8-7627-4d8c-9f7d-d2ee6ffd8ca0 · outbound

This paper cites Reliable fidelity and diversity metrics for generative models,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Reliable fidelity and diversity metrics for generative models,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:42.986288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.592522Z digest=sha256:5a665f197defbc6ecd2e4e61a73f8ba820ad484419865270b672f600aa6ae60c

Observation 94ed3f79-cf9d-4933-9880-15011d8c22bb · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Large-scale contrastive language-audio pretraining with feature fu- sion and keyword-to-caption augmentation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:36:42.972018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.595514Z digest=sha256:9824ba679384a5ae6cbf61cf2827b5653a7c4fd16241cbd70d0d5fba3ef9095e

Observation 8a87dcdc-06b6-4acc-8217-f75e2e86a756 · outbound

This paper cites Efficient Training of Audio Transformers with Patchout.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Efficient Training of Audio Transformers with Patchout

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:42.599014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:42.599014Z digest=sha256:669ba9ea3ee2bd6c37d1853b3afa3b4c276db24a4f346a80dcd30e0d7a3b0c7b

Pith citing papers

Observation 9994dfab-e618-4b92-91d6-f403c9decf4b · inbound

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections cites this paper.

Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:36:42.957902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:36:42.391185Z digest=sha256:afba87e97f53e72e6e2daf43cf0f6a7094863a5a2f5c212fcc33ebe456474bb2