Pith. sign in

REVIEW 2 cited by

Vision-to-Music Generation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.21254 v1 pith:MRMREBFY submitted 2025-03-27 cs.CV cs.AIcs.MMcs.SDeess.AS

classification cs.CVcs.AIcs.MMcs.SDeess.AS
keywords generationvision-to-musicmusicresearchfieldchallengesexistingfurther
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Vision-to-music Generation, including video-to-music and image-to-music tasks, is a significant branch of multimodal artificial intelligence demonstrating vast application prospects in fields such as film scoring, short video creation, and dance music synthesis. However, compared to the rapid development of modalities like text and images, research in vision-to-music is still in its preliminary stage due to its complex internal structure and the difficulty of modeling dynamic relationships with video. Existing surveys focus on general music generation without comprehensive discussion on vision-to-music. In this paper, we systematically review the research progress in the field of vision-to-music generation. We first analyze the technical characteristics and core challenges for three input types: general videos, human movement videos, and images, as well as two output types of symbolic music and audio music. We then summarize the existing methodologies on vision-to-music generation from the architecture perspective. A detailed review of common datasets and evaluation metrics is provided. Finally, we discuss current challenges and promising directions for future research. We hope our survey can inspire further innovation in vision-to-music generation and the broader field of multimodal generation in academic research and industrial applications. To follow latest works and foster further innovation in this field, we are continuously maintaining a GitHub repository at https://github.com/wzk1015/Awesome-Vision-to-Music-Generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ML Defender (aRGus NDR): An Open-Source Embedded ML NIDS for Botnet and Anomalous Traffic Detection in Resource-Constrained Organizations

    cs.CR 2026-04 unverdicted novelty 7.0 of 10

    ML Defender achieves F1=0.9985 on CTU-13 Neris botnet detection with a dual fast-detector plus random forest model, outperforming Suricata (zero alerts) and Zeek (F1=0.042) in a three-paradigm comparison.

  2. Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A new public-domain film dataset and a frame-by-frame dialogue-conditioning module improve video-to-music generation on paired-fidelity metrics.

Pith tools