Pith. sign in

REVIEW 5 cited by

Generative Disco: Text-to-Video Generation for Music Visualization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.08551 v2 pith:CJAHRWH7 submitted 2023-04-17 cs.HC cs.AI

classification cs.HCcs.AI
keywords musicgenerativediscogeneratedgenerationhelpsholdsintervals
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Visuals can enhance our experience of music, owing to the way they can amplify the emotions and messages conveyed within it. However, creating music visualization is a complex, time-consuming, and resource-intensive process. We introduce Generative Disco, a generative AI system that helps generate music visualizations with large language models and text-to-video generation. The system helps users visualize music in intervals by finding prompts to describe the images that intervals start and end on and interpolating between them to the beat of the music. We introduce design patterns for improving these generated videos: transitions, which express shifts in color, time, subject, or style, and holds, which help focus the video on subjects. A study with professionals showed that transitions and holds were a highly expressive framework that enabled them to build coherent visual narratives. We conclude on the generalizability of these patterns and the potential of generated video for creative professionals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Not Like Us, Hunty: Measuring Perceptions and Behavioral Effects of Minoritized Anthropomorphic Cues in LLMs

    cs.HC 2025-05 conditional novelty 6.0 of 10

    An experiment with 985 participants found that LLM agents using AAE or Queer slang did not increase reliance or trust over a standard English agent, and AAE speakers significantly preferred the standard English agent.

  2. VideoDiff: Human-AI Video Co-Creation with Alternatives

    cs.HC 2025-02 conditional novelty 6.0 of 10

    Aligned timeline and transcript views for multiple AI-generated video edits let creators compare and refine alternatives roughly twice as fast, with lower workload and higher final-video satisfaction, in a within-subj...

  3. LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    LeviTor controls 3D object trajectories in generated videos by feeding K-means clustered mask points with estimated depth into a video diffusion model.

  4. Examining the Usage of Generative AI Models in Student Learning Activities for Software Programming

    cs.SE 2025-11 conditional novelty 5.0 of 10

    ChatGPT helps students pass programming tests but not understand concepts; both heavy reliance and minimal use lead to weaker learning.

  5. Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching

    cs.HC 2025-01 conditional novelty 5.0 of 10

    A sketch-driven text-to-image interface with analogical inspiration and sketch scaffolding increases designers' self-reported inspiration, exploration, and co-creation compared to a ControlNet baseline.

Pith tools