Pith. sign in

REVIEW 6 cited by

VampNet: Music Generation via Masked Acoustic Token Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04686 v2 pith:U5H6NMPK submitted 2023-07-10 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords vampnetmusicacousticcoherentcompressionduringinpaintingmasked
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by applying a variety of masking approaches (called prompts) during inference. VampNet is non-autoregressive, leveraging a bidirectional transformer architecture that attends to all tokens in a forward pass. With just 36 sampling passes, VampNet can generate coherent high-fidelity musical waveforms. We show that by prompting VampNet in various ways, we can apply it to tasks like music compression, inpainting, outpainting, continuation, and looping with variation (vamping). Appropriately prompted, VampNet is capable of maintaining style, genre, instrumentation, and other high-level aspects of the music. This flexible prompting capability makes VampNet a powerful music co-creation tool. Code and audio samples are available online.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    LaDA-Band generates complete vocal-aligned instrumental accompaniments with discrete masked diffusion, claiming simultaneous gains in audio fidelity, long-range coherence, and orchestration quality.

  2. SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

    eess.AS 2024-12 conditional novelty 6.0 of 10

    A language-model song generator adapted for multi-task editing, supporting segment-wise and track-wise modifications alongside full-song generation.

  3. Watermarking Training Data of Music Generation Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Injected audio watermarks can be detected in music generated by a fine-tuned MusicGen model; simple tones work best, and a neural watermark requires dozens of repeated embeddings.

  4. IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling

    eess.AS 2025-05 conditional novelty 5.0 of 10

    On AudioCaps, IMPACT reports the best Fréchet Distance and Fréchet Audio Distance among the compared systems while generating audio faster than diffusion baselines.

  5. Video-Guided Foley Sound Generation with Multimodal Controls

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.

  6. Improving Controllability and Editability for Pretrained Text-to-Music Generation Models

    cs.SD 2024-11 conditional novelty 2.0 of 10

    A thesis compilation presenting three complementary approaches to improving editing and control of pretrained text-to-music models, with Instruct-MusicGen demonstrating the strongest stem-level editing results.

Pith tools