ArtifactNet extracts codec residuals from spectrograms with a 4M-parameter network to detect AI music at F1=0.9829 and 1.49% FPR on unseen tracks from 22 generators, outperforming larger baselines.
Music source separation in the waveform domain
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
The Spheres dataset provides multitrack orchestral recordings with isolated instrument stems and acoustic characterizations to support supervised machine learning for music source separation in the classical domain.
EnCodec is an end-to-end trained streaming neural audio codec that uses a single multiscale spectrogram discriminator and a gradient-normalizing loss balancer to achieve higher fidelity than prior methods at the same bitrates for 24 kHz mono and 48 kHz stereo audio.
A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18-HQ tests.
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
Diffusion-based refinement followed by consistency distillation improves music source separation quality and inference speed across U-Net and BS-RoFormer backbones on Slakh2100 and MUSDB18.
MaineCoon is presented as the first 22B-parameter real-time streaming audio-visual autoregressive model optimized for social-interactive applications, using novel training techniques and an agentic inference framework.
A shared continuous-latent flow model generates music from text/vision or extracts a target source from a mixture via visual-audio alignment, gated modulation, and dynamic modality masking.
DTT-BSR+ is a generative-then-regression cascade for music source restoration that reports MMSNR gains over single-stage DTT-BSR and X-LANCE on most stems while noting a distribution-vs-reconstruction trade-off via FAD.
A text-to-music model is improved by conditioning on and selecting with a human preference reward, where expert iteration on top outputs contributes the largest measured gains on 100 Song Describer prompts.
citing papers explorer
-
ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics
ArtifactNet extracts codec residuals from spectrograms with a 4M-parameter network to detect AI music at F1=0.9829 and 1.49% FPR on unseen tracks from 22 generators, outperforming larger baselines.
-
The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
The Spheres dataset provides multitrack orchestral recordings with isolated instrument stems and acoustic characterizations to support supervised machine learning for music source separation in the classical domain.
-
High Fidelity Neural Audio Compression
EnCodec is an end-to-end trained streaming neural audio codec that uses a single multiscale spectrogram discriminator and a gradient-normalizing loss balancer to achieve higher fidelity than prior methods at the same bitrates for 24 kHz mono and 48 kHz stereo audio.
-
Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models
A Conformer-conditioned decoder-only language model generates discrete tokens via a neural audio codec to separate four music stems, reaching near state-of-the-art perceptual quality and top NISQA on vocals in MUSDB18-HQ tests.
-
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
A FiLM-conditioned transformer masker on DAC codec latents performs text-guided sound separation with claimed efficiency, but the main comparison against AudioSep is confounded by asymmetric input processing.
-
Improving Music Source Separation with Diffusion and Consistency Refinement
Diffusion-based refinement followed by consistency distillation improves music source separation quality and inference speed across U-Net and BS-RoFormer backbones on Slakh2100 and MUSDB18.
-
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
MaineCoon is presented as the first 22B-parameter real-time streaming audio-visual autoregressive model optimized for social-interactive applications, using novel training techniques and an agentic inference framework.
-
MAGE: Modality-Agnostic Music Generation and Target-Source Extraction
A shared continuous-latent flow model generates music from text/vision or extracts a target source from a mixture via visual-audio alignment, gated modulation, and dynamic modality masking.
-
DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration
DTT-BSR+ is a generative-then-regression cascade for music source restoration that reports MMSNR gains over single-stage DTT-BSR and X-LANCE on most stems while noting a distribution-vs-reconstruction trade-off via FAD.
-
Improving Text-to-Music Generation with Human Preference Rewards
A text-to-music model is improved by conditioning on and selecting with a human preference reward, where expert iteration on top outputs contributes the largest measured gains on 100 Song Describer prompts.