Pith. sign in

REVIEW 2 cited by

Naturalistic Music Decoding from EEG Data via Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09062 v6 pith:Y6HHOQRO submitted 2024-05-15 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords musicdatamodelsdecodingdiffusionlatentnaturalisticneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this article, we explore the potential of using latent diffusion models, a family of powerful generative models, for the task of reconstructing naturalistic music from electroencephalogram (EEG) recordings. Unlike simpler music with limited timbres, such as MIDI-generated tunes or monophonic pieces, the focus here is on intricate music featuring a diverse array of instruments, voices, and effects, rich in harmonics and timbre. This study represents an initial foray into achieving general music reconstruction of high-quality using non-invasive EEG data, employing an end-to-end training approach directly on raw data without the need for manual pre-processing and channel selection. We train our models on the public NMED-T dataset and perform quantitative evaluation proposing neural embedding-based metrics. Our work contributes to the ongoing research in neural decoding and brain-computer interfaces, offering insights into the feasibility of using EEG data for complex auditory information reconstruction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decoding Speaker-Normalized Pitch from EEG for Mandarin Perception

    cs.SD 2025-05 conditional novelty 6.0 of 10

    EEG decoding of Mandarin pitch is more accurate for speaker-normalized than raw pitch across multiple speakers, suggesting the brain encodes relative, speaker-independent pitch.

  2. Predicting Artificial Neural Network Representations to Learn Recognition Model for Music Identification from Brain Recordings

    q-bio.NC 2024-12 conditional novelty 5.0 of 10

    Training EEG encoders with an auxiliary InfoNCE loss that predicts a co-trained music encoder's representation improves 10-song EEG identification accuracy on the NMED-T dataset.

Pith tools