REVIEW 8 cited by
DreamDiffusion: Generating High-Quality Images from Brain EEG Signals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces DreamDiffusion, a novel method for generating high-quality images directly from brain electroencephalogram (EEG) signals, without the need to translate thoughts into text. DreamDiffusion leverages pre-trained text-to-image models and employs temporal masked signal modeling to pre-train the EEG encoder for effective and robust EEG representations. Additionally, the method further leverages the CLIP image encoder to provide extra supervision to better align EEG, text, and image embeddings with limited EEG-image pairs. Overall, the proposed method overcomes the challenges of using EEG signals for image generation, such as noise, limited information, and individual differences, and achieves promising results. Quantitative and qualitative results demonstrate the effectiveness of the proposed method as a significant step towards portable and low-cost ``thoughts-to-image'', with potential applications in neuroscience and computer vision. The code is available here \url{https://github.com/bbaaii/DreamDiffusion}.
Forward citations
Cited by 8 Pith papers
-
3D-Telepathy: Reconstructing 3D Objects from EEG Signals
3D-Telepathy reconstructs 3D objects from EEG signals by combining a dual self-attention EEG encoder with stable diffusion and variational score distillation into a NeRF, and reports best 2D-frame metrics among compar...
-
Category-aware EEG image generation based on wavelet transform and contrast semantic loss
A DWT-gated transformer EEG encoder with CLIP alignment and category-aware clustering loss generates semantic images via a pre-trained diffusion model, achieving 43% max single-subject top-1 classification accuracy an...
-
MoTime: A Dataset Suite for Multimodal Time Series Forecasting
MoTime provides a large multimodal forecasting benchmark and shows that external text or images can improve forecasts in some datasets, especially cold-start and sparse settings, though gains are inconsistent.
-
What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment
Joint temporal-spectral-spatial EEG encoding with contrastive CLIP alignment sets new SOTA on THINGS-EEG zero-shot decoding, including the first systematic cross-session results.
-
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
WorldWeaver reduces temporal drift in long-horizon video generation by jointly modeling RGB and depth perceptual conditions with segmented noise scheduling.
-
Decoding Visual Neural Representations by Multimodal with Dynamic Balancing
A multimodal EEG-image-text contrastive framework with dynamic gradient balancing and stochastic noise improves zero-shot object recognition from EEG on ThingsEEG, raising top-1 accuracy from 13.8% to 15.8%.
-
Foundation Models for Cross-Domain EEG Analysis Application: A Survey
A survey that organizes EEG foundation-model research into five output-modality categories: native EEG, text, vision, audio, and multimodal fusion, with a claim to be the first such comprehensive taxonomy.
-
CATVis: Context-Aware Thought Visualization
CATVis combines a Conformer EEG classifier, CLIP-based caption retrieval and re-ranking, and Stable Diffusion to generate images from EEG, reporting large gains over prior work.
Discussion (0). Continue with ORCID to comment.