Pith. sign in

REVIEW 8 cited by

DreamDiffusion: Generating High-Quality Images from Brain EEG Signals

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.16934 v2 pith:NSEV2GGB submitted 2023-06-29 cs.CV

classification cs.CV
keywords dreamdiffusionmethodimagesignalsbrainencodergeneratinghigh-quality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces DreamDiffusion, a novel method for generating high-quality images directly from brain electroencephalogram (EEG) signals, without the need to translate thoughts into text. DreamDiffusion leverages pre-trained text-to-image models and employs temporal masked signal modeling to pre-train the EEG encoder for effective and robust EEG representations. Additionally, the method further leverages the CLIP image encoder to provide extra supervision to better align EEG, text, and image embeddings with limited EEG-image pairs. Overall, the proposed method overcomes the challenges of using EEG signals for image generation, such as noise, limited information, and individual differences, and achieves promising results. Quantitative and qualitative results demonstrate the effectiveness of the proposed method as a significant step towards portable and low-cost ``thoughts-to-image'', with potential applications in neuroscience and computer vision. The code is available here \url{https://github.com/bbaaii/DreamDiffusion}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3D-Telepathy: Reconstructing 3D Objects from EEG Signals

    cs.CV 2025-06 conditional novelty 6.0 of 10

    3D-Telepathy reconstructs 3D objects from EEG signals by combining a dual self-attention EEG encoder with stable diffusion and variational score distillation into a NeRF, and reports best 2D-frame metrics among compar...

  2. Category-aware EEG image generation based on wavelet transform and contrast semantic loss

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A DWT-gated transformer EEG encoder with CLIP alignment and category-aware clustering loss generates semantic images via a pre-trained diffusion model, achieving 43% max single-subject top-1 classification accuracy an...

  3. MoTime: A Dataset Suite for Multimodal Time Series Forecasting

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MoTime provides a large multimodal forecasting benchmark and shows that external text or images can improve forecasts in some datasets, especially cold-start and sparse settings, though gains are inconsistent.

  4. What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

    cs.CV 2026-06 unverdicted novelty 5.5 of 10

    Joint temporal-spectral-spatial EEG encoding with contrastive CLIP alignment sets new SOTA on THINGS-EEG zero-shot decoding, including the first systematic cross-session results.

  5. WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    WorldWeaver reduces temporal drift in long-horizon video generation by jointly modeling RGB and depth perceptual conditions with segmented noise scheduling.

  6. Decoding Visual Neural Representations by Multimodal with Dynamic Balancing

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A multimodal EEG-image-text contrastive framework with dynamic gradient balancing and stochastic noise improves zero-shot object recognition from EEG on ThingsEEG, raising top-1 accuracy from 13.8% to 15.8%.

  7. Foundation Models for Cross-Domain EEG Analysis Application: A Survey

    cs.HC 2025-08 conditional novelty 4.0 of 10

    A survey that organizes EEG foundation-model research into five output-modality categories: native EEG, text, vision, audio, and multimodal fusion, with a claim to be the first such comprehensive taxonomy.

  8. CATVis: Context-Aware Thought Visualization

    cs.CV 2025-07 reject novelty 4.0 of 10

    CATVis combines a Conformer EEG classifier, CLIP-based caption retrieval and re-ranking, and Stable Diffusion to generate images from EEG, reporting large gains over prior work.

Pith tools