Pith. sign in

REVIEW 7 cited by

Visual Decoding and Reconstruction via EEG Embeddings with Guided Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07721 v7 pith:7HCSR3G5 submitted 2024-03-12 cs.HC eess.SPq-bio.NC

classification cs.HCeess.SPq-bio.NC
keywords reconstructionvisualdecodingembeddingimageclipdiffusionframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How to decode human vision through neural signals has attracted a long-standing interest in neuroscience and machine learning. Modern contrastive learning and generative models improved the performance of visual decoding and reconstruction based on functional Magnetic Resonance Imaging (fMRI). However, the high cost and low temporal resolution of fMRI limit their applications in brain-computer interfaces (BCIs), prompting a high need for visual decoding based on electroencephalography (EEG). In this study, we present an end-to-end EEG-based visual reconstruction zero-shot framework, consisting of a tailored brain encoder, called the Adaptive Thinking Mapper (ATM), which projects neural signals from different sources into the shared subspace as the clip embedding, and a two-stage multi-pipe EEG-to-image generation strategy. In stage one, EEG is embedded to align the high-level clip embedding, and then the prior diffusion model refines EEG embedding into image priors. A blurry image also decoded from EEG for maintaining the low-level feature. In stage two, we input both the high-level clip embedding, the blurry image and caption from EEG latent to a pre-trained diffusion model. Furthermore, we analyzed the impacts of different time windows and brain regions on decoding and reconstruction. The versatility of our framework is demonstrated in the magnetoencephalogram (MEG) data modality. The experimental results indicate that our EEG-based visual zero-shot framework achieves SOTA performance in classification, retrieval and reconstruction, highlighting the portability, low cost, and high temporal resolution of EEG, enabling a wide range of BCI applications. Our code is available at https://github.com/ncclab-sustech/EEG_Image_decode.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding

    cs.CL 2026-02 unverdicted novelty 6.0 of 10

    SemKey predicts four semantic attributes from EEG and conditions a frozen LLM on them, beating prior decoders on new semantic-alignment metrics while leaving true word-level accuracy low (2.7% content recall).

  2. EEG2TEXT-CN: An Exploratory Study of Open-Vocabulary Chinese Text-EEG Alignment via Large Language Model and Contrastive Learning on ChineseEEG

    cs.CL 2025-06 reject novelty 6.0 of 10

    A Chinese EEG-to-text system that aligns 128-channel EEG with per-character text embeddings achieves BLEU-1 6.38% on a held-out subject, claimed as the first open-vocabulary EEG-to-Chinese decoder.

  3. Category-aware EEG image generation based on wavelet transform and contrast semantic loss

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A DWT-gated transformer EEG encoder with CLIP alignment and category-aware clustering loss generates semantic images via a pre-trained diffusion model, achieving 43% max single-subject top-1 classification accuracy an...

  4. DynaMind: Reconstructing Dynamic Visual Scenes from EEG by Aligning Temporal Dynamics and Multimodal Semantics to Guided Diffusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    DynaMind reconstructs videos from EEG by combining region-aware semantic mapping, a temporal blueprint, and dual-guidance diffusion, outperforming EEG2Video on SEED-DV in most comparisons.

  5. Transformer-based EEG Decoding: A Survey

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A survey that classifies Transformer-based EEG decoding models into backbone, hybrid, and customized categories and reviews their applications and limitations.

  6. Uncovering the EEG Temporal Representation of Low-dimensional Object Properties

    cs.HC 2025-07 conditional novelty 4.0 of 10

    Using a pre-trained EEG decoder and temporal masking, the authors find concept-specific activation windows and prototypical temporal clusters in THINGS-EEG data.

  7. Bridging Brain with Foundation Models through Self-Supervised Learning

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A PRISMA-based survey maps self-supervised learning techniques, brain foundation models, datasets, and evaluation protocols for EEG and related neural signals, including a skeptical review of EEG-to-text decoding.

Pith tools