Pith. sign in

REVIEW 11 cited by

MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11207 v2 pith:YERLA64Q submitted 2024-03-17 cs.CV cs.AIq-bio.NC

classification cs.CVcs.AIq-bio.NC
keywords dataspaceclipreconstructionssubjecttrainingbrainfmri
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reconstructions of visual perception from brain activity have improved tremendously, but the practical utility of such methods has been limited. This is because such models are trained independently per subject where each subject requires dozens of hours of expensive fMRI training data to attain high-quality results. The present work showcases high-quality reconstructions using only 1 hour of fMRI training data. We pretrain our model across 7 subjects and then fine-tune on minimal data from a new subject. Our novel functional alignment procedure linearly maps all brain data to a shared-subject latent space, followed by a shared non-linear mapping to CLIP image space. We then map from CLIP space to pixel space by fine-tuning Stable Diffusion XL to accept CLIP latents as inputs instead of text. This approach improves out-of-subject generalization with limited training data and also attains state-of-the-art image retrieval and reconstruction metrics compared to single-subject approaches. MindEye2 demonstrates how accurate reconstructions of perception are possible from a single visit to the MRI facility. All code is available on GitHub.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fast Whole-Brain, Geometry-Aware Functional Alignment for Cross-Subject Decoding

    q-bio.NC 2026-07 conditional novelty 6.5 of 10

    SpectralOT regularizes entropic optimal transport with the first three Laplace-Beltrami eigenmodes of cortical geometry to produce fast, parsimonious whole-brain functional alignments that improve cross-subject decoding.

  2. Real-time Reconstruction of Human Visual Perception from fMRI

    cs.CV 2026-07 conditional novelty 6.0 of 10

    First demonstration that single-trial visual images can be decoded from fMRI in near-real-time (about 10-15 seconds) with roughly one hour of fine-tuning data.

  3. Predicted Cortex Is Not a Domain-General Prior: A Matched-Control Audit of Brain-Encoding Features for Video Memorability

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Predicted cortical responses from a brain-encoding model beat their own visual backbone on VideoMem but lose on Memento10k, so they are a dataset-specific memorability representation, not a domain-general prior.

  4. BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A new automated pipeline decomposes fMRI activity into components and labels them with visual concepts, claiming thousands of interpretable patterns across the human visual cortex.

  5. Hi-DREAM: Brain-Inspired Hierarchical Diffusion for fMRI-to-Image Reconstruction via ROI Encoder and VisuAl Mapping

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Hi-DREAM conditions a latent diffusion U-Net on early/mid/late visual-ROI streams via a multi-scale cortical pyramid and depth-matched ControlNet, reporting state-of-the-art semantic metrics on NSD fMRI-to-image recon...

  6. MindShot: Multi-Shot Video Reconstruction from fMRI with LLM Decoding

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    MindShot reconstructs multiple video shots from fMRI by first predicting shot boundaries, then using LLM-generated captions of keyframes to decode visual content.

  7. Optimizing fMRI Data Acquisition for Decoding Natural Speech with Limited Participants

    q-bio.NC 2025-05 conditional novelty 6.0 of 10

    In a small cohort, fMRI decoders improve with more data per participant, and multi-subject training or shared stimuli add no benefit, so deep phenotyping is the recommended acquisition strategy.

  8. Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single-stage, LoRA-finetuned diffusion model decodes seen images directly from continuous BOLD fMRI time series and beats previous pipelines on semantic metrics.

  9. VoxelFormer: Parameter-Efficient Multi-Subject Visual Decoding from fMRI

    cs.CV 2025-09 reject novelty 5.0 of 10

    A 39M-parameter transformer with token merging and query compression gives around 74% top-1 fMRI-to-image retrieval across six training subjects, far below the ~98% of larger state-of-the-art decoders.

  10. Threat Vectors and the State of the Art in Defense Methods for Security in Neurotechnology

    cs.CR 2026-07 conditional novelty 4.5 of 10

    Neurosecurity lags BCI capability; a full-stack attack-surface taxonomy plus transferable defenses from cyber, hardware security, and private ML can close many gaps immediately.

  11. Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review

    cs.AI 2025-10 conditional novelty 1.0 of 10

    This paper is a survey: it organizes existing foundation-model work in neuroscience into five application domains and lists public datasets, without presenting new experiments.

Pith tools