CVJun 24

What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

arXiv:2606.257186.1
Predicted impact top 75% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in brain-computer interfaces and neural decoding, this work improves semantic alignment and generalization in EEG-based visual decoding by explicitly modeling multiview neural structure.

The paper introduces a multiview EEG representation learning framework for zero-shot visual decoding from EEG, achieving state-of-the-art results: 54.8% Top-1 and 85.6% Top-5 accuracy in within-subject, 15.3% Top-1 and 45.4% Top-5 in cross-subject, and 40.8% Top-1 and 78.0% Top-5 in cross-session settings on THINGS-EEG.

Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging due to the low signal-to-noise ratio, non-stationarity, and limited spatial resolution of EEG. Existing EEG-vision alignment methods often rely on holistic EEG embeddings, which can obscure the complementary temporal, spectral, and spatial structure underlying visual perception. We introduce a unified multiview EEG representation learning framework for aligning brain responses with visual semantic embeddings. Our method builds an EEG encoder that jointly models three complementary views: input-conditioned state-space temporal dynamics, learnable wavelet-based spectral decomposition for sample-adaptive frequency modeling, and attention-modulated graph learning for structured electrode interactions. The resulting multiview EEG embeddings are fused and aligned with pretrained visual representations in a shared semantic space using contrastive learning with EEG-specific regularization, enabling 200-way zero-shot visual classification. Experiments on THINGS-EEG benchmark show that our method achieves state-of-the-art performance, with 54.8% Top-1 and 85.6% Top-5 accuracy in the within-subject setting and 15.3% Top-1 and 45.4% Top-5 accuracy in the cross-subject setting. We further present the first systematic cross-session EEG-image decoding evaluation, achieving 40.8% Top-1 and 78.0% Top-5 accuracy. These results suggest that explicitly modeling multiview neural structure improves both semantic alignment and generalization in EEG-based visual decoding.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes