Shengyu Gong, Weiming Zeng, Yueyang Li +4 more
Brain-computer interfaces that decode what someone is looking at perform well on tidy laboratory stimuli and fall apart on real photographs. The gap is the whole problem, since nobody wants an interface that only works on simplified images.
The authors trace the failure to what the standard training objective optimises. Multimodal contrastive learning aligns representations by geometric distance, pulling matching pairs together in a shared space. That says nothing about whether the alignment is semantically coherent, and it treats every person's brain as though it encodes the same thing the same way.
Both omissions matter here more than in most applications. People differ in how they represent what they see and in what they attend to within a scene, so an objective assuming otherwise is fitting an average nobody actually is. Making the model subject-aware is an acknowledgement that inter-subject variability is signal, not noise to be averaged away.
Non-invasive brain-computer interfaces exhibit significant performance degradation when moving from controlled laboratory stimuli to real-world natural images. This degradation occurs because conventional multimodal contrastive representation learning models focus exclusively on optimizing geometric distance alignment, thereby failing to account for semantic consistency and inter-subject variability in neural representation and selective attention. As a result, these models…
What Holds Back Brain-Computer Interfaces? Uncovering Challenges and Opportunities in BCI-controlled Games for Cerebral Palsy Rehabilitation
arXiv (BCI) · June 24, 2026Towards Robust EEG Decoding Based on Riemannian Self-Attention
arXiv (BCI) · June 24, 2026BrainAgent: A Large Language Model-Driven Multi-Agent Framework for Autonomous Brain Signal Understanding
arXiv (EEG) · June 23, 2026EEG Interpretation Across Chant Listening: A Single-Subject Pilot Investigation Using Spectral and Functional Connectivity Analysis