Byeongseo Bok, Futa Waseda, Jun Liu +1 more
Brain decoding often works by matching fMRI responses to the internal representations of an AI model, and CLIP, with its joint image-and-text space, has become a favorite target. But CLIP was never built to line up with brains, so some of what it encodes is shortcut features that don't correspond to anything the brain does, which caps the alignment.
The question here is simple: would an adversarially robust version of CLIP align better? Robust training strips out brittle shortcut features and keeps more perceptually meaningful ones, which might sit closer to real neural activity. Testing this by swapping only the target representation, on two fMRI datasets, the robust variants consistently improved image retrieval and zero-shot classification. Attribution analysis showed robust and standard models weight features quite differently, suggesting robustness reorganizes what the representation cares about.
Based on the abstract, so the paper will have the datasets and metrics spelled out.
Brain decoding aims to uncover neural mechanisms by inferring stimulus-related representations from brain signals. In fMRI studies, this is typically achieved by mapping fMRI responses to the latent representations of computational models. Recently, CLIP has become a popular choice for brain decoding due to its rich vision--language embedding space. However, aligning fMRI signals with CLIP representations remains challenging. As CLIP is not explicitly optimized for neural…
SpindleFlexNet: Flexible sleep spindles detection for EEG signals based on an adaptive one-dimensional RetinaNet-based framework
arXiv (BCI) · June 18, 2026B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet
arXiv (EEG) · June 18, 2026Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU
arXiv (BCI) · June 17, 2026SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface