Incorporating multi-modal prompt learning into foundation models enhances predictability of visual fMRI responses to dynamic natural stimuli.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41191973.
- Also identified by DOI 10.1088/1741-2552/ae1bd9.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
<i>Objective</i>. Modeling neural encoding of visual stimuli often uses deep neural networks (DNNs) to predict human brain response to external stimuli. However, each DNN depends on networks tailored for computer vision tasks, resulting in suboptimal brain correspondence. On the other hand, when end-to-end optimizing the encoding process for specific brain regions, challenges like training difficulties arise. Additionally, these models mostly focus on visual information processing, while the human brain integrates multi-modal information such as language to achieve a comprehensive understanding.<i>Approach</i>. To address these limitations, this paper proposes a multi-modal prompt learning (PL) model for neural encoding of dynamic natural stimuli. Specifically, we leverage the powerful representation ability of pre-trained foundation models and fine-tune them using our multi-modal prompts. These prompts, which include textual and visual prompts tailored to each specific regions of interest, can adapt foundation models to neural encoding tasks with fewer trainable parameters. We use the CLIP For video Clip retrieval (CLIP4clip) and Video Masked Autoencoder V2 (videoMAEv2) for feature extraction with backbone freezing, refine the representations via PL, and map the fused multi-modal features to predict voxel-wise brain responses.<i>Main results</i>. Extensive experiments on two functional magnetic resonance imaging video datasets demonstrate that our method outperforms existing fine-tuning methods and public models.<i>Significance</i>. This work highlights the potential of prompt-based fine-tuning strategies in bridging the gap between foundation models and neural encoding tasks.
Medical subject headings
- Magnetic Resonance Imaging
- Photic Stimulation
- Visual Perception
- Neural Networks, Computer
- Brain
- Deep Learning