Memory like the human brain: A framework for decoding multimodal learning of brain-visual-linguistic features.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41833171.
- Also identified by DOI 10.1016/j.media.2026.104038.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Decoding human visual neural representations is scientifically important for advancing research on brain-like intelligence. Existing research typically aligns neural signals captured by fMRI or EEG with visual and linguistic features, thereby enabling models to decode brain activity into unseen visual categories. However, current methods still suffer from two fundamental challenges: (1) Representation drift. The alignment between learned features and new patterns degrades during continuous training, unlike the stable retention of knowledge in the human brain. (2) The incomplete modeling of common and individual representations. The model frequently fails to effectively disentangle common semantics from modality-specific information within image and text pairs. To address these issues, we propose a novel framework named MLHuB that mimics the human brain's learning mechanism. Firstly, we propose a memory unit responsible for reading and updating the learned text-image features as a way to consolidate acquired knowledge. Secondly, we compute common and individual features between text and images via orthogonal projection and utilize intra-modality mutual information maximization to regularize the learning of text-image pairs, encouraging the model to explore unseen knowledge. Finally, we integrate both intra- and inter-modality mutual information maximization to learn a more consistent joint representation between modalities. Extensive experiments on three datasets demonstrate that our proposed MLHuB achieves state-of-the-art performance.