Cognitive prototype learning: Towards self-supervised and semantic-aware few-shot open-set sound recognition.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42019217.
- Also identified by DOI 10.1016/j.neunet.2026.109008.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Environmental sound classification based on deep learning has achieved remarkable success, but current models often face challenges such as data scarcity and category openness. Humans recognize novel sounds with limited samples by leveraging pattern discovery and semantic reasoning, but computational models often lack such cognitive capabilities. To bridge this gap, we propose a cognitive prototype learning framework that integrates self-supervised learning with semantic knowledge guidance for few-shot open-set recognition (FSOR). It comprises three key modules: (1) a self-supervised perception module that learns general acoustic patterns through contrastive learning; (2) a semantic context module that derives semantic embeddings from class names using a pretrained model; and (3) a semantic prototype generation module that produces compact and semantically consistent prototypes by dynamically fusing acoustic and semantic features. The synergy of these three modules mitigates overfitting of known class, reduces semantic noise, and enhances prototype compactness. Experiments on ESC-50 dataset show that the proposed method achieves 64.50% accuracy and 66.29% AUROC under 5-way 1-shot, and 78.36% accuracy and 75.09% AUROC under 5-way 5-shot settings, outperforming representative FSOR methods in both closed-set classification and open-set detection. These results demonstrate the effectiveness of the proposed framework in bridging the gap between machine intelligence and human cognitive reasoning.