Efficient fine-tuning of large-scale vision-language models for visual marketing analysis: From brand logo detection to aesthetic preference prediction.

Qin, Suxiang; Wei, Songwen · PLoS One · 2026

basic_science · Level V

Where this comes from

Abstract

Visual marketing analysis has emerged as a critical research domain at the intersection of computer vision, natural language processing, and consumer behavior modeling. This study addresses three fundamental challenges in this field: the accurate and robust detection of brand logos in complex commercial visual scenes, the construction of a unified model for understanding both visual content and accompanying text through fine-grained vision-language alignment, and the quantification of audience aesthetic preferences for data-driven marketing effectiveness prediction. We propose Brand-Aesthetic Vision-Language Assistant (BAVLA), a novel framework comprising Multi-granularity Context-aware Brand Detection and Fusion Module (MCBF), Aesthetic-aware Vision-Language Alignment and Reasoning Module (AVLR), and Task-aware Progressive Efficient Tuning strategy (TaPET). Compared to existing methods, the proposed MCBF module improves logo detection mAP by 6.2% and 2.5% over Faster R-CNN and YOLOv8, respectively, on the Flickr Logo-27 dataset. Furthermore, the complete BAVLA framework achieves superior aesthetic prediction performance, surpassing previous best methods (NIMA, AestheticCNN) by 0.084 and 0.055 in PLCC, respectively, on the AVA dataset, while attaining 82.4% accuracy in marketing effectiveness classification. These findings validate the effectiveness of the proposed modules and training strategy in advancing visual marketing analysis capabilities.

Medical subject headings