Towards robust remote sensing visual question answering with spectral expert adaptation and group-relative optimization.

Yuan, Junjiang; Li, Zhe; Zhang, Tingbin · Neural Netw · 2026

Where this comes from

Abstract

Transformer-based vision-language large language models (V-L LLMs) have advanced multimodal understanding on natural-image benchmarks, but their use in remote sensing (RS) visual question answering (VQA) is still limited by domain shift, high adaptation cost, and the need to reason over spatially organized scenes. We present SpectralLoRA-R1 for RS-VQA. The method has two parts. First, mixture-of-experts SpectralLoRA (MoE-SpectralLoRA) adapts frozen pretrained weights in a fixed spectral basis: low-rank residuals are parameterized by coefficient matrices between the leading left and right singular vectors of each pretrained weight matrix. Several SpectralLoRA experts are combined through a lightweight MoE router, allowing different image-question pairs to use different spectral coefficient mixtures. Second, we apply generalized reinforcement policy optimization (GRPO) with verifiable rewards to tune the model's answer format, answer consistency, and final-answer correctness. Across multiple RS-VQA benchmarks, SpectralLoRA-R1 improves over previous methods while keeping the backbone frozen. The results indicate that fixed-basis spectral adaptation and verifier-guided post-training can be combined to improve RS vision-language understanding without full model fine-tuning.