Towards robust remote sensing visual question answering with spectral expert adaptation and group-relative optimization.
Where this comes from
- Record sourced from PubMed, PMID 42424798.
- Also identified by DOI 10.1016/j.neunet.2026.109308.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Transformer-based vision-language large language models (V-L LLMs) have advanced multimodal understanding on natural-image benchmarks, but their use in remote sensing (RS) visual question answering (VQA) is still limited by domain shift, high adaptation cost, and the need to reason over spatially organized scenes. We present SpectralLoRA-R1 for RS-VQA. The method has two parts. First, mixture-of-experts SpectralLoRA (MoE-SpectralLoRA) adapts frozen pretrained weights in a fixed spectral basis: low-rank residuals are parameterized by coefficient matrices between the leading left and right singular vectors of each pretrained weight matrix. Several SpectralLoRA experts are combined through a lightweight MoE router, allowing different image-question pairs to use different spectral coefficient mixtures. Second, we apply generalized reinforcement policy optimization (GRPO) with verifiable rewards to tune the model's answer format, answer consistency, and final-answer correctness. Across multiple RS-VQA benchmarks, SpectralLoRA-R1 improves over previous methods while keeping the backbone frozen. The results indicate that fixed-basis spectral adaptation and verifier-guided post-training can be combined to improve RS vision-language understanding without full model fine-tuning.