RAW-CLIP Fusion: Unleashing Semantic-Aware Denoising for Sensor-Agnostic Low-Light Imaging.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42139123.
- Also identified by DOI 10.1109/TIP.2026.3692025.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Denoising images captured under extreme low-light conditions remains a persistent challenge in computational photography, primarily due to low signal-to-noise ratios and sensor-specific noise characteristics. These variations often require persensor noise calibration to achieve effective denoising. Although recent calibration-free methods aim to reduce this dependency through synthetic noise modeling or few-shot fine-tuning, their performance often degrades in extreme low-light scenarios across different sensors due to mismatches between synthetic and real-world noise. To address this gap, we introduce CLIP-Guided Denoising (CLD), the first framework to leverage large-scale vision models pretrained on sRGB images for cross-domain feature fusion, effectively guiding RAW image denoising across diverse sensors. Although not trained on RAW data, CLIP embeddings offer semantically robust and noise-invariant features that help guide the denoising network to focus on the underlying image content rather than fitting to specific noise distributions. Extensive experiments on the SID and ELD datasets demonstrate that CLD achieves state-of-the-art performance in calibration-free settings, significantly outperforming prior methods under extreme low-light conditions and achieving robust generalization across unseen sensor domains.