A pipeline for enabling Nearshore Infrared Video Super-resolution to learn more high-frequency foreground information.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40349427.
- Also identified by DOI 10.1016/j.neunet.2025.107547.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
A key challenge in Nearshore Infrared Video Super-resolution (NIVSR) is the limited high-frequency foreground information. The most common approach is to fuse frames in order to learn cross-temporal information. However, existing methods struggle to achieve pixel-level reconstruction of foreground features in infrared video with limited detail. This factor is further amplified due to the transformation of the image into patches in the Super-Resolution (SR) process. This paper presents a novel spatial and temporal network, TASNet, designed to improve reconstruction quality. TASNet models the video in terms of both spatial and temporal features, facilitating their interaction. The Efficient Foreground Information Perception (EFIP) module leverages feature variations to emphasize foreground information in the current frame. Temporal-Difference Learning (TDL) learns information from different frames and integrates it using learnable weights. Additionally, a strategy utilizing the long-context comprehension of Visual Transformers (ViT) is introduced to mitigate temporal discrepancies between frames. The method is simple, robust, and surpasses State-of-the-art (SOTA) techniques in benchmark experiments (TASNet: 28.33 Peak Signal-to-Noise Ratio (PSNR), 0.9122 Structural Similarity Index Measure (SSIM); RBPN: 27.27 PSNR, 0.9024 (SSIM). The source code is in the https://github.com/Yuanlin-Zhao/TASNet.
Medical subject headings
- Infrared Rays
- Video Recording
- Neural Networks, Computer
- Image Processing, Computer-Assisted
- Machine Learning