Vision Transformer and NeoPulseNet: A Dual Approach for Accurate rPPG Signal Extraction in Neonates.

Anil, Aravind A; Karthik, Srinivasa; Sivaprakasam, Mohanasankar; Joseph, Jayaraj · IEEE J Biomed Health Inform · 2025

basic_science · Level V

Where this comes from

Abstract

Non-contact heart rate (HR) monitoring via camera offers a safer alternative to traditional wired methods in neonates. The first step in this process is accurate segmentation of skin pixels on the neonate's face, which poses challenges due to interference from caregivers' skin. To address this, we employed a vision transformer trained specifically on our neonatal dataset. Following skin segmentation, HR was extracted using two approaches: (1) traditional rPPG algorithms (POS, CHROM, ICA, LGI, PBV) and (2) a deep learning model trained on surrogate ground truth derived from these algorithms. For this purpose, we developed a novel architecture, NeoPulseNet, which integrates 1D-CNNs and transformer layers. The results demonstrated that surrogate training yielded a lower mean absolute error (MAE) compared to the classical methods. The best MAE of 7.85 bpm was achieved by selecting the optimal surrogate ground-truth rPPG for each video segment, rather than applying a single algorithm uniformly across all segments. NeoPulseNet was further validated under varying neonatal conditions, including differences in skin melanin, head position relative to the camera, lighting variations, and motion. Across these conditions, it consistently produced MAE values around or below 10 bpm, which is considered clinically acceptable. Finally, computational efficiency was assessed on low-end hardware using 5-second video segments. NeoPulseNet achieved a processing time of 13 ms, comparable to classical rPPG algorithms and significantly faster than state-of-the-art deep learning methods, which required nearly 1 second.