July 06, 2026
Keywords: Apnoea of prematurity; neonatal intensive care; video-based respiratory monitoring; non-contact monitoring; machine learning; multimodal data fusion.
Pre-term birth, defined as delivery before 37 weeks of gestation, remains a leading cause of neonatal morbidity and mortality worldwide [1]–[3]. Respiratory complications such as apnoea, driven by immature respiratory control and lung development, are major contributors to mortality in this population. Apnoea of prematurity (AOP) affects up to 85% of infants born at or before 34 weeks of gestation [4], [5]. Episodes of cessation of breathing (COBE) are a defining feature of AOP and may result in hypoxaemia (reduced oxygen levels in the blood), bradycardia, neurological injury, and long-term developmental complications, including cognitive impairment and motor dysfunction [6], [7]. Accurate and reliable detection of COBE episodes is therefore critical for timely clinical intervention and improved outcomes.
Although polysomnography remains the gold standard for diagnosing apnoea [8], [9], it is impractical for continuous monitoring in the neonatal intensive care unit (NICU) due to its complexity, requirement for specialised equipment, and need for trained personnel. Routine respiratory monitoring and apnoea detection in the NICU rely primarily on contact-based techniques, including impedance pneumography (IP) and pulse oximetry [10], [11]. Impedance pneumography uses chest electrodes to measure changes in transthoracic electrical impedance that occur due to cyclic changes in lung air volume during breathing [12], [13]. These impedance variations provide an indirect measure of respiratory effort and are commonly used to estimate respiratory rate (RR) and identify respiratory pauses. Pulse oximetry is an optical technique used to estimate peripheral oxygen saturation (\(SpO_{2}\)) and identify desaturation events associated with COBE episodes [14].
Despite their widespread use, contact-based monitoring systems are subject to several limitations. Motion artefacts, electrode displacement, and cardiac-related fluctuations in impedance signals can generate false alarms or obscure respiratory pauses [12], [14], [15]. The attachment and removal of adhesive sensors may also contribute to skin irritation, discomfort, and epidermal injury in pre-term infants with immature and highly sensitive skin [16], [17]. Contact sensors can also interfere with routine caregiving practices such as kangaroo care, which promotes physiological stability and bonding [18], [19]. These constraints are particularly pronounced in low-resource settings, where limited monitoring infrastructure and a shortage of clinical staff may delay timely recognition of respiratory events [20], [21]. Collectively, these challenges highlight the need for monitoring approaches that enhance robustness while reducing physical burden.
Non-contact physiological monitoring has emerged as a promising alternative approach. Technologies such as thermal imaging, Doppler radar, WiFi-based sensing, and video camera-based monitoring can capture respiratory-related signals without direct skin contact [22]–[24]. Thermal cameras detect respiration-induced temperature fluctuations near the airway [25]. Doppler radar systems estimate respiration by tracking chest or abdominal wall movements during breathing [26]. WiFi-based methods infer respiration from breathing-induced variations in wireless signal propagation characteristics [27]. While these modalities reduce sensor-related discomfort, they may be costly, susceptible to environmental interference, and remain insufficiently validated in neonatal populations [28], [29].
Visible-spectrum (RGB) video cameras offer a comparatively accessible and scalable approach. Subtle thoracic or abdominal motion can be extracted through image processing and machine learning techniques [30]–[32]. In adult cohorts, video-based methods have achieved high sensitivity for respiratory event detection under controlled conditions [32]–[35]. However, translating video-based respiratory monitoring to pre-term infants introduces distinct technical challenges. Neonatal respiratory patterns are more irregular than in adults, with breath-to-breath variability, making RR estimation more challenging [36], [37]. Cardiac activity can also introduce artefacts into respiratory signals by superimposing cardiac-induced fluctuations, reducing the accuracy of respiratory signal extraction [38], [39]. Unlike adult sleep studies, AOP may occur across sleep-wake states, including wakefulness and rapid eye movement (REM) sleep, requiring respiratory monitoring across a broader range of physiological states [40]–[43]. These physiological challenges are compounded by the NICU environment, where infants are frequently repositioned or partially occluded by bedding or caregivers, and incubator structures can introduce reflections and visual interference [44]. Variable lighting and spontaneous limb movements further complicate feature extraction, requiring methods that are robust to real-world NICU conditions. Most prior neonatal studies rely on small sample sizes, observer-dependent annotations, or controlled laboratory settings, limiting generalisability to routine clinical environments [22], [39].
Contact-based sensors quantify physiological parameters such as RR and \(SpO_{2}\) but do not capture visual context such as posture changes, caregiver interactions, or transient motion patterns. Vision-based monitoring may provide complementary information by capturing surface motion cues that help distinguish respiratory pauses from artefactual fluctuations in physiological signals. Integrating contact and non-contact modalities may therefore improve the robustness of COBE event detection by combining physiological measurements with contextual visual information. However, the extent to which video-derived information provides complementary value beyond conventional physiological monitoring for COBE detection remains unclear.
In this study, we propose video camera-based and hybrid frameworks for COBE detection in pre-term infants. Using synchronised video and physiological recordings acquired under routine NICU conditions, we evaluate whether video-derived signals contain clinically relevant information for COBE detection and whether their integration with conventional physiological measurements improves detection performance.
We used data from a clinical study of pre-term infants admitted to the NICU at the John Radcliffe Hospital, Oxford, UK. The study was conducted in a collaboration between Oxford University Hospitals NHS Foundation Trust and the Oxford Biomedical Research Centre (BRC), with approval from the South Central - Oxford Research Ethics Committee (13/SC/0597). Written informed consent was obtained from parents prior to participation.
Infants born at less than 37 weeks gestation and requiring high-dependency care were eligible for inclusion. The cohort comprised 30 pre-term infants (mean gestational age 31.1 \(\pm\) 1.8 weeks) monitored for up to 7 consecutive days during daytime under routine clinical conditions. There were a total of across 90 recording sessions, yielding approximately 426 hours of synchronised video and physiological recordings. A 3-CCD JAI AT-200CL digital camera (JAI A/S, Denmark) was positioned through a modified opening in the canopy of a Giraffe OmniBed Carestation incubator (General Electric, Connecticut, USA) (figure 1a). The camera captured continuous 24-bit RGB video at 20 frames per second with a resolution of 1628 \(\times\) 1236 pixels. A representative video frame is shown in figure 1b. Standard physiological signals, including heart rate (HR), respiratory rate (RR), and \(SpO_2\), were recorded using a Philips IntelliVue MX800 patient monitor (Philips, Amsterdam, Netherlands) shown in figure 1c. The monitor was equipped with a Masimo SET \(SpO_2\) Vuelink IntelliVue measurement module (Masimo, California, USA) for oxygen saturation monitoring . Video acquisition was temporarily paused during selected clinical procedures, including phototherapy, cannula insertion, or kangaroo care. The full cohort design and data acquisition protocol have been described previously in Villarroel et al. [31].
Previous analysis on our dataset demonstrated substantial discrepancies between monitor-provided respiratory rate (\(RR_{philips}\)) and manual breath counting, indicating that the monitor values were insufficiently reliable for use as a reference signal [45]. In addition, the Philips IntelliVue MX800 derives respiratory rate using proprietary signal-processing algorithms and does not provide a corresponding measure of signal quality [46]. While the processing pipeline and averaging window are not publicly documented, monitor-provided RR estimates exhibited reduced short-term variability (figure 2b). Such behaviour is consistent with temporal averaging approaches commonly used in bedside monitoring systems to provide stable respiratory measurements. These approaches may reduce sensitivity to transient respiratory fluctuations and short respiratory pauses such as those during COBE episodes [37], [47]. Therefore, respiratory rate was recomputed directly from the raw IP waveform (\(RR_{ip}\)).
To remove baseline drift and high-frequency noise, the IP signal was bandpass filtered using an 8th-order high-pass Butterworth filter (0.033 Hz) and a 6th-order low-pass Butterworth filter (2.83 Hz), preserving respiratory frequencies corresponding to approximately 2–170 breaths/min. Respiratory cycles were then identified using the Mean Average Curve (MAC) peak detector [48], and respiratory rate was estimated using breath counting within a 10-second sliding window with a 1-second step size. Signal quality assessment was applied to individual respiratory cycles based on physiological plausibility, spectral characteristics, and waveform purity. Respiratory cycles that failed these quality criteria were excluded from RR estimation. Figure 2 shows RR estimation over a 5-minute segment containing a period of cessation of breathing. Full details of the RR recomputation procedure are provided in Appendix 6.
COBE labels were derived from physiological recordings according to established neonatal clinical criteria for apnoea [40], [49]. A COBE event was defined as either (i) a respiratory pause (\(RR_{ip}\) < 20 breaths/min) lasting at least 20 seconds, or (ii) a shorter respiratory pause of at least 10 seconds accompanied by bradycardia (HR < 100 beats/min) or oxygen desaturation (\(SpO_2\) < 80% lasting at least 10 seconds).
Potential candidate COBE events were identified through automated screening of periods where oxygen saturation (\(SpO_{2}\)) remained below 80% for at least 10 seconds. Consecutive desaturation episodes occurring within 20 seconds of each other were merged to avoid splitting prolonged physiological events into multiple candidate segments because of brief transient recoveries in oxygen saturation. This approach is consistent with observations that apnoea-related events may occur in temporal clusters during periods of physiological instability [50]. This procedure identified 620 candidate desaturation-associated segments. To ensure representation of COBE events occurring both with and without oxygen desaturation, an additional 620 non-desaturation segments were randomly sampled from the same recording sessions for manual review.
Each candidate segment was independently reviewed by three annotators (one clinician and two biomedical engineers) using \(RR_{ip}\), \(SpO_2\), ECG, PPG, and IP waveforms. The reviewers examined 5-minute contextual windows centred on each segment to determine whether a COBE episode had occurred. The decision workflow used to standardise annotation is shown in Appendix 7.
Confirmed COBE events were extracted as 80-second segments comprising 60 seconds preceding and 20 seconds following the onset of the annotated event, providing temporal context before and after the event. Normal breathing events were extracted as 80-second segments centred on the midpoint of the corresponding 5-minute windows labelled as normal breathing. Segments containing missing physiological data resulting from sensor disconnection, or for which the three reviewers did not agree, were excluded from further analysis. The final dataset comprised 346 COBE segments and 608 normal breathing segments derived from 24 infants. Inter-rater reliability was substantial for desaturation-associated events (Fleiss’ \(\kappa = 0.80\)) and moderate for non-desaturation events (\(\kappa = 0.57\)), reflecting the greater ambiguity in identifying respiratory pauses in the absence of associated desaturation.
The annotated 80-second segments were processed using an automated region-of-interest (ROI) detection pipeline based on MediaPipe Pose, which uses the BlazePose pose-estimation model [51]. Although BlazePose is primarily trained on adult human data, it was applied to estimate anatomical landmarks in neonatal recordings for dynamic ROI localisation.
BlazePose identifies 33 anatomical landmarks corresponding to major body joints. For ROI definition, only the left and right shoulders and hips were utilised. These landmarks were used to define two regions of interest, illustrated in figure 3: (i) a torso ROI capturing global body motion, and (ii) a respiratory ROI (RR ROI) targeting localised breathing-related motion.
The torso centre was computed as the midpoint between the shoulder midpoint and the hip midpoint. A quadrilateral defined by the shoulder and hip landmarks was used to estimate torso orientation as shown in figure 3 (a). A minimum-area bounding rectangle aligned with the torso axis was fitted and expanded by 10% to accommodate variations in posture.
To preserve gradual motion while reducing abrupt frame-to-frame shifts in the estimated landmark positions, the bounding box centre was smoothed using a recursive exponential moving average (EMA) filter defined as
\[\hat{\mathbf{c}}_{t} = \alpha_{t}\mathbf{c}_{t} + (1-\alpha_{t})\hat{\mathbf{c}}_{t-1}.\]
where \(\mathbf{c}_{t}\) is the current bounding box centre, \(\hat{\mathbf{c}}_{t}\) is the smoothed centre, and \(\alpha_{t}\) is an adaptive smoothing coefficient. The smoothing coefficient (\(\alpha_t\)) was adapted according to the frame-to-frame displacement of the bounding box centre. The displacement thresholds and corresponding smoothing coefficients were selected empirically based on visual inspection of ROI stability across representative neonatal recordings. The frame-to-frame displacement was defined as
\[d_t = \left\| \mathbf{c}_t - \hat{\mathbf{c}}_{t-1} \right\|_2,\]
where \(d_t\) is the Euclidean distance (in pixels) between the current bounding box centre (\(\mathbf{c}_t\)) and the smoothed centre from the previous frame (\(\hat{\mathbf{c}}_{t-1}\)). The smoothing coefficient was then selected according to
\[\alpha_t= \begin{cases} 0.2, & d_t \leq 2.55,\\ 0.6, & 2.55 < d_t \leq 7,\\ 1.0, & d_t > 7. \end{cases}\]
Lower smoothing coefficients placed greater emphasis on the previous bounding box centre to suppress landmark localisation jitter, whereas higher coefficients increased the influence of the current landmark estimate, allowing rapid ROI repositioning during substantial body motion. The resulting rectangle constituted the torso ROI and captured large-scale torso motion.
A \(75 \times 75\) pixel RR ROI was positioned over the lower abdominal region where breathing-induced motion is typically most pronounced. The RR ROI position was updated dynamically and smoothed using the same recursive filter applied to the torso ROI to reduce frame-to-frame jitter. The ROI size was selected based on previous work showing that small ROIs are more likely to contain breathing-related motion [46], [52]. Because respiratory movements are spatially localised, a small ROI focused on the abdominal region reduces interference from non-respiratory body movements and background motion. Figure 3 (b) illustrates the resulting torso and RR ROIs.


Figure 3: Dynamic ROI localisation using BlazePose landmarks. (a) Left and right shoulder and hip landmarks (blue circular markers) connected to form a torso quadrilateral used to estimate torso orientation; (b) The resulting torso ROI (blue). A \(75 \times 75\) pixel respiratory ROI (red) was positioned over the lower abdominal region to capture respiratory-related motion..
Using the dynamically tracked torso and respiratory ROIs, two camera-derived time-series signals were extracted from each 80-second video segment. First, a frame-difference (FD) signal was computed from the torso ROI to quantify global infant motion. This signal was computed as the mean absolute pixel-wise intensity difference between consecutive frames within the torso ROI. The resulting FD signal captures both respiratory motion and non-respiratory body movement. Second, a pixel-intensity-based respiratory signal (\(PPGi_{rr}\)) was extracted from the respiratory ROI. For each frame, the mean pixel intensity within the \(75 \times 75\) ROI was computed, generating a time-series that reflects respiratory-induced abdominal motion. Figures 4 presents representative examples of the extracted signals alongside physiological waveforms.
All extracted neonatal video segments were visually reviewed to identify recordings unsuitable for respiratory analysis. Segments were excluded if the infant’s torso was obstructed from the camera view during clinical interventions or kangaroo care, or when the infant was moved out of frame.
Of the annotated 346 COBE segments, 100 were excluded, leaving 246 COBE segments suitable for camera-based analysis. Similarly, 165 of the 608 normal breathing segments were excluded, resulting in 443 retained normal breathing segments. The final dataset therefore comprised 246 COBE episodes and 443 normal breathing episodes from 23 infants, for a total of 689 80-second segments.
The 80-second segments were then subdivided into overlapping 20-second windows with a 10-second step size, producing seven windows per 80-second segment. A 20-second window was sufficiently long to fully contain the shortest clinically relevant COBE episodes (10 seconds), while the 10-second step provided 50% overlap between adjacent windows, allowing events occurring near window boundaries to be captured within at least one analysis window. Each window was assigned a binary class label (COBE or non-COBE) based on the expert annotations, as illustrated in figure 4a. This procedure produced 4,823 windows, including 755 COBE and 4,068 non-COBE samples. Table 1 summarises the final machine learning dataset.
| Count | |
|---|---|
| Contributing infants | 23 |
| COBE segments (80 s) | 246 |
| Normal breathing segments (80 s) | 443 |
| Total 80-second segments | 689 |
| Positive 20-second windows | 755 |
| Negative 20-second windows | 4,068 |
| Total analysis windows | 4,823 |
Each 20-second window was treated as an individual training instance. To prevent data leakage, infant-level separation was preserved throughout dataset partitioning such that windows from the same infant were not present in both training and test sets. The dataset of 23 infants was divided into a training set (19 infants) and an independent test set (4 infants). The training set was further partitioned into five cross-validation folds, with each fold containing data from 3–4 infants. Dataset partitions were selected to maintain demographic balance with respect to gestational age, sex, and the distribution of COBE and non-COBE windows.
| Dataset | Positive windows | Negative windows | Total windows |
|---|---|---|---|
| Training | ,456 | ,095 | |
| Test | |||
| Total | ,068 | ,823 |
Figure 5 provides an overview of the machine learning approaches evaluated in this study. Residual Networks (ResNets) and ConvNeXt were investigated as the deep learning architectures for COBE detection. Both architectures were evaluated using three input configurations. Models were trained using individual camera-derived signals, combined camera-derived signals, and hybrid combinations of camera-derived and physiological signals. First, camera-only models were trained using the FD and \(PPGi_{rr}\) signals extracted from neonatal video recordings. Second, multimodal camera models combined FD and \(PPGi_{rr}\) using a late-fusion strategy. Finally, hybrid models integrated the camera-derived signals with physiological signals (IP, ECG-derived respiration (EDR), and the PPG envelope) using late fusion. While physiological signals provide high temporal resolution, they are susceptible to motion artefacts and electrode displacement. Conversely, camera-derived signals are non-contact but sensitive to occlusion and lighting variability. The hybrid framework was therefore designed to leverage complementary strengths across modalities and assess whether their integration improved COBE detection performance.
For multimodal and hybrid experiments, late fusion was employed. Each input modality was processed by an independent feature-extraction branch before the resulting feature representations were concatenated for classification. Although modalities were processed separately during feature extraction, all branches were trained jointly in an end-to-end manner using a shared classification objective, allowing gradients to optimise modality-specific feature representations simultaneously.
All models were trained and evaluated using the same dataset partitioning, windowing strategy, and evaluation protocol. Camera-only models were first evaluated using individual and combined camera-derived signals, after which the best-performing camera configurations were integrated with physiological signals in the hybrid framework.
To evaluate the feasibility of camera-only COBE detection, ResNet architectures were adapted for one-dimensional time-series analysis. All two-dimensional convolutional, pooling, and normalisation layers were replaced with their one-dimensional counterparts to process the two extracted camera-derived signals (FD and \(PPGi_{rr}\)).
Following Carter et al. [53], the number of feature channels per stage was reduced from the conventional [64,128,256,512] to [32,32,64,64], reflecting the lower feature complexity of one-dimensional signals compared with two-dimensional images. Adaptive average pooling was applied prior to classification to retain information from different regions of the feature map while maintaining a fixed-dimensional representation suitable for the classifier. The final classifier consisted of a multi-layer perceptron (MLP) replacing the standard single fully connected layer, enabling modelling of non-linear temporal feature interactions. Residual skip connections were retained throughout to maintain stable gradient propagation. The resulting architecture for camera-based COBE detection is shown in figure 6.
Three standard ResNet variants (ResNet-18, ResNet-34, and ResNet-50) were evaluated to assess the influence of network depth on camera-based COBE detection performance. Models were trained using the Adam optimiser with learning rates selected within the range \(9 \times 10^{-5}\) to \(1 \times 10^{-4}\) and exponential decay applied per epoch. Five-fold cross-validation was performed on the training set for model selection.
To evaluate whether combining camera-derived motion and respiratory information improved COBE detection, FD and \(PPGi_{rr}\) were also evaluated jointly using a late-fusion strategy. Each signal was processed by an independent ResNet branch, and the resulting feature representations were concatenated prior to classification as shown in figure 7. This configuration enabled comparison between models trained on FD alone, \(PPGi_{rr}\) alone, and the combined camera-derived signals.
Physiological respiratory signals were obtained from IP, ECG, and PPG recordings using the signal-processing pipeline described in Appendix 9. The filtered IP waveform was used directly as an input respiratory signal. An ECG-derived respiration (EDR) signal was computed from respiratory-induced variations in successive ECG R-peak amplitudes. Spline interpolation was then applied to obtain a continuous respiratory waveform. A respiratory envelope was derived from the PPG signal using peak detection followed by cubic spline interpolation. These approaches are consistent with established methods for extracting respiratory information from ECG and PPG signals [54]–[56]. The resulting respiratory signals were resampled to a common frequency of 60 Hz to facilitate multimodal analysis.
The hybrid network consisted of two modality-specific branches. The physiological branch employed a ConvNeXt architecture [57], which extends conventional convolutional networks through the use of grouped convolutions, Gaussian Error Linear Unit (GELU) activations, and inverted bottleneck blocks. The camera branch employed the adapted ResNet architecture described in Section 2.8. Each branch independently extracted modality-specific feature representations, which were subsequently concatenated using a late-fusion strategy. Late fusion was selected to preserve modality-specific feature extraction while enabling joint end-to-end optimisation of both branches. The fusion head comprised two fully connected layers with GELU activations, enabling modelling of non-linear interactions between modalities. The final layer output the probability of a COBE episode for each analysis window. The overall hybrid architecture is illustrated in Figure 8.
Both branches were trained jointly in an end-to-end manner using a shared weighted cross-entropy (CE) loss objective, defined as
\[\mathcal{L}_{\mathrm{CE}} = -\frac{1}{n} \sum_{i=1}^{n} \sum_{c=1}^{C} w_c\,y_{i,c}\log(\hat{p}_{i,c}),\]
where \(n\) is the batch size, \(C=2\) is the number of classes, \(w_c\) is the weight assigned to class \(c\), \(y_{i,c}\) is the one-hot encoded ground-truth label for sample \(i\), and \(\hat{p}_{i,c}\) is the predicted probability that sample \(i\) belongs to class \(c\). The ConvNeXt branch was optimised using AdamW with a weight decay of 0.05, label smoothing of 0.1, exponential moving averaging of model weights, and a cosine learning-rate schedule with a five-epoch warm-up period. The ResNet branch was optimised using Adam. Gradients from the fusion head were propagated through both branches during training, enabling learning of complementary modality-specific representations.
Model development was performed using five-fold cross-validation on the training set. Following cross-validation, the best-performing configuration was retrained using the complete training set and evaluated on the independent test set. Class imbalance was addressed using class-weighted cross-entropy loss, with higher loss weights assigned to the minority (COBE) class during training. Early stopping was applied when validation loss failed to improve for five consecutive epochs, preventing unnecessary training once performance had stabilised. Performance was evaluated using balanced accuracy, true positive rate (TPR), false positive rate (FPR), precision, F1 score, and Cohen’s \(\kappa\). Balanced accuracy was used as the primary metric to account for class imbalance and was calculated as:
\[\label{bal95accuracy} \text{Balanced Accuracy} = \frac{\text{TPR}_{\mathrm{COBE}}+\text{TPR}_{\mathrm{Normal}}}{2} \times 100\%\tag{1}\]
Cross-validation results are reported as mean \(\pm\) standard deviation across folds. Model selection prioritised configurations achieving the highest mean cross-validation balanced accuracy while favouring lower variance when performance was comparable.
We evaluated the performance of camera-only ResNet models and hybrid multimodal models combining camera-derived and physiological signals. Results are presented for cross-validation as mean \(\pm\) standard deviation, followed by evaluation on an independent test set.
Cross-validation results for the ResNet models trained on individual camera-derived signals are summarised in Table ¿tbl:resnet95cv?. Among the single-signal models, the highest balanced accuracies were 63.3% \(\pm\) 4.3 for ResNet-50 trained on the FD signal and 68.7% \(\pm\) 5.8 for ResNet-34 trained on the \(PPGi_{rr}\) signal. The multimodal model combining FD and \(PPGi_{rr}\) achieved a balanced accuracy of 68.6% \(\pm\) 3.6.
|>
p2.2cm|>
p1.2cm|c|c|c|c|c|c|c|c|
& &
& & & & & & & & &
& 18 & 0.57 \(\pm\) 0.17 & 0.32 \(\pm\) 0.10 & 0.27 \(\pm\) 0.08 & 0.36 \(\pm\) 0.09 & 0.18 \(\pm\) 0.07 & 68.6 \(\pm\) 10.4 & 57.1 \(\pm\) 17.4 & 62.8 \(\pm\) 4.9
& 34 & 0.67 \(\pm\) 0.10 & 0.41 \(\pm\) 0.07 & 0.25 \(\pm\) 0.05 & 0.36 \(\pm\) 0.06 & 0.15 \(\pm\) 0.03 & 59.0 \(\pm\) 6.8 & 66.7 \(\pm\) 9.9 & 62.8 \(\pm\) 2.5
& 50 & 0.60 \(\pm\) 0.12 & 0.33 \(\pm\) 0.07 & 0.27 \(\pm\) 0.08 &
0.36 \(\pm\) 0.08 & 0.18 \(\pm\) 0.07 & 66.5 \(\pm\) 7.2 & 60.1 \(\pm\) 12.3 & 63.3 \(\pm\) 4.3
& 18 & 0.77 \(\pm\) 0.09 & 0.38 \(\pm\) 0.06 & 0.29 \(\pm\) 0.07 & 0.41 \(\pm\) 0.08 & 0.23 \(\pm\) 0.07 & 61.7 \(\pm\) 6.2 & 76.7 \(\pm\) 9.3 & 69.2 \(\pm\) 5.2
& 34 & 0.71 \(\pm\) 0.14 & 0.34 \(\pm\) 0.07 & 0.30 \(\pm\) 0.07 &
0.42 \(\pm\) 0.08 & 0.24 \(\pm\) 0.07 & 66.3 \(\pm\) 6.9 & 71.0 \(\pm\) 14.4 & 68.7 \(\pm\) 5.8
& 50 & 0.70 \(\pm\) 0.04 & 0.33 \(\pm\) 0.07 & 0.30 \(\pm\) 0.09 & 0.41 \(\pm\) 0.08 & 0.24 \(\pm\) 0.07 & 67.5 \(\pm\) 6.8 & 69.6 \(\pm\) 4.3 & 68.5 \(\pm\) 2.8
\(FD\) + \(PPGi_{rr}\) & & 0.76 \(\pm\) 0.07 & 0.39 \(\pm\) 0.05 & 0.28 \(\pm\) 0.07 & 0.41 \(\pm\) 0.07 & 0.22 \(\pm\) 0.06 & 61.4 \(\pm\) 4.8 & 75.8 \(\pm\) 6.8 & 68.6 \(\pm\) 3.6
Table ¿tbl:resnet95test? reports performance of the best ResNet variants on the independent test set. The \(PPGi_{rr}\)-based ResNet-34 achieved balanced accuracy of 76.9% on the independent test set (TPR = 0.80).
|>
p2.5cm|c|c|c|c|c|c|c|c|
&
& & & & & & & &
\(FD\) & 0.52 & 0.24 & 0.31 & 0.39 & 0.22 & 76.5 & 51.7 & 64.1
\(PPGi_{rr}\) & 0.80 & 0.26 & 0.38 & 0.52 & 0.37 & 73.6 & 80.2 & 76.9
\(FD\) + \(PPGi_{rr}\) & 0.70 & 0.24 & 0.37 & 0.49 & 0.34 & 75.9 & 69.8 & 72.9
Table ¿tbl:hybrid95table? summarises the performance of the hybrid models evaluated on the independent test set across different input combinations. Each configuration included at least one camera-derived signal combined with physiological information from IP, the PPG envelope, EDR, or their combination. Extended hybrid results, including additional signal combinations and cross-validation outcomes, are provided in the supplementary material (Tables ¿tbl:hybrid95cvtable? - ¿tbl:hybrid95testtable?).
|>
p3.5cm|c|c|c|c|c|c|c|c|
&
& & & & & & & &
\(FD\) + IP & 0.81 & 0.08 & 0.68 & 0.74 & 0.68 & 92.1 & 81.0 & 86.6
\(FD\) + \(PPG\) & 0.78 & 0.41 & 0.28 & 0.41 & 0.22 & 59.1 & 78.5 & 68.8
\(FD\) + \(EDR\) & 0.89 & 0.32 & 0.36 & 0.51 & 0.36 & 67.7 & 88.8 & 78.2
\(PPGi_{rr}\) + IP & 0.92 & 0.11 & 0.63 & 0.75 & 0.68 & 88.9 & 92.2 &
90.6
\(PPGi_{rr}\) + \(PPG\) & 0.76 & 0.22 & 0.42 & 0.54 & 0.41 & 78.4 & 75.9 & 77.1
\(PPGi_{rr}\) + \(EDR\) & 0.85 & 0.37 & 0.32 & 0.47 & 0.29 & 63.3 & 85.3 & 74.3
\(FD\) + \(PPGi_{rr}\) + IP & 0.91 & 0.12 & 0.60 & 0.72 & 0.65 & 87.7 & 90.5 & 89.1
\(FD\) + \(PPGi_{rr}\) + \(PPG\) & 0.80 & 0.29 & 0.36 & 0.50 & 0.34 & 71.0 & 80.2 & 75.6
\(FD\) + \(PPGi_{rr}\) + \(EDR\) & 0.72 & 0.19 & 0.44 & 0.55 & 0.42 & 81.0 & 72.4 & 76.7
\(FD\) + \(PPGi_{rr}\) + IP + PPG + EDR & & & & & & & &
This study evaluated the feasibility of detecting apnoea-related cessation of breathing (COBE) in pre-term infants using video camera-derived signals, both independently and in combination with physiological measurements. Following video quality assessment, 689 annotated 80-second segments from 23 infants were retained for analysis. The results demonstrate that non-contact video-derived features contain clinically relevant information for distinguishing COBE episodes from normal breathing, supporting the feasibility of video-based COBE detection in the NICU.
Camera-only models demonstrated that respiratory motion extracted from video can be used to detect COBE events. Among the camera-derived signals, the pixel-intensity respiratory signal (\(PPGi_{rr}\)) consistently outperformed the frame-difference (FD) signal, achieving a test accuracy of 76.9% compared to 64.1% for FD. While FD captures global torso motion, including non-respiratory movement, \(PPGi_{rr}\) focuses on localised abdominal motion and more directly reflects breathing-related dynamics. This suggests that respiration-specific motion features provide more discriminative information than general movement alone. These findings support the feasibility of non-contact respiratory monitoring in neonatal environments.
Combining FD and \(PPGi_{rr}\) did not improve performance beyond \(PPGi_{rr}\) alone (72.9% vs 76.9%), indicating that global motion information may introduce variability that obscures respiration-specific patterns. This may reflect the difficulty of distinguishing respiratory from non-respiratory motion in smaller datasets, highlighting the importance of feature selection in multimodal settings, particularly when training data are limited.
The strongest performance was achieved using hybrid models integrating video camera-derived signals with impedance pneumography (IP). The combination of \(PPGi_{rr}\) and IP achieved a test accuracy of 90.6%, exceeding the performance of video camera-only models and indicating a clear benefit of multimodal integration. For comparison, contact-only machine learning models evaluated on the same dataset achieved a maximum test accuracy of 88.7% [58]. The hybrid configuration therefore achieved comparable and slightly higher performance while incorporating visual information. Non-contact sensing can provide complementary respiratory information when combined with conventional physiological monitoring, particularly IP. While IP provides a direct measure of respiratory effort, it remains susceptible to motion artefacts and electrode displacement. Camera-derived signals contribute contextual motion cues and additional respiratory dynamics that enhance classification performance when integrated within a unified model.
In contrast, combining camera-derived signals with PPG or ECG alone yielded more modest improvements. For example, combining \(PPGi_{rr}\) with PPG resulted in only a marginal increase in performance compared to \(PPGi_{rr}\) alone, suggesting overlap in the physiological information captured by these signals. Similarly, combining FD with ECG improved performance relative to models trained on either signal alone. Notably, indiscriminate fusion of all available signals did not further improve performance (84%), indicating that adding redundant inputs may not introduce informative features. These observations emphasise that multimodal integration must be selective and physiologically motivated.
A primary limitation of this study is the modest dataset size, reflecting the challenges of acquiring high-quality neonatal video recordings in clinical environments. Although segment quality assessment was necessary to ensure reliable motion extraction, it reduced the number of analysable COBE events and limits evaluation across a wider range of clinical conditions. Additionally, BlazePose was originally trained primarily on adult data. Although it enabled ROI tracking for most analysed recordings, dedicated neonatal pose-estimation models may further improve robustness.
The hybrid framework was evaluated retrospectively. Future work should assess performance under varying lighting conditions, occlusion patterns, and caregiving activities. Larger multi-centre datasets will be essential to evaluate generalisability across diverse neonatal populations and incubator configurations.
This study demonstrates the feasibility of video camera-based detection of COBE in pre-term infants and shows that video-derived respiratory features contain clinically meaningful information. Video camera-only models achieved moderate detection performance, demonstrating feasibility. Hybrid integration with impedance pneumography improved accuracy, supporting the complementary value of non-contact sensing. These findings indicate that video camera-based monitoring can augment conventional NICU respiratory monitoring by providing additional contextual and motion-related information. Rather than replacing contact-based sensors, multimodal fusion offers a strategy for enhancing detection robustness in dynamic neonatal environments.
A reference respiratory rate (\(RR_{ip}\)) was recomputed directly from the raw impedance pneumography (IP) signal. The IP waveform was first resampled from 62.5 Hz to 24 Hz using cubic spline interpolation and de-trended to remove the DC component. An 8th-order high-pass Butterworth filter (0.033 Hz) and a 6th-order low-pass Butterworth filter (2.83 Hz) were then applied to remove baseline drift and high-frequency noise while preserving the physiological respiratory frequency range observed in the study population.
Respiratory cycles were identified using a peak-and-trough detection algorithm based on the Mean Average Curve (MAC) method described by Lu et al. [48]. The MAC signal provided a dynamic threshold for peak detection, with respiratory peaks defined as local maxima above the threshold and troughs identified as local minima between successive peaks. Peaks with amplitudes below 20% of the median respiratory amplitude were rejected.
Signal quality assessment was performed for each detected respiratory cycle using the methodology proposed by Li et al. [59]. Three complementary signal quality indices (SQIs) were computed. First, a physiological-bounding SQI (\(SQI_{phys}\)) assessed whether the instantaneous respiratory rate fell within a plausible range for pre-term infants (2–170 breaths/min). Second, a spectral-concentration SQI (\(SQI_{bin}\)) quantified the proportion of signal power concentrated around the dominant respiratory frequency. \(SQI_{bin}\) was assigned a value of 1 when at least 50% of the spectral power was contained within a 0.3 Hz bandwidth centred on the dominant frequency [46]. Third, the Spectral Purity Index (SPI) quantified the extent to which the respiratory waveform was dominated by a single periodic frequency component, with lower values indicating increased waveform irregularity, noise, or motion artefacts.
The three measures were combined to produce a breath-level signal quality index (\(SQI_{breath}\)). Respiratory cycles failing any of the quality criteria were excluded from respiratory rate estimation. This quality-control procedure reduced the influence of motion artefacts, electrode disturbances, and non-respiratory fluctuations on the recomputed respiratory rate.
After detecting the peaks and troughs of the respiratory signal and assessing the signal quality for each breath, respiratory rate was estimated using a breath-counting approach within a 10-second sliding window with a step size of 1 second. The number of accepted respiratory cycles within each window was converted to breaths per minute by scaling to a 60-second interval, producing the continuous \(RR_{ip}\) time series.
To support respiratory rate estimation during both normal breathing and COBE episodes, a motion signal was first computed from the torso ROI following the frame-difference approach described by Cattani et al. [22]. The absolute pixel-wise difference between consecutive video frames was calculated within the torso ROI and averaged across pixels within the ROI to quantify infant motion.
A camera-derived respiratory rate (\(RR_{cam}\)) was then estimated from the respiratory signal (\(PPGi_{rr}\)) extracted from the abdominal ROI. The respiratory signal was obtained from the green colour channel, which has been shown to capture respiratory-related motion effectively in skin regions [46], [60].
The \(PPGi_{rr}\) signal was de-trended and filtered using a motion-dependent filter-switching approach. Periods of normal activity and low activity associated with respiratory pauses were identified using the computed motion signal, and different filter sets were applied to each activity state. During normal activity, an 8th-order high-pass filter (0.42 Hz) and a 6th-order low-pass filter (2.75 Hz) were applied, corresponding to respiratory rates between 25 and 140 breaths/min. During periods of low activity, a 6th-order high-pass filter (0.20 Hz) and a 2nd-order low-pass filter (1.42 Hz) were applied, corresponding to respiratory rates between 12 and 85 breaths/min.
Respiratory cycles were identified using the Mean Average Curve (MAC) peak detector developed by Lu et al. [48] and the Boxed Slope Sum Function (BSSF) peak detector described by Zong et al. [61]. Signal quality assessment was performed using two complementary measures. First, an activity-based signal quality index (\(SQI_{act}\)) quantified the extent of motion artefacts using the frame-difference signal. Second, a peak-agreement signal quality index (\(SQI_{peak}\)) quantified agreement between respiratory cycles identified by the MAC detector and the BSSF. These measures were combined to produce a breath-level signal quality index (\(SQI_{breath}\)), and respiratory cycles failing the quality criteria were excluded from respiratory rate estimation.
Respiratory rate was estimated using a breath-counting approach within a 10-second sliding window with a 1-second step size. Each window was expanded to include the complete first and last respiratory cycles before breath counting was performed. The number of accepted respiratory cycles within the expanded window was converted to breaths per minute, producing the continuous \(RR_{cam}\) time series. The resulting RR estimates were subsequently aligned with the reference respiratory rate (\(RR_{ip}\)) using cross-correlation prior to dataset construction and analysis.
The IP signal was filtered using an 8th-order high-pass Butterworth filter (0.08 Hz) and a 6th-order low-pass Butterworth filter (2.75 Hz) to remove baseline drift and high-frequency noise while preserving the neonatal respiratory frequency range. The filtered waveform was resampled to 60 Hz and used directly as a measure of respiratory activity.
The ECG signal was filtered using an 8th-order high-pass Butterworth filter (0.67 Hz) and a 2nd-order low-pass Butterworth filter (4 Hz). R-peaks were identified using the Pan–Tompkins algorithm [62]. Respiratory information was obtained from variations in successive ECG R-peak amplitudes, which were interpolated using cubic splines to generate a continuous ECG-derived respiration (EDR) waveform. The resulting signal was resampled to 60 Hz.
The PPG signal was filtered using an 8th-order high-pass Butterworth filter (0.08 Hz) and a 2nd-order low-pass Butterworth filter (2.75 Hz). Successive PPG peaks were identified and interpolated using cubic splines to generate a continuous respiratory envelope representing respiratory-induced modulation of the PPG waveform. The resulting signal was resampled to 60 Hz.
Tables ¿tbl:hybrid95cvtable? and ¿tbl:hybrid95testtable? show the performances of hybrid models under different input combinations. Each hybrid model combines camera-derived signals with physiological inputs (IP, \(ECG\), and PPG envelope) to detect COBE episodes.
|>
p3.5cm|c|c|c|c|c|c|c|c|
&
& & & & & & & &
\(FD\) + IP & 0.86 \(\pm\) 0.04 & 0.20 \(\pm\) 0.02 & 0.46 \(\pm\)
0.10 & 0.59 \(\pm\) 0.09 & 0.48 \(\pm\) 0.09 & 79.9 \(\pm\) 2.4 & 85.4
\(\pm\) 4.3 & 82.6 \(\pm\) 2.6
\(FD\) + \(PPG\) & 0.78 \(\pm\) 0.10 & 0.43 \(\pm\) 0.08 & 0.26 \(\pm\) 0.04
& 0.39 \(\pm\) 0.04 & 0.19 \(\pm\) 0.06 & 56.6 \(\pm\) 7.5 & 78.1 \(\pm\) 9.7 & 67.4 \(\pm\) 6.1
\(FD\) + \(ECG\) & 0.78 \(\pm\) 0.08 & 0.39 \(\pm\) 0.08 & 0.29 \(\pm\) 0.06
& 0.42 \(\pm\) 0.06 & 0.23 \(\pm\) 0.07 & 61.2 \(\pm\) 8.1 & 78.1 \(\pm\) 8.8 & 69.6 \(\pm\) 5.0
\(PPGi_{rr}\) + IP &0.81 \(\pm\) 0.08 & 0.23 \(\pm\) 0.05 & 0.42 \(\pm\) 0.12 & 0.54 \(\pm\) 0.11 & 0.42 \(\pm\) 0.11 &77.4 \(\pm\)
4.5 & 81.0 \(\pm\) 8.0 & 79.2 \(\pm\) 3.9
\(PPGi_{rr}\) + \(PPG\) & 0.67 \(\pm\) 0.05 & 0.37 \(\pm\) 0.07 & 0.27 \(\pm\)
0.05 & 0.37 \(\pm\) 0.05 & 0.18 \(\pm\) 0.07 & 62.6 \(\pm\) 6.8 & 67.1 \(\pm\) 5.4 & 64.8 \(\pm\) 5.4
\(PPGi_{rr}\) + \(ECG\) & 0.68 \(\pm\) 0.05 & 0.35 \(\pm\) 0.06 & 0.28 \(\pm\)
0.06 & 0.39 \(\pm\) 0.06 & 0.21 \(\pm\) 0.06 & 65.0 \(\pm\) 6.2 & 68.3 \(\pm\) 5.1 & 66.7 \(\pm\) 3.5
\(FD\) + \(PPG\) + IP & 0.73 \(\pm\) 0.06 & 0.21 \(\pm\) 0.03 & 0.40 \(\pm\)
0.09 & 0.51 \(\pm\) 0.08 & 0.39 \(\pm\) 0.07 & 78.6 \(\pm\) 2.7 & 73.0 \(\pm\) 5.9 & 75.8 \(\pm\) 2.4
\(FD\) + \(ECG\) + IP & 0.88 \(\pm\) 0.04 & 0.20 \(\pm\) 0.03 & 0.47
\(\pm\) 0.09 & 0.61 \(\pm\) 0.08 & 0.50 \(\pm\) 0.08 & 80.2 \(\pm\) 2.8 & 87.8 \(\pm\) 4.3 & 84.0 \(\pm\) 2.0
\(FD\) + \(ECG\) + \(PPG\) & 0.79 \(\pm\) 0.08 & 0.39 \(\pm\) 0.06 & 0.28 \(\pm\) 0.05 & 0.42 \(\pm\) 0.05 & 0.23 \(\pm\) 0.05 & 60.8 \(\pm\) 5.7 & 79.0 \(\pm\) 7.6 & 69.9 \(\pm\) 4.5
\(PPGi_{rr}\) + \(PPG\) + IP & 0.80 \(\pm\) 0.07 & 0.21 \(\pm\) 0.03 & 0.43 \(\pm\) 0.08 & 0.55 \(\pm\) 0.08 & 0.43 \(\pm\) 0.07 & 78.9 \(\pm\) 3.1 & 79.7 \(\pm\) 6.8 & 79.3 \(\pm\) 2.4
\(PPGi_{rr}\) +\(ECG\) + IP & 0.78 \(\pm\) 0.08 & 0.21 \(\pm\) 0.03 &
0.43 \(\pm\) 0.10 & 0.55 \(\pm\) 0.10 & 0.43 \(\pm\) 0.10 & 79.3 \(\pm\) 3.1 & 77.6 \(\pm\) 7.5 & 78.5 \(\pm\) 3.8
\(PPGi_{rr}\) +\(ECG\) + \(PPG\) & 0.71 \(\pm\) 0.07 & 0.36 \(\pm\) 0.05 & 0.28
\(\pm\) 0.05 & 0.40 \(\pm\) 0.05 & 0.21 \(\pm\) 0.05 & 63.5 \(\pm\) 5.2 & 71.0 \(\pm\) 6.9 & 67.2 \(\pm\) 3.8
\(FD\) + \(PPGi_{rr}\) + IP & 0.80 \(\pm\) 0.08 & 0.23 \(\pm\) 0.03 & 0.41 \(\pm\) 0.11 & 0.54 \(\pm\) 0.11 & 0.41 \(\pm\) 0.12 & 77.0 \(\pm\) 3.5 & 79.8 \(\pm\) 7.5 & 78.4 \(\pm\) 4.5
\(FD\) + \(PPGi_{rr}\) + \(PPG\) & 0.84 \(\pm\) 0.08 & 0.42 \(\pm\) 0.09 & 0.29
\(\pm\) 0.06 & 0.43 \(\pm\) 0.06 & 0.24 \(\pm\) 0.08 & 58.6 \(\pm\) 8.6 & 83.6 \(\pm\) 7.8 & 71.1 \(\pm\) 5.7
\(FD\) + \(PPGi_{rr}\) + \(ECG\) & 0.74 \(\pm\) 0.05 & 0.38 \(\pm\) 0.07 & 0.28
\(\pm\) 0.06 & 0.40 \(\pm\) 0.06 & 0.21 \(\pm\) 0.06 & 62.1 \(\pm\) 6.9 & 73.8 \(\pm\) 4.9 & 68.0 \(\pm\) 3.8
\(FD\) + \(PPGi_{rr}\) + IP + \(PPG\) + \(ECG\) & 0.90 \(\pm\) 0.03 & 0.19 \(\pm\) 0.02 & 0.48 \(\pm\) 0.07 & 0.62 \(\pm\) 0.07 & 0.52 \(\pm\) 0.06 & 80.6 \(\pm\) 2.1 & 90.2 \(\pm\) 2.5 & 85.4 \(\pm\) 1.7
|>
p3.5cm|c|c|c|c|c|c|c|c|
&
& & & & & & & &
\(FD\) + IP & 0.81 & 0.08 & 0.68 & 0.74 & 0.68 & 92.1 & 81.0 &
86.6
\(FD\) + \(PPG\) & 0.78 & 0.41 & 0.28 & 0.41 & 0.22 & 59.1 & 78.5 & 68.8
\(FD\) + \(ECG\) & 0.89 & 0.32 & 0.36 & 0.51 & 0.36 & 67.7 & 88.8 & 78.2
\(PPGi_{rr}\) + IP & 0.92 & 0.11 & 0.63 & 0.75 & 0.7 & 88.9 & 92.2 &
90.6
\(PPGi_{rr}\) + \(PPG\) & 0.76 & 0.22 & 0.42 & 0.54 & 0.41 & 78.4 & 75.9 & 77.1
\(PPGi_{rr}\) + \(ECG\) & 0.85 & 0.37 & 0.32 & 0.47 & 0.29 & 63.3 & 85.3 & 74.3
\(FD\) + \(PPG\) + IP & 0.85 & 0.07 & 0.71 & 0.78 & 0.73 & 93.0 & 85.3 & 89.2
\(FD\) + \(ECG\) + IP & 0.91 & 0.12 & 0.61 & 0.73 & 0.66 & 88.1
& 91.4 & 89.7
\(FD\) + \(ECG\) + \(PPG\) & 0.85 & 0.32 & 0.35 & 0.50 & 0.34 & 68.2 & 85.3 & 76.8
\(PPGi_{rr}\) + \(PPG\) + IP & 0.91 & 0.13 & 0.58 & 0.71 & 0.63 & 86.8 & 90.5 & 88.7
\(PPGi_{rr}\) +\(ECG\) + IP & 0.91 & 0.13 & 0.59 & 0.71 & 0.64 &
86.8 & 91.4 & 89.1
\(PPGi_{rr}\) +\(ECG\) + \(PPG\) & 0.88 & 0.25 & 0.41 & 0.56 & 0.43 & 74.5 & 87.9 & 81.2
\(FD\) + \(PPGi_{rr}\) + IP & 0.91 & 0.12 & 0.60 & 0.72 & 0.65 & 87.7 & 90.5 & 89.1
\(FD\) + \(PPGi_{rr}\) + \(PPG\) & 0.80 & 0.29 & 0.36 & 0.50 & 0.34 & 71.0 & 80.2 & 75.6
\(FD\) + \(PPGi_{rr}\) + \(ECG\) & 0.72 & 0.19 & 0.44 & 0.55 & 0.42 & 81.0 & 72.4 & 76.7
\(FD\) + \(PPGi_{rr}\) + IP + \(PPG\) + \(ECG\) & 0.76 & 0.08 & 0.66 & 0.71 & 0.64 & 92.1 & 75.9
& 84.0