Applied AI

Low-cost single-lead ECG for vascular-age prediction

Co-authored a peer-reviewed Springer study; contributed data collection, the signal and feature pipeline, and model training and evaluation.

Role
Co-author with equal contribution, third of seven authors. My part
Period
Published 2025
Status
Published in Circuits, Systems, and Signal Processing (Springer), volume 44, issue 8, pages 5852-5875, 2025.
Outcome
Random forest R² 0.99 on 6,131 segmented samples, and R² 0.87 on 42 unsegmented samples with transfer learning.

Read the paper

01 / Problem

The study asks whether a low-cost single-lead ECG module can predict vascular age, and whether it reveals smoking-induced changes in the ECG. The cohort was 42 apparently healthy subjects aged 18 to 30, 20 of them light but habitual smokers. That is a small cohort, and one lead carries less information than a twelve-lead ECG. Others built the ECG module, so my work starts at the captured signal. The engineering problem was to get a model that means something out of 42 people.

My role: Co-author with equal contribution, third of seven authors. My part: data collection, the signal and feature pipeline, model training and evaluation. Not the ECG hardware.

02 / System

Captured signal 1 Preprocessing 2 Segmentation 3 Features 4 Models 5 Evaluation 6 Captured signal 1 Preprocessing 2 Segmentation 3 Features 4 Models 5 Evaluation 6
One recording, from capture to evaluation. I did not build the capture hardware. I collected data and built everything after it.

Select a component to read what it does and how it fails.

  1. Captured signal. Single-lead ECG from a low-cost module that others built. My work starts at the captured signal; every stage after it inherits its noise.
  2. Preprocessing. Cleans the raw signal before anything is measured. A weak cleaning step corrupts every feature after it.
  3. Segmentation. Overlapping 5-second windows with a 1-second stride turn 42 samples into 6,131. The windows overlap, so they are not independent.
  4. Features. 21 features: 13 from the ECG, such as intervals and QRS duration, and 8 demographic or clinical. Every interval definition can hide an error.
  5. Models. Regression baselines, a decision tree, random forest, 1D-CNN and ResNet-18 transfer learning. Random forest was the paper's best model in all three setups.
  6. Evaluation. Three setups: segmented, unsegmented, and unsegmented with transfer learning from a public PPG dataset. All three are reported, including the weak one.

03 / Decisions

Segment the recordings

Decision
Overlapping 5-second windows with a 1-second stride, which turned 42 subject-level samples into 6,131.
Rejected
Relying on one sample per subject alone, though the paper reports that setup too.
Why
42 rows is too few to train most models on. Windows give the models thousands of rows to learn from.
Cost
The windows overlap and come from the same 42 people, so they are not 6,131 independent samples. What broke, below, follows from this.

Transfer learning for the small set

Decision
Pre-train on a public PPG dataset, then fine-tune on this ECG data, for the 42-sample set.
Rejected
Training on the 42 samples alone.
Why
Alone, the random forest scored R² 0.26. With the pre-training it scored R² 0.87 on the same 42 samples.
Cost
The result leans on a dataset I did not collect, and 42 people are still 42 people.

Engineered features and classical models first

Decision
21 engineered features feeding classical models, with the random forest as the headline model.
Rejected
A deep network as the headline model, though I implemented a 1D-CNN and ResNet-18 too.
Why
The random forest was the paper's best model in all three setups, on features I could inspect.
Cost
Every interval and duration comes from my pipeline, so a mistake there would sit silently under every model.

04 / What broke

A weak result on the small set

Symptom
Trained on the 42 unsegmented samples alone, the random forest scored R² 0.26 (MSE 3.56).
Cause
42 rows, one per person, gave the model too little to learn from.
Fix
Transfer learning from a public PPG dataset lifted it to R² 0.87 (MSE 0.99) on the same 42 samples. The weak number stays in the paper's table.

A headline number to read with care

Symptom
The segmented result, R² 0.99 on 6,131 samples, is the figure people quote, and the one to read with most care.
Cause
The 6,131 samples are overlapping windows cut from the same 42 people, so they are not 6,131 independent observations.
Fix
Windowing adds rows, not people, so nothing inside this dataset fixes it. I read R² 0.87 on 42 samples, one per person, as the more conservative figure, and I would want more subjects before trusting either.

05 / Outcome

  • Across 42 subjects and 21 features, the random forest reached MSE 0.07 on the 6,131 segmented samples and MSE 0.99 on the 42 unsegmented samples with transfer learning.
  • I implemented and evaluated ResNet-18 transfer learning as well. The paper prints no accuracy figure for it, so I quote none.

06 / The rule I took from this

Report a result with its sample, or do not report it.

Rule 9 of 9

Stack

  • Python
  • Keras
  • scikit-learn
  • NeuroKit2
  • SciPy
  • random forest
  • ResNet-18 transfer learning

Next