Comparing the performance of machine learning and conventional models for predicting atherosclerotic cardiovascular disease in a general Chinese population
Student10:55CCAI
paperi.ai
0:00 / 0:00
Zihao Fan, Zhi Du, Jinrong Fu, Ying Zhou, Pengyu Zhang, Chuning Shi, Yingxian Sun
What if the usual cardiovascular risk calculators are leaving useful heart signals on the table? This study tests whether machine learning can combine routine factors with detailed ECG and echocardiography data to improve prediction.
What if the usual cardiovascular risk calculators are leaving useful heart signals on the table? This study tests whether machine learning can combine routine factors with detailed ECG and echocardiography data to improve prediction. Atherosclerotic cardiovascular disease, or ASCVD, is defined here as nonfatal acute myocardial infarction, coronary heart disease death, and stroke, and it has become the leading cause of morbidity and mortality worldwide.
In China, cardiovascular disease accounts for 38.9 percent of deaths among females and 35.5 percent among males, with approximately 61 percent of these deaths attributed to ASCVD. The clinical challenge is that identifying high-risk individuals is critical for primary prevention, while personalized cardiovascular risk assessment remains difficult in practice.
The Pooled Cohort Equations are widely used, but they were derived primarily from non-Hispanic White and African-American populations, and overestimation or underestimation has been reported in specific populations. China-PAR was designed for Chinese adults and approved by Chinese guidelines in 2019, although more evidence is needed to confirm its generality across different populations.
Recent evidence suggests that, beyond traditional risk factors such as age, sex, and smoking, electrocardiography and echocardiography parameters can help predict cardiovascular disease by detecting subclinical impairment before cardiac symptoms appear. Abnormality in the P wave, PR intervals, and left ventricular ejection fraction has been associated with adverse cardiovascular outcomes.
Yet the massive amounts of parameters reflecting cardiac electrophysiology, structure, and function were unlikely to be included in traditional prediction models. A trainable model for different populations may therefore provide novel insights for ASCVD prevention.
The study established machine-learning risk prediction models from a dataset integrating demographic, behavioral, psychological, electrocardiograph, and echocardiography variables in a community-based general population in Northeast China. It then compared machine-learning algorithms with the traditional Cox regression models PCE and China-PAR to evaluate which method provided superior predictive performance.
The Northeast China Rural Cardiovascular Health Study is a prospective population cohort using multistage, stratified, random cluster sampling, with 11,956 rural residents aged at least 35 years recruited in Liaoning Province between January 9 and August 23, 2013. The study collected demographics, physical status and vital signs, medical histories, echocardiography data, ECG examinations, and laboratory data.
Participants were included if they were between 35 and 85 years old and had no history of cardiovascular disease. Of the 11,956 participants assessed for eligibility, 10,349, or 86.6 percent, completed at least one follow-up visit. Figure one traces how the study cohort was assembled for the final analyses.
From the eleven thousand nine hundred fifty-six-person total cohort, one thousand six hundred seven were lost to follow-up, while seven hundred forty were excluded for cardiovascular disease history or missing electrocardiography or echocardiography data. The resulting nine thousand six hundred nine participants were divided into a seven thousand six hundred eighty-eight-person training cohort and a one thousand nine hundred twenty-one-person test cohort, clarifying the basis for model development and evaluation.
Standard twelve-lead ECGs were recorded in the resting supine position at baseline and analyzed automatically with the MUSE Cardiology Information System and the GE Marquette twelve-SL ECG analysis program. The GE system processed a total of 645 parameters from the unprocessed digital ECG data.
Two hundred one parameters, including relative coordinate points and calculated values, were excluded. The remaining 444 ECG variables were used for analysis. The primary endpoints included stroke and coronary heart disease, with each participant followed through health status, hospital admissions, outpatient diagnoses, and deaths from 2015 to 2018.
Two physicians independently reviewed medical records, categorized events, and specified event dates. Stroke was defined as sudden focal neurological dysfunction lasting 24 hours or until death, or lasting less than 24 hours with a clinically relevant brain lesion.
Coronary heart disease included myocardial infarction, resuscitated cardiac arrest, definite angina, probable angina followed by revascularization, and coronary heart disease death. The datasets were randomized into training and testing sets, with 80 percent used for training and 20 percent for testing.
Model development tried several classifiers: artificial neural network, random forest, gradient boosting machine, K nearest neighbors, Adaptive Boosting, support vector machine, and Categorical Boosting. The models used an optimal subset, stratified ten-fold cross-validation on the training set, and grid search to determine appropriate hyperparameters.
To address class imbalance, more weight was assigned to minority-class samples, increasing their misclassification cost. Performance was evaluated using discrimination, calibration, net benefit, and net reclassification improvement against PCE, China-PAR, recalibrated PCE, and recalibrated China-PAR.
Table one compares baseline clinical characteristics for the total cohort of nine thousand six hundred nine participants and the four hundred thirty-one participants with ASCVD. It reports continuous measures as means with standard deviations and categorical measures as counts and percentages, including age, blood pressure, laboratory values, heart rates, smoking, and drinking status.
The authors used t-tests or Mann–Whitney U tests for continuous variables, and chi-squared or Fisher’s exact tests for categorical variables, establishing the clinical context for later analyses. Figure 2 compares discrimination and calibration for PCE and China-PAR before and after recalibration.
All of the conventional models showed moderate discrimination, with China-PAR having the highest discrimination and an AUC of 0.780. However, every model showed poor calibration, with a Hosmer–Lemeshow chi-squared value greater than 18 and a p value below 0.05. The Brier score ranged from 0.043 to 0.057, while MCC ranged from 0.186 to 0.194.
Stepwise model building and the RFE algorithm reduced the final machine-learning ASCVD models to 30 key predictor variables. Figure 3 compares discrimination and calibration among the established machine-learning classifiers. The artificial neural network outperformed the other classifiers and had the greatest AUC value and consistency.
Its AUC was 0.800, higher than the China-PAR and PCE models, although the reported p values were 0.12 and 0.08. Figure three compares the machine-learning models in the test cohort using discrimination and calibration. In panel A, receiver operating characteristic curves show the ANN model with an AUC of zero point eight zero zero, alongside Catboost at zero point seven eight seven, GBM at zero point seven seven four, KNN at zero point seven six seven, RF at zero point seven five nine, Adaboost at zero point seven two seven, and SVM at zero point six nine seven.
Panel B presents Hosmer–Lemeshow calibration plots, showing how observed and predicted risks correspond across models. Figure two compares discrimination and calibration for the original China-PAR and PCE models, alongside their recalibrated versions. In panel A, the receiver operating characteristic curves cluster closely, with AUC values from zero point seven seven seven to zero point seven eight zero, indicating moderate discrimination.
Panel B shows the calibration plots and reports Hosmer–Lemeshow chi-squared values from eighteen point six to one hundred twenty-six point six, with the accompanying text noting poor calibration for all models. Figure four presents decision curves across threshold probabilities from zero to twenty percent, comparing PCE, China-PAR, ReChina-PAR, and several machine-learning models.
The vertical axis, net benefit, is shown alongside “all” and “none” reference strategies, illustrating how clinical usefulness changes as the high-risk threshold changes. This matters because the chart evaluates these prediction approaches in a decision-focused way, complementing the authors’ reported calibration and overall-performance assessment of the ANN model.
The limitations include that the machine-learning algorithm cannot assess the independent effects of each variable on events and may make it difficult to identify specific treatments to reduce individual risk. Longer follow-up is needed because ASCVD is chronic and progressive.
The study excluded 84 variables with more than 10 percent missing data that might have had predictive value, and missing-data imputation might bias the analysis. External validation studies are required to demonstrate the accuracy of the model’s predictions in diverse populations.
The models used initial follow-up blood pressure and glucose, but changes during follow-up were not taken into account. In this cohort, the machine-learning artificial neural network showed the strongest overall discrimination among the tested models, but the result still needs longer follow-up and external validation before broad clinical use.
A derivative work by Paperi · AI-generated script, voice and captions
· pages and figures unaltered
Made with Paperi.
Drop in a research PDF — get a narrated video walkthrough like this one,
with highlights that follow the narration. Free to start.
Renée M. Ferrari, Jennifer Leeman, Alison T. Brenner, Sara Y. Correa, Teri L. Malo, Alexis Moore, Meghan C. O’Leary, Connor M. Randolph, Shana Ratner, Leah Frerichs, Deeonna E. Farr, Seth D. Crockett, Stephanie B. Wheeler, Kristen Hassmiller Lich, Evan Beasley, Michelle Hogsed, Ashley Bland, Claudia Richardson, Mike Newcomer, Daniel S. Reuland
Colorectal cancer screening can save lives, yet the people most likely to be missed are often those with lower incomes or no insurance. This study asked how to make screening fit the places serving them.Colorectal cancer screening can save lives, yet it remains especially underused in underserved communities. This project’s surprising solution was not simply a better test—it was redesigning who does the work and how the pieces connect.
Beibei Xiong, Christine Stirling, Daniel X. Bailey, Melinda Martin‐Khan
A hospital can have skilled professionals, a national care standard, and still fail to give every patient one joined-up plan. This study found that confidence was high—but the support needed to make care consistent was not.Australian hospital professionals generally felt confident delivering comprehensive care—but the weakest point was creating one shared care plan, and nearly one in five reported that none of their unit’s patients had one.
Cancer care is not only about choosing a treatment. Patients and families also need someone to help connect medical decisions, daily problems, and the long stretch after treatment. This study asks whether oncology nurses could fill that missing role in Israel.Israel faces a large cancer burden, yet oncology clinical nurse specialists have not been established there. This study asks what nurses and physicians think such a role could contribute—and where its boundaries should be drawn.