Hannah Ruth B. Labana, Marc Nathaniel Valeros, Angie M. Ceniza-Canillo
The human heart is a fist-sized muscular organ that serves as the primary component of the cardiovascular system. Its primary function is to facilitate the circulation of blood throughout the body. Plaque buildup along the artery walls, which can result in a heart attack, is the leading cause of Coronary Heart Disease, the most prevalent cardiovascular disease. As there is no cure for Coronary Heart Disease, prevention is imperative by identifying its risk factors. This research seeks to leverage artificial intelligence by building a machine learning model capable of predicting coronary heart disease using non-laboratory risk factors. The study used a data set comprising 319,795 observations and 18 attributes. As the data set is heavily imbalanced, Synthetic Minority Oversampling Technique (SMOTE) was applied to the training set. Three machine learning models were trained and tested, namely Random Forest (RF), Extreme Gradient Boosting (XGBoost), and RF-XGBoost. Furthermore, a five-fold cross-validation approach was implemented to prevent overfitting. Once the training and testing phase were completed, the three models were compared to each other and it resulted in the XGBoost model outputting the most appropriate metrics in predicting Coronary Heart Disease with an accuracy of 73.50%, recall of 70.39%, an area under precision-recall curve of 47.01%, a geometric mean of 72.08%, and a Matthew's correlation coefficient of 27.55%. As it dealt with a highly imbalanced data set, it only yielded an f1-score of 32.32%. © 2023 IEEE.
University of San Carlos Cebu City, Department of Computer, Information Sciences and Mathematics, Cebu, Philippines