Back to projects

Predicting Heart Disease Using Machine Learning

Machine Learning·2026
Predicting Heart Disease Using Machine Learning

Exploratory analysis comparing five ML models for heart disease risk prediction on a heavily imbalanced clinical dataset of 130,000 records.

Technologies:

Machine LearningPythonScikit-learnLightGBMSMOTE

Details:

This study investigates heart disease risk prediction using a clinical survey dataset of 129,998 records across 16 features. The target variable was severely imbalanced, with only 9.4% positive heart disease cases versus 90.6% negative, making minority class recall the primary evaluation metric throughout. The dataset contained binary health indicators (HighBP, HighChol, Stroke, Smoker), continuous variables (BMI, Age, PhysHlth, MentHlth), and one ordinal feature (Diabetes). EDA revealed that Age, HighBP, and Stroke had the strongest Pearson correlations with the target. Missing BMI values were imputed using modal BMI within corresponding age groups to respect the skewed distribution. Duplicate rows were retained initially to preserve minority class representation, then removed in the final improvement stage. Features with low predictive value including BMI, Education, Income, and PhysActivity were dropped after correlation analysis. Five models were trained and evaluated across three class balancing strategies: no resampling (baseline), SMOTE, and SMOTEENN with hyperparameter tuning via GridSearchCV. The models were Logistic Regression, Random Forest, SVM, LightGBM, and a sequential ANN. The Diabetes column was one-hot encoded in the final stage to allow each severity category to contribute independently. Across the baseline, all models defaulted heavily toward the majority class. Logistic Regression achieved 89% accuracy but only 26% minority recall, missing most actual heart disease cases. The ANN performed best at baseline with minority recall of 0.82 and AUC of 0.83. After applying SMOTE, LR+SMOTE recorded the strongest minority recall of 0.75 and AUC of 0.80. In the final SMOTEENN stage with hyperparameter tuning, Logistic Regression achieved the highest AUC of 0.826, correctly ranking a heart disease patient above a healthy individual in 82.6% of cases, confirming it as the most discriminative model overall.