Heart Disease Prediction and Analysis
- Role
- Full-Stack / ML Project
- Timeline
- Academic Data Mining Project
- Status
- Completed
- Category
- AI / Machine Learning
A full-stack machine-learning application that predicts heart-disease risk from clinical and physiological features, with a Python ML pipeline, FastAPI backend, and React frontend.
This project applies data mining and machine-learning techniques to predict the presence of heart disease from 11 clinical features. It compares five classification algorithms, selects the best performer through stratified cross-validation and grid-search tuning, and exposes the trained model through a FastAPI backend and interactive React frontend.
The problem
Heart disease is one of the leading causes of death worldwide. Early detection from clinical data - blood pressure, cholesterol, resting ECG, exercise-induced angina, and similar features - can significantly aid clinical decision-making. This project provides an end-to-end ML pipeline that makes model comparison, evaluation metrics, and prediction results accessible through a web interface.
Approach
Exploratory data analysis
The dataset is analysed for target distribution, feature correlations, and class balance. Visualisations include a correlation heatmap, ROC curve with AUC score, confusion matrix, feature importance chart, and distribution plots for all numeric and categorical features.
Preprocessing and feature engineering
Missing values are handled, categorical variables encoded, and numeric features scaled. A data pipeline ensures consistent transformations between training and prediction.
Model comparison
Five classification algorithms are evaluated side-by-side: Naive Bayes, Logistic Regression, Decision Tree, K-Nearest Neighbors, and Random Forest. Each is assessed using stratified K-fold cross-validation to avoid data-split bias.
Hyperparameter tuning
The best-performing models are tuned with GridSearchCV. Random Forest achieved the highest results: 89.13% test accuracy, 90.48% F1 score, and 92.65% ROC-AUC on the hold-out test set.
API and frontend
A FastAPI backend with Pydantic schemas exposes a prediction endpoint. The React/TypeScript/Vite frontend provides an interactive prediction form, a model performance dashboard, and a gallery of generated visualisation figures.
Stack
Machine Learning
Backend
Frontend
What it changes
- Random Forest best model: 89.13% test accuracy, 90.48% F1 score, 92.65% ROC-AUC.
- Five-model comparison with stratified cross-validation and grid-search tuning.
- Interactive prediction form and model performance dashboard.
- Permutation-based feature importance and full visualisation gallery.
This project is an educational data mining exercise. It does not constitute medical diagnosis, clinical advice, or a substitute for professional medical evaluation. Do not use model predictions as the basis for clinical decisions.