This project aims to predict customer churn using machine learning models, helping financial institutions better understand customer behavior and reduce attrition. The analysis was done on real-world datasets, with detailed preprocessing, feature selection, and model evaluation steps.
- Objective: Predict whether a customer will churn based on demographic and behavioral features.
- Dataset Size: 175,000+ rows from two datasets (merged)
- Tools Used: Python, Pandas, Scikit-learn, Matplotlib, Seaborn
- ML Models: Random Forest, SVM, Decision Tree, Naive Bayes
-
π§Ή Data Cleaning & Preparation:
Merged two datasets, handled missing values, encoded categorical variables, and applied feature scaling (MinMaxScaler, LabelEncoder). -
π EDA & Feature Selection:
Performed exploratory analysis using Seaborn and Matplotlib. Used SelectKBest (Chi-Square), RFECV, and FAMD for dimensionality reduction and feature selection. -
π§ Model Training & Evaluation:
Trained four classifiers and evaluated using Accuracy, Precision, Recall, and F1-Score.
β Best Result: Random Forest with 85% Accuracy and 84% F1-Score -
π‘ Insights Discovered:
Customers who were older, less active, or held fewer products were more likely to churn, providing actionable insights for retention.
- Languages: Python
- Libraries: Pandas, NumPy, Scikit-learn, Seaborn, Matplotlib
- Techniques: Feature Encoding, Scaling, Chi-Square Selection, RFECV, FAMD, Classification Metrics