This project presents an end-to-end Customer Churn Analysis and Machine Learning classification workflow using customer data from a telecommunications company.
The objective was to identify the main factors associated with customer churn and develop predictive models capable of identifying customers with a higher risk of leaving the company.
The project combines:
- Data exploration
- Data cleaning and preprocessing
- Feature engineering
- Exploratory Data Analysis (EDA)
- Data visualization
- Machine Learning classification
- Model evaluation and comparison
- Feature importance analysis
- Business interpretation
- Customer retention insights
Rather than focusing only on predictive performance, the project connects Machine Learning results with business-oriented customer retention decisions.
The analysis was designed to answer several business questions:
- Which customers are more likely to churn?
- Which contract types present the highest churn risk?
- How does customer tenure affect churn?
- Are higher monthly charges associated with customer churn?
- Which internet services are associated with higher churn?
- Do Online Security and Tech Support influence customer retention?
- Are certain payment methods associated with lower churn?
- Which customer characteristics are the strongest churn indicators?
- Can Machine Learning models identify customers at risk of leaving?
- How can these findings support customer retention strategies?
- Python
- Pandas
- NumPy
- Matplotlib
- Seaborn
- Scikit-Learn
- Jupyter Notebook
- Data Cleaning
- Data Preprocessing
- Feature Engineering
- Exploratory Data Analysis
- Data Visualization
- Machine Learning
- Classification
- Logistic Regression
- Random Forest
- Model Evaluation
- Feature Importance
- Business Analysis
- Analytical Storytelling
The dataset contains customer information from a telecommunications company.
The available variables include information related to:
- Customer demographics
- Contract type
- Customer tenure
- Internet services
- Additional services
- Payment methods
- Monthly charges
- Customer account information
- Churn status
The target variable used for the classification problem is:
Churn
Yesβ Customer left the companyNoβ Customer remained with the company
The objective of the Machine Learning workflow is therefore to identify patterns associated with customers who are more likely to churn.
The Exploratory Data Analysis investigated the relationship between customer churn and several customer characteristics.
The analysis focused particularly on:
- Contract Type
- Customer Tenure
- Monthly Charges
- Internet Service
- Online Security
- Tech Support
- Payment Method
- Senior Citizen Status
The objective was not only to identify statistical patterns, but also to understand which customer characteristics could have practical implications for retention strategies.
The project follows a complete Machine Learning workflow:
- Data Cleaning
- Data Preprocessing
- Feature Engineering
- Categorical Variable Encoding
- Train/Test Split
- Logistic Regression
- Random Forest Classifier
- Model Evaluation
- Model Comparison
- Feature Importance Analysis
- Business Interpretation
This workflow transforms raw customer information into a structured classification problem and evaluates multiple models using consistent performance metrics.
Two classification algorithms were implemented and compared.
Logistic Regression was used as an interpretable baseline classification model.
Its advantages for this problem include:
- Simple interpretation
- Efficient training
- Probabilistic classification
- Strong baseline performance
- Useful benchmark for more complex models
A Random Forest Classifier was also implemented to evaluate whether an ensemble model could improve predictive performance.
Its main characteristics include:
- Ability to capture non-linear relationships
- Interaction between multiple variables
- Robustness to complex feature relationships
- Feature importance estimation
The models were evaluated using several classification metrics:
- Accuracy
- Precision
- Recall
- F1 Score
- ROC AUC
| Model | Accuracy | Precision | Recall | F1 Score | ROC AUC |
|---|---|---|---|---|---|
| Logistic Regression | 0.80 | 0.64 | 0.57 | 0.60 | 0.836 |
| Random Forest | 0.79 | 0.63 | 0.52 | 0.57 | 0.818 |
The results show that Logistic Regression slightly outperformed Random Forest across the evaluated metrics.
Despite being the simpler model, Logistic Regression achieved the highest ROC AUC at 0.836, compared with 0.818 for Random Forest.
This demonstrates an important Machine Learning principle: a more complex model does not necessarily produce better predictive performance.
The following visualization compares the main evaluation metrics for both classification models.
Logistic Regression achieved slightly stronger overall performance, particularly in Recall, F1 Score, and ROC AUC.
For a churn prediction problem, Recall is particularly relevant because failing to identify customers who are actually at risk of leaving may result in missed retention opportunities.
The ROC Curve provides another perspective on the models' ability to distinguish between customers who churn and customers who remain.
The Logistic Regression model achieved the highest ROC AUC:
ROC AUC = 0.836
This indicates a solid ability to discriminate between churn and non-churn customers.
The Random Forest model was also used to explore the relative importance of different customer characteristics.
Feature importance analysis helps connect predictive modeling with business interpretation by identifying which variables contribute most strongly to the model's decisions.
These results can help prioritize the customer characteristics that deserve greater attention when designing retention strategies.
The analysis revealed several relevant patterns associated with customer churn.
Customers with month-to-month contracts are significantly more likely to churn.
Longer-term contracts appear to be associated with stronger customer retention.
Customers with short tenure present the highest churn risk.
This suggests that the early stages of the customer relationship may represent a critical retention period.
Higher monthly charges are associated with increased churn.
Customers paying higher monthly amounts may therefore require additional attention from a retention perspective.
Customers using Fiber Optic services show higher churn levels than customers using DSL.
This may indicate differences in pricing, expectations, service experience, or customer characteristics.
Customers without Online Security are considerably more likely to churn.
Customers without Tech Support also show higher churn rates.
These additional services appear to be associated with stronger customer retention.
Automatic payment methods are associated with lower churn.
This suggests that payment behavior may also provide useful information when identifying customers at risk.
Senior citizens present higher churn rates than other customer groups.
This segment may therefore require differentiated retention strategies or customer support approaches.
Based on the analysis, several retention initiatives could be considered:
- Prioritize customers with month-to-month contracts for retention campaigns.
- Monitor customers during the early stages of their tenure.
- Investigate the relationship between higher monthly charges and customer dissatisfaction.
- Evaluate the customer experience associated with Fiber Optic services.
- Promote services such as Online Security and Tech Support where appropriate.
- Encourage convenient automatic payment methods.
- Develop targeted retention strategies for customer groups showing higher churn risk.
- Use predictive churn scores to prioritize customers for proactive retention actions.
The objective is not simply to predict churn, but to transform predictive results into actionable business decisions.
Project_02_Customer_Churn/
βββ data/ βββ images/ β βββ feature_importance.png β βββ model_comparison.png β βββ roc_curve.png βββ notebooks/ β βββ 02_customer_churn_analysis.ipynb βββ requirements.txt βββ README.md βββ .gitignore
This project demonstrates the development of a complete Machine Learning classification workflow, from customer data exploration and preprocessing to predictive modeling, evaluation, and business interpretation.
The project demonstrates practical skills in:
- Python data analysis
- Data cleaning and preprocessing
- Feature engineering
- Exploratory Data Analysis
- Data visualization
- Classification modeling
- Logistic Regression
- Random Forest
- Model evaluation
- ROC AUC analysis
- Feature importance analysis
- Customer churn analysis
- Business insight generation
- Analytical storytelling
Most importantly, the project demonstrates the ability to connect Machine Learning results with a real business problem: customer retention.
As Project 02 of the portfolio, it extends the exploratory analytics foundation developed in Project 01 by introducing predictive modeling and model evaluation.
MartΓn Panelo
Data Analyst | Geophysicist | Scientific Computing
Analytical professional combining data analytics, scientific computing, and geoscience experience, with a focus on Python, SQL, Power BI, data visualization, and business-oriented problem solving.
- GitHub: PaneloMartin
- LinkedIn: MartΓn Panelo


