This project performs Exploratory Data Analysis (EDA) on IPL ball-by-ball datasets from 2008 to 2023 to uncover meaningful insights about player performance, team statistics, match outcomes, and venue trends.
Using Python-based data analysis libraries, the project transforms raw cricket data into visual and statistical insights that support data-driven understanding of the Indian Premier League.
The IPL generates millions of ball-by-ball records every season, making it difficult to manually identify long-term trends and player performance.
This project analyzes historical IPL data to answer key cricket-related questions using data cleaning, feature engineering, visualization, and exploratory analysis techniques.
- IPL Ball-by-Ball Data (2008–2023)
- Match-level and delivery-level datasets
- Python
- Pandas
- NumPy
- Matplotlib
- Jupyter Notebook
✔ Data Cleaning
✔ Handling Missing Values
✔ Feature Engineering
✔ Data Transformation
✔ Aggregation & Filtering
Who scored the highest runs in each IPL season?
Which batsman hit the highest number of sixes every season?
Who took the most wickets in each season?
Which bowler bowled the highest number of dot balls?
Highest target set in every IPL season.
Lowest target recorded in every season.
Number of championships won by each franchise.
Overall winning percentage while batting first compared to chasing.
Fastest half-century in every IPL season.
Stadiums that hosted the highest number of IPL matches.
- Exploratory Data Analysis (EDA)
- Data Cleaning
- Missing Value Handling
- Feature Engineering
- Data Visualization
- Statistical Analysis
- Python Programming
- Analytical Thinking
- Interactive Power BI Dashboard
- Streamlit Web Application
- Predictive Match Outcome Model
- Player Performance Forecasting
- Machine Learning Integration
IPL-Data-Analysis/
│
├── data/
├── notebooks/
├── images/
├── README.md
└── requirements.txt
Aarathi T V
Computer Science Engineering Student
Passionate about Data Analytics, AI, Software Development, and Project Management.








