A comprehensive data analysis project exploring Indian Premier League (IPL) ball-by-ball data using Python, Pandas, NumPy, and Matplotlib.
This project analyzes over 260,000 delivery records across multiple IPL seasons to uncover team performance trends, scoring patterns, batting dominance, bowling pressure, and match-phase dynamics through data cleaning, feature engineering, statistical analysis, and visualization.
The Indian Premier League (IPL) generates massive amounts of ball-by-ball data every season. This project leverages data science techniques to extract meaningful insights from historical IPL matches and answer key cricket analytics questions.
The analysis focuses on:
- Team scoring performance
- Boundary distribution
- Match phase scoring patterns
- Top run-scorers
- Bowling pressure through dot balls
- Statistical outlier detection
Which teams have scored the most total runs in IPL history, and how are boundaries distributed across teams?
How do scoring rates differ across the three match phases?
- Powerplay (Overs 1–6)
- Middle Overs (Overs 7–15)
- Death Overs (Overs 16–20)
Who are the top run-scorers in IPL history, and which bowling teams apply the most pressure through dot balls?
- Python
- Pandas
- NumPy
- Matplotlib
- Jupyter Notebook
Dataset: IPL Ball-by-Ball Dataset
Source: Kaggle
The dataset contains:
- 260,000+ ball-by-ball records
- Batting and bowling teams
- Batter and bowler information
- Runs scored
- Extras
- Wickets and dismissals
- Match over information
The dataset underwent several preprocessing steps:
- Handled missing values
- Filled null wicket-related fields with meaningful labels
- Removed super-over deliveries
- Standardized franchise names
- Removed duplicate records
- Corrected data types for numerical analysis
Examples:
- Delhi Daredevils → Delhi Capitals
- Kings XI Punjab → Punjab Kings
- Royal Challengers Bangalore → Royal Challengers Bengaluru
Four new analytical features were created:
| Feature | Description |
|---|---|
| over_phase | Categorizes overs into Powerplay, Middle, and Death |
| is_boundary | Flags fours and sixes |
| is_dot_ball | Identifies pressure deliveries |
| run_category | Categorizes runs into Dot, Single, Double, Triple, Four, and Six |
- Total runs scored by each IPL franchise
- Boundary frequency across teams
- Average runs per ball by phase
- Total runs scored by phase
- Top 10 IPL run-scorers
- Contribution comparison
- Dot-ball percentage by bowling team
- Mean and standard deviation calculations
- Run distribution analysis
- Correlation analysis
- Z-score based outlier detection
The project includes:
- Team total runs bar chart
- Top 10 batters bar chart
- Dot-ball percentage pie chart
- Boundary distribution histogram
- Phase-wise scoring comparison charts
- Mumbai Indians and Chennai Super Kings have scored the highest total runs in IPL history.
- Death Overs (16–20) produce the highest scoring rate.
- Middle Overs are generally the most restrictive phase.
- A small group of elite batters contribute a significant portion of total IPL runs.
- Virat Kohli consistently appears among the leading run scorers.
- Strong bowling teams maintain higher dot-ball percentages.
- Dot-ball percentage serves as a useful indicator of bowling pressure.
- Z-score analysis identified rare high-impact deliveries that can influence aggregate statistics.
IPL-Ball-by-Ball-Analysis/
│
├── Dataset.csv
├── Pandas_Analysis.ipynb
├── NumPy_Analysis.ipynb
├── Matplotlib_Visualizations.ipynb
├── images/
│ ├── team_runs.png
│ ├── top_batters.png
│ ├── dot_ball_percentage.png
│ └── boundary_distribution.png
│
└── README.md
This project showcases:
- Data Cleaning
- Exploratory Data Analysis (EDA)
- Data Wrangling
- Feature Engineering
- Statistical Analysis
- Data Visualization
- Python Programming
- Pandas
- NumPy
- Matplotlib
- Predictive modeling for match outcomes
- Player performance forecasting
- Venue-based analysis
- Toss impact analysis
- Advanced visualizations using Seaborn and Plotly
- Machine learning models for run prediction
- Mahdiat Rahman
- Md. Golam Rabbani Sajib
- Rabeta Tanjum Arni
IPL Complete Dataset (2008–2020)
Source: https://www.kaggle.com/datasets/patrickb1912/ipl-complete-dataset-20082020