A machine learning project that classifies movie reviews as positive or negative using Natural Language Processing (NLP) techniques and a Multinomial Naive Bayes classifier.
This project implements a sentiment analysis classifier that can automatically determine whether a movie review expresses positive or negative sentiment. The model is trained on the IMDB Movie Reviews Dataset and achieves 85.4% accuracy on the test set.
- Comprehensive Text Preprocessing: HTML tag removal, URL cleaning, case normalization, and special character handling
- TF-IDF Vectorization: Converts text into numerical features for machine learning
- Hyperparameter Optimization: Uses validation set to find optimal alpha parameter for Naive Bayes
- Robust Evaluation: Includes train/validation/test split, K-fold cross-validation, and confusion matrix analysis
- Real-time Predictions: Can classify new movie reviews instantly
- Test Accuracy: 85.42%
- Cross-Validation Accuracy: 84.89% (±0.002)
- Balanced Performance: Equal precision and recall for both positive and negative classes
- Model Type: Multinomial Naive Bayes with α = 2.0
precision recall f1-score support
negative 0.85 0.85 0.85 2500
positive 0.85 0.85 0.85 2500
accuracy 0.85 5000
macro avg 0.85 0.85 0.85 5000
weighted avg 0.85 0.85 0.85 5000
- Python 3.x
- scikit-learn: Machine learning algorithms and evaluation metrics
- pandas: Data manipulation and analysis
- numpy: Numerical computing
- matplotlib: Data visualization
- re: Regular expressions for text preprocessing
sentiment-analysis-classifier/
│
├── Imdb_review_pred.ipynb # Main Jupyter notebook with complete pipeline
├── README.md # Project documentation
└── IMDB Dataset.csv # Dataset (not included in repo)
pip install pandas numpy scikit-learn matplotlib reThis project uses the IMDB Movie Reviews Dataset. You'll need to download it and place it in the project directory as IMDB Dataset.csv. The dataset should contain:
review: Movie review textsentiment: Labels ('positive' or 'negative')
- Clone the repository
git clone https://github.com/yourusername/sentiment-analysis-classifier.git
cd sentiment-analysis-classifier-
Download the IMDB Dataset from Kaggle and place it in the project directory
-
Open and run the Jupyter notebook
jupyter notebook Imdb_review_pred.ipynb-
Data Loading & Exploration
- Load IMDB dataset with 50,000 movie reviews
- Balanced dataset (25,000 positive, 25,000 negative)
-
Text Preprocessing
- Convert to lowercase
- Remove HTML tags and URLs
- Keep only alphabetic characters
- Remove extra whitespace
-
Feature Engineering
- TF-IDF vectorization with max 5,000 features
- Transform text into numerical vectors
-
Data Splitting
- 80% training (40,000 samples)
- 10% validation (5,000 samples)
- 10% testing (5,000 samples)
-
Model Training & Optimization
- Test multiple alpha values [0.01, 0.1, 0.5, 1.0, 2.0]
- Select best performing model on validation set
- Train final model on combined train+validation data
-
Evaluation
- Test set evaluation
- 5-fold cross-validation
- Confusion matrix analysis
# Example predictions
new_reviews = [
"This movie was absolutely amazing, a masterpiece!",
"Terrible film. Waste of time and money.",
"It was okay, not the best but not the worst either."
]
# The model predicts:
# "This movie was absolutely amazing, a masterpiece!" → positive
# "Terrible film. Waste of time and money." → negative
# "It was okay, not the best but not the worst either." → negative- Balanced Performance: The model performs equally well on both positive and negative reviews
- Robust Generalization: Cross-validation shows consistent performance across different data splits
- Effective Preprocessing: Simple text cleaning techniques prove sufficient for good performance
- Optimal Hyperparameters: Alpha = 2.0 provides the best regularization for this dataset
- Implement more advanced text preprocessing (stemming, lemmatization)
- Experiment with different algorithms (SVM, Random Forest, Neural Networks)
- Add n-gram features (bigrams, trigrams)
- Implement ensemble methods
- Create a web interface for real-time predictions
- Add support for neutral sentiment classification
This project demonstrates:
- Text Preprocessing: Essential NLP techniques for cleaning and preparing text data
- Feature Engineering: Converting text to numerical features using TF-IDF
- Model Selection: Systematic approach to hyperparameter tuning
- Evaluation Methods: Comprehensive model assessment using multiple metrics
- Machine Learning Pipeline: End-to-end implementation from data to deployment-ready model
Feel free to fork this project and submit pull requests for any improvements. Some areas where contributions would be welcome:
- Additional preprocessing techniques
- Alternative algorithms implementation
- Performance optimization
- Documentation improvements
This project is open source and available under the MIT License.
⭐ If you found this project helpful, please give it a star!