Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

🎬 IMDB Movie Review Sentiment Analysis Classifier

A machine learning project that classifies movie reviews as positive or negative using Natural Language Processing (NLP) techniques and a Multinomial Naive Bayes classifier.

📊 Project Overview

This project implements a sentiment analysis classifier that can automatically determine whether a movie review expresses positive or negative sentiment. The model is trained on the IMDB Movie Reviews Dataset and achieves 85.4% accuracy on the test set.

🔍 Key Features

  • Comprehensive Text Preprocessing: HTML tag removal, URL cleaning, case normalization, and special character handling
  • TF-IDF Vectorization: Converts text into numerical features for machine learning
  • Hyperparameter Optimization: Uses validation set to find optimal alpha parameter for Naive Bayes
  • Robust Evaluation: Includes train/validation/test split, K-fold cross-validation, and confusion matrix analysis
  • Real-time Predictions: Can classify new movie reviews instantly

📈 Model Performance

  • Test Accuracy: 85.42%
  • Cross-Validation Accuracy: 84.89% (±0.002)
  • Balanced Performance: Equal precision and recall for both positive and negative classes
  • Model Type: Multinomial Naive Bayes with α = 2.0

Performance Metrics

              precision    recall  f1-score   support

    negative       0.85      0.85      0.85      2500
    positive       0.85      0.85      0.85      2500

    accuracy                           0.85      5000
   macro avg       0.85      0.85      0.85      5000
weighted avg       0.85      0.85      0.85      5000

🛠️ Technology Stack

  • Python 3.x
  • scikit-learn: Machine learning algorithms and evaluation metrics
  • pandas: Data manipulation and analysis
  • numpy: Numerical computing
  • matplotlib: Data visualization
  • re: Regular expressions for text preprocessing

📁 Project Structure

sentiment-analysis-classifier/
│
├── Imdb_review_pred.ipynb    # Main Jupyter notebook with complete pipeline
├── README.md                 # Project documentation
└── IMDB Dataset.csv          # Dataset (not included in repo)

🚀 Getting Started

Prerequisites

pip install pandas numpy scikit-learn matplotlib re

Dataset

This project uses the IMDB Movie Reviews Dataset. You'll need to download it and place it in the project directory as IMDB Dataset.csv. The dataset should contain:

  • review: Movie review text
  • sentiment: Labels ('positive' or 'negative')

Running the Code

  1. Clone the repository
git clone https://github.com/yourusername/sentiment-analysis-classifier.git
cd sentiment-analysis-classifier
  1. Download the IMDB Dataset from Kaggle and place it in the project directory

  2. Open and run the Jupyter notebook

jupyter notebook Imdb_review_pred.ipynb

🔄 Model Pipeline

  1. Data Loading & Exploration

    • Load IMDB dataset with 50,000 movie reviews
    • Balanced dataset (25,000 positive, 25,000 negative)
  2. Text Preprocessing

    • Convert to lowercase
    • Remove HTML tags and URLs
    • Keep only alphabetic characters
    • Remove extra whitespace
  3. Feature Engineering

    • TF-IDF vectorization with max 5,000 features
    • Transform text into numerical vectors
  4. Data Splitting

    • 80% training (40,000 samples)
    • 10% validation (5,000 samples)
    • 10% testing (5,000 samples)
  5. Model Training & Optimization

    • Test multiple alpha values [0.01, 0.1, 0.5, 1.0, 2.0]
    • Select best performing model on validation set
    • Train final model on combined train+validation data
  6. Evaluation

    • Test set evaluation
    • 5-fold cross-validation
    • Confusion matrix analysis

💡 Usage Example

# Example predictions
new_reviews = [
    "This movie was absolutely amazing, a masterpiece!",
    "Terrible film. Waste of time and money.",
    "It was okay, not the best but not the worst either."
]

# The model predicts:
# "This movie was absolutely amazing, a masterpiece!" → positive
# "Terrible film. Waste of time and money." → negative  
# "It was okay, not the best but not the worst either." → negative

📊 Key Insights

  • Balanced Performance: The model performs equally well on both positive and negative reviews
  • Robust Generalization: Cross-validation shows consistent performance across different data splits
  • Effective Preprocessing: Simple text cleaning techniques prove sufficient for good performance
  • Optimal Hyperparameters: Alpha = 2.0 provides the best regularization for this dataset

🔮 Future Improvements

  • Implement more advanced text preprocessing (stemming, lemmatization)
  • Experiment with different algorithms (SVM, Random Forest, Neural Networks)
  • Add n-gram features (bigrams, trigrams)
  • Implement ensemble methods
  • Create a web interface for real-time predictions
  • Add support for neutral sentiment classification

📚 Learning Outcomes

This project demonstrates:

  • Text Preprocessing: Essential NLP techniques for cleaning and preparing text data
  • Feature Engineering: Converting text to numerical features using TF-IDF
  • Model Selection: Systematic approach to hyperparameter tuning
  • Evaluation Methods: Comprehensive model assessment using multiple metrics
  • Machine Learning Pipeline: End-to-end implementation from data to deployment-ready model

🤝 Contributing

Feel free to fork this project and submit pull requests for any improvements. Some areas where contributions would be welcome:

  • Additional preprocessing techniques
  • Alternative algorithms implementation
  • Performance optimization
  • Documentation improvements

📄 License

This project is open source and available under the MIT License.


⭐ If you found this project helpful, please give it a star!

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages