Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

ย 

History

6 Commits
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿ“ง Spam Mail Classifier

Open In Colab Python Streamlit scikit-learn License

๐Ÿ›ก๏ธ An intelligent Machine Learning application that classifies emails as Spam or Ham (Not Spam) using Natural Language Processing and Logistic Regression.


๐Ÿ“‘ Table of Contents


๐ŸŽฏ Overview

Spam emails are a constant nuisance, cluttering inboxes and sometimes posing security threats. This project implements a Machine Learning-based Spam Classifier that can automatically detect and filter spam emails with high accuracy.

Why This Project?

  • ๐Ÿ“ฌ Email Overload: Over 45% of all emails sent daily are spam
  • ๐Ÿ”’ Security: Spam often contains phishing attempts and malware
  • โฑ๏ธ Time Saving: Automated filtering saves hours of manual sorting
  • ๐Ÿค– AI-Powered: Uses NLP techniques for intelligent classification

๐ŸŽฌ Demo

Web Application Interface

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚             ๐Ÿ“ง Spam Email Classifier                        โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                             โ”‚
โ”‚  Enter your email message here:                             โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚ Congratulations! You've won $1,000,000! Click here  โ”‚   โ”‚
โ”‚  โ”‚ to claim your prize now. Limited time offer!!!      โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ”‚                                                             โ”‚
โ”‚               [ ๐Ÿ” Predict ]                                โ”‚
โ”‚                                                             โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   โ”‚
โ”‚  โ”‚  ๐Ÿšซ Spam Email Detected!                            โ”‚   โ”‚
โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚
โ”‚                                                             โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

โœจ Features

Feature Description
๐Ÿ”ค Text Processing TF-IDF vectorization for feature extraction
๐Ÿงน Data Cleaning Removes duplicates and handles null values
โš–๏ธ Balanced Dataset Handles class imbalance through sampling
๐ŸŒ Web Interface Interactive Streamlit app for real-time predictions
๐Ÿ“Š Visualization Decision boundary plots using PCA
๐Ÿ’พ Model Persistence Saved model for easy deployment

๐Ÿ”ฌ How It Works

                         โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                         โ”‚   Input Email    โ”‚
                         โ”‚    Message       โ”‚
                         โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                  โ”‚
                                  โ–ผ
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚    Text Preprocessing   โ”‚
                    โ”‚  โ€ข Lowercase conversion โ”‚
                    โ”‚  โ€ข Stop words removal   โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ”‚
                                 โ–ผ
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚   TF-IDF Vectorization  โ”‚
                    โ”‚  Transform text to      โ”‚
                    โ”‚  numerical features     โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ”‚
                                 โ–ผ
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚  Logistic Regression    โ”‚
                    โ”‚     Classification      โ”‚
                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                                 โ”‚
                    โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
                    โ”‚                         โ”‚
                    โ–ผ                         โ–ผ
           โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”          โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
           โ”‚  ๐Ÿšซ SPAM     โ”‚          โ”‚  โœ… HAM      โ”‚
           โ”‚  (Class 0)   โ”‚          โ”‚  (Class 1)   โ”‚
           โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜          โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ“Š Dataset

The model is trained on a mail dataset containing email messages labeled as spam or ham.

Data Processing Steps:

  1. Load Data - Read CSV file with email messages
  2. Handle Nulls - Check and manage missing values
  3. Remove Duplicates - Clean duplicate entries
  4. Label Encoding - Convert categories (Spam=0, Ham=1)
  5. Balance Classes - Sample to handle class imbalance

Label Distribution:

Category Label Description
Ham 1 Legitimate email (Not Spam)
Spam 0 Unwanted/Junk email

๐Ÿ—๏ธ Model Pipeline

1. Feature Extraction - TF-IDF Vectorizer

TfidfVectorizer(
    min_df=1,           # Minimum document frequency
    stop_words='english', # Remove common English words
    lowercase=True       # Convert to lowercase
)

What is TF-IDF?

  • TF (Term Frequency): How often a word appears in a document
  • IDF (Inverse Document Frequency): How important a word is across all documents
  • TF-IDF: Combines both to give weight to important words

2. Classification - Logistic Regression

A linear classifier that predicts the probability of an email being spam or ham based on the TF-IDF features.


๐Ÿš€ Installation

Prerequisites

  • Python 3.8 or higher
  • pip package manager

Quick Start

  1. Clone the repository

    git clone https://github.com/Lakshya2031/Spam_mail_detector_2.git
    cd Spam_mail_detector_2
  2. Create virtual environment (recommended)

    python -m venv venv
    
    # Windows
    venv\Scripts\activate
    
    # macOS/Linux
    source venv/bin/activate
  3. Install dependencies

    pip install -r requirements.txt

๐Ÿ’ป Usage

๐ŸŒ Run the Web Application

streamlit run "app_py (4).py"

Access the app at http://localhost:8501

๐Ÿ““ Explore the Notebook

Option 1: Open in Google Colab (Recommended)

  • Click the "Open in Colab" badge at the top

Option 2: Run locally

jupyter notebook Mail_Detection.ipynb

๐Ÿงช Quick Test

import pickle

# Load model and vectorizer
model = pickle.load(open('spam_model1.pkl', 'rb'))
vectorizer = pickle.load(open('feature_extraction1.pkl', 'rb'))

# Test message
message = "Congratulations! You won a free iPhone! Click here now!"
prediction = model.predict(vectorizer.transform([message]))

print("Spam" if prediction[0] == 0 else "Ham")

๐Ÿ“ˆ Model Performance

Metric Value
Algorithm Logistic Regression
Feature Extraction TF-IDF Vectorizer
Test Accuracy ~96%
Train-Test Split 80-20
Random State 3

Visualization

The notebook includes a Decision Boundary Plot using PCA dimensionality reduction to visualize how the model separates spam from ham emails.


๐Ÿ“ Project Structure

Spam_mail_detector_2/
โ”‚
โ”œโ”€โ”€ ๐Ÿ““ Mail_Detection.ipynb      # Jupyter notebook with full analysis
โ”œโ”€โ”€ ๐Ÿ app_py (4).py             # Streamlit web application
โ”œโ”€โ”€ ๐Ÿ“ฆ spam_model1.pkl           # Trained Logistic Regression model
โ”œโ”€โ”€ ๐Ÿ“ฆ feature_extraction1.pkl   # Fitted TF-IDF Vectorizer
โ”œโ”€โ”€ ๐Ÿ“‹ requirements.txt          # Python dependencies
โ””โ”€โ”€ ๐Ÿ“– readme.md                 # Project documentation

๐Ÿ› ๏ธ Technologies Used

Category Technology Purpose
Language Python 3.8+ Core programming
Data Processing Pandas, NumPy Data manipulation
NLP TF-IDF Vectorizer Text feature extraction
Machine Learning Scikit-learn Model training & evaluation
Visualization Matplotlib, mlxtend Plotting decision boundaries
Web Framework Streamlit Interactive web app
Serialization Pickle Model persistence

๐Ÿ”ฎ Future Improvements

  • Add more ML algorithms (Naive Bayes, SVM, Random Forest)
  • Implement deep learning with LSTM/Transformers
  • Add email header analysis
  • Create browser extension
  • Deploy on cloud (Heroku/AWS/GCP)
  • Add batch email processing
  • Implement model retraining pipeline

๐Ÿค Contributing

Contributions make the open-source community amazing! Here's how to contribute:

  1. Fork the repository
  2. Create your feature branch
    git checkout -b feature/AmazingFeature
  3. Commit your changes
    git commit -m 'Add some AmazingFeature'
  4. Push to the branch
    git push origin feature/AmazingFeature
  5. Open a Pull Request

๐Ÿ“œ License

This project is licensed under the MIT License - see the LICENSE file for details.


๐Ÿ‘ค Author

Lakshya

GitHub


โญ Support

If you found this project helpful, please consider giving it a โญ on GitHub!


๐Ÿ“ง Contact

Have questions or suggestions? Feel free to:


๐Ÿ›ก๏ธ Stay Safe from Spam with AI! ๐Ÿ›ก๏ธ

Made with โค๏ธ using Machine Learning

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages