Skip to content

About

Supervised NLP pipeline categorising job descriptions with TF-IDF and Naive Bayes

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Job Posting Classifier

Supervised NLP pipeline categorising job descriptions with TF-IDF and Naive Bayes

Domain Task Classes

Python scikit--learn pandas TF--IDF Naive_Bayes


Overview

A supervised text classification pipeline that reads free-text job descriptions and assigns each to one of six employment sectors: IT, Healthcare, Marketing, Finance, Sales and Admin.

The model pairing is deliberate. Multinomial Naive Bayes over TF-IDF features is the canonical baseline for sparse high-dimensional text, and it is genuinely hard to beat on this kind of task without substantially more data — it trains in seconds, needs no GPU, and its conditional independence assumption holds up unusually well for bag-of-words document classification.

Domain & Techniques

Layer Implementation
Feature Extraction TF-IDF vectorisation capped at 5,000 features with English stop-word removal, down-weighting terms common across every posting
Classifier Multinomial Naive Bayes — the standard generative baseline for discrete count/frequency features
Data Handling pandas ingestion with null-dropping and restriction to the six labelled sectors
Validation 80/20 stratifiable train/test split at a fixed random seed for reproducible runs
Evaluation Per-class precision, recall and F1 via classification_report — the right lens when sector classes are imbalanced and raw accuracy would mislead
Similarity Cosine similarity over the TF-IDF space supports nearest-posting retrieval

Pipeline

jobposts.csv
     |
     v
[ pandas: dropna, filter to 6 sectors ]
     |
     v
[ train_test_split  80 / 20, seed=42 ]
     |
     +-----------------------------+
     |                             |
     v                             v
TfidfVectorizer.fit_transform   .transform
(max_features=5000,             (test set — fit only on train,
 stop_words='english')           so no leakage)
     |                             |
     v                             |
MultinomialNB.fit  <---------------+
     |
     v
predict --> classification_report (precision / recall / F1 per sector)

Dataset

The pipeline expects jobposts.csv in the repository root with at least these columns:

Column Description
job_description Free-text posting body
category Sector label — one of IT, Healthcare, Marketing, Finance, Sales, Admin

Rows with nulls in either column are dropped, and rows outside the six sector labels are filtered out.

Usage

pip install pandas scikit-learn
python main.py

The script prints a per-class precision / recall / F1 report for the held-out test split.

Repository Layout

Path Purpose
main.py End-to-end pipeline — load, vectorise, train, evaluate

Project Status

Implemented: the full supervised pipeline from CSV ingestion through TF-IDF vectorisation, Naive Bayes training, and per-class evaluation, with leakage-safe fitting (the vectoriser is fit on train only).

Roadmap: persist the fitted vectoriser and model with joblib so inference does not require retraining; add cross-validation in place of the single split; and benchmark a linear SVM and a transformer embedding baseline against the Naive Bayes reference to quantify what the extra complexity actually buys on this dataset.


Part of the AI + Robotics engineering portfolio of Divyansh Sachdev
90+ national & international competition wins · IIT / NIT / IIIT podiums

About

Supervised NLP pipeline categorising job descriptions with TF-IDF and Naive Bayes

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages