Supervised NLP pipeline categorising job descriptions with TF-IDF and Naive Bayes
A supervised text classification pipeline that reads free-text job descriptions and assigns each to one of six employment sectors: IT, Healthcare, Marketing, Finance, Sales and Admin.
The model pairing is deliberate. Multinomial Naive Bayes over TF-IDF features is the canonical baseline for sparse high-dimensional text, and it is genuinely hard to beat on this kind of task without substantially more data — it trains in seconds, needs no GPU, and its conditional independence assumption holds up unusually well for bag-of-words document classification.
| Layer | Implementation |
|---|---|
| Feature Extraction | TF-IDF vectorisation capped at 5,000 features with English stop-word removal, down-weighting terms common across every posting |
| Classifier | Multinomial Naive Bayes — the standard generative baseline for discrete count/frequency features |
| Data Handling | pandas ingestion with null-dropping and restriction to the six labelled sectors |
| Validation | 80/20 stratifiable train/test split at a fixed random seed for reproducible runs |
| Evaluation | Per-class precision, recall and F1 via classification_report — the right lens when sector classes are imbalanced and raw accuracy would mislead |
| Similarity | Cosine similarity over the TF-IDF space supports nearest-posting retrieval |
jobposts.csv
|
v
[ pandas: dropna, filter to 6 sectors ]
|
v
[ train_test_split 80 / 20, seed=42 ]
|
+-----------------------------+
| |
v v
TfidfVectorizer.fit_transform .transform
(max_features=5000, (test set — fit only on train,
stop_words='english') so no leakage)
| |
v |
MultinomialNB.fit <---------------+
|
v
predict --> classification_report (precision / recall / F1 per sector)
The pipeline expects jobposts.csv in the repository root with at least these columns:
| Column | Description |
|---|---|
job_description |
Free-text posting body |
category |
Sector label — one of IT, Healthcare, Marketing, Finance, Sales, Admin |
Rows with nulls in either column are dropped, and rows outside the six sector labels are filtered out.
pip install pandas scikit-learn
python main.pyThe script prints a per-class precision / recall / F1 report for the held-out test split.
| Path | Purpose |
|---|---|
main.py |
End-to-end pipeline — load, vectorise, train, evaluate |
Implemented: the full supervised pipeline from CSV ingestion through TF-IDF vectorisation, Naive Bayes training, and per-class evaluation, with leakage-safe fitting (the vectoriser is fit on train only).
Roadmap: persist the fitted vectoriser and model with joblib so inference does not require
retraining; add cross-validation in place of the single split; and benchmark a linear SVM and a
transformer embedding baseline against the Naive Bayes reference to quantify what the extra
complexity actually buys on this dataset.
90+ national & international competition wins · IIT / NIT / IIIT podiums