Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Sindhi Sentiment Analysis — v2 (Expanded Dataset)

Classical ML sentiment classifier for Sindhi text (Positive / Negative / Neutral), retrained on an expanded, fully human-labeled dataset.

Live demo: https://huggingface.co/spaces/alinawazmahar/sindhi-sentiment

What changed in this update

v1 (original) v2 (this update)
Dataset size 1,909 sentences 4,420 sentences
Labeling Semi-supervised: ~848 hand-labeled + ~1,061 pseudo-labeled from news (confidence ≥ 0.70–0.72) Fully human-labeled, no pseudo-labeling
Best model Logistic Regression LinearSVC (Logistic Regression also retrained, slightly behind)
Test accuracy 94.8% 91.7%
Macro F1 — 0.918

Why accuracy went down, and why that's the more honest number

The v1 dataset's test set included pseudo-labeled news sentences that were filtered to a confidence threshold — meaning the evaluation set itself was biased toward "easier," more clear-cut examples a prior model was already confident about. v2 removes that filter entirely: every sentence is human-labeled, and the dataset is more linguistically diverse as a result.

The drop from 94.8% to 91.7% reflects evaluation on a harder, more representative set — not a regression in the model. As supporting evidence: Neutral, historically the hardest class to classify, has the highest per-class F1 in v2 (0.932), suggesting the larger dataset specifically helped the model generalize better on the cases that matter most.

Class Precision Recall F1
Negative 0.92 0.92 0.918
Neutral 0.92 0.95 0.932
Positive 0.91 0.89 0.902

Pipeline

  1. Preprocessing — Unicode NFC normalization, diacritic removal, control character stripping
  2. Feature engineering — Dual TF-IDF:
    • Character n-grams (2–6), 25,000 features — captures Sindhi morphology (root + suffix patterns)
    • Word n-grams (1–2), 25,000 features — captures phrases and collocations
    • Concatenated via scipy.sparse.hstack → 50,000-dim feature space
  3. Models trained — Logistic Regression, LinearSVC (wrapped in CalibratedClassifierCV for probability outputs), Complement Naive Bayes
  4. Hyperparameter tuning — GridSearchCV, 5-fold stratified CV, optimizing accuracy
  5. Evaluation — held-out 20% test split (stratified), never touched during tuning

Dataset

sindhi_sentiment_v5_final.csv — 4,420 sentences, balanced across 3 classes:

Class Count
Positive 1,501
Negative 1,500
Neutral 1,419

Split: 70% train / 10% validation / 20% test (stratified).

Usage

pip install scikit-learn==1.7.2 joblib scipy numpy pandas
python train.py

Outputs (in models/sentiment/):

  • sentiment_model.joblib — best model (LinearSVC) + both TF-IDF vectorizers + label map
  • all_models_ensemble.joblib — all 3 tuned models, for ensemble/majority-vote use cases
  • training_report.json — full metrics (CV scores, test accuracy, per-class F1, confusion matrices)

Inference example

import joblib
from scipy.sparse import hstack

bundle = joblib.load("models/sentiment/sentiment_model.joblib")
clf, char_vec, word_vec, label_map = bundle["clf"], bundle["char_vec"], bundle["word_vec"], bundle["label_map_inv"]

text = "اڄ جو ڏينهن تمام سٺو آهي"
X = hstack([char_vec.transform([text]), word_vec.transform([text])])
pred = clf.predict(X)[0]
print(label_map[pred])  # Positive

Limitations

  • Classical ML (TF-IDF + linear classifiers), not a transformer — fast and interpretable, but lacks deep contextual understanding
  • Trained on news-corpus-style text (Kawish, AwamiAwaz) and generated/augmented sentences — may not generalize as well to social media or conversational Sindhi
  • No sarcasm or implicit sentiment handling

Author

Ali Nawaz · 24-BSCS-06 · Shaikh Ayaz University Kaggle · HuggingFace · GitHub

About

Sentiment classifier for Sindhi text (dual TF-IDF + LinearSVC) — v2: 91.7% acc / 0.918 macro F1 on a fully human-labeled 4,420-sentence dataset. Includes a transparent writeup of why accuracy dropped from v1's biased eval set.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages