Classical ML sentiment classifier for Sindhi text (Positive / Negative / Neutral), retrained on an expanded, fully human-labeled dataset.
Live demo: https://huggingface.co/spaces/alinawazmahar/sindhi-sentiment
| v1 (original) | v2 (this update) | |
|---|---|---|
| Dataset size | 1,909 sentences | 4,420 sentences |
| Labeling | Semi-supervised: ~848 hand-labeled + ~1,061 pseudo-labeled from news (confidence ≥ 0.70–0.72) | Fully human-labeled, no pseudo-labeling |
| Best model | Logistic Regression | LinearSVC (Logistic Regression also retrained, slightly behind) |
| Test accuracy | 94.8% | 91.7% |
| Macro F1 | — | 0.918 |
The v1 dataset's test set included pseudo-labeled news sentences that were filtered to a confidence threshold — meaning the evaluation set itself was biased toward "easier," more clear-cut examples a prior model was already confident about. v2 removes that filter entirely: every sentence is human-labeled, and the dataset is more linguistically diverse as a result.
The drop from 94.8% to 91.7% reflects evaluation on a harder, more representative set — not a regression in the model. As supporting evidence: Neutral, historically the hardest class to classify, has the highest per-class F1 in v2 (0.932), suggesting the larger dataset specifically helped the model generalize better on the cases that matter most.
| Class | Precision | Recall | F1 |
|---|---|---|---|
| Negative | 0.92 | 0.92 | 0.918 |
| Neutral | 0.92 | 0.95 | 0.932 |
| Positive | 0.91 | 0.89 | 0.902 |
- Preprocessing — Unicode NFC normalization, diacritic removal, control character stripping
- Feature engineering — Dual TF-IDF:
- Character n-grams (2–6), 25,000 features — captures Sindhi morphology (root + suffix patterns)
- Word n-grams (1–2), 25,000 features — captures phrases and collocations
- Concatenated via
scipy.sparse.hstack→ 50,000-dim feature space
- Models trained — Logistic Regression, LinearSVC (wrapped in
CalibratedClassifierCVfor probability outputs), Complement Naive Bayes - Hyperparameter tuning —
GridSearchCV, 5-fold stratified CV, optimizing accuracy - Evaluation — held-out 20% test split (stratified), never touched during tuning
sindhi_sentiment_v5_final.csv — 4,420 sentences, balanced across 3 classes:
| Class | Count |
|---|---|
| Positive | 1,501 |
| Negative | 1,500 |
| Neutral | 1,419 |
Split: 70% train / 10% validation / 20% test (stratified).
pip install scikit-learn==1.7.2 joblib scipy numpy pandas
python train.pyOutputs (in models/sentiment/):
sentiment_model.joblib— best model (LinearSVC) + both TF-IDF vectorizers + label mapall_models_ensemble.joblib— all 3 tuned models, for ensemble/majority-vote use casestraining_report.json— full metrics (CV scores, test accuracy, per-class F1, confusion matrices)
import joblib
from scipy.sparse import hstack
bundle = joblib.load("models/sentiment/sentiment_model.joblib")
clf, char_vec, word_vec, label_map = bundle["clf"], bundle["char_vec"], bundle["word_vec"], bundle["label_map_inv"]
text = "اڄ جو ڏينهن تمام سٺو آهي"
X = hstack([char_vec.transform([text]), word_vec.transform([text])])
pred = clf.predict(X)[0]
print(label_map[pred]) # Positive- Classical ML (TF-IDF + linear classifiers), not a transformer — fast and interpretable, but lacks deep contextual understanding
- Trained on news-corpus-style text (Kawish, AwamiAwaz) and generated/augmented sentences — may not generalize as well to social media or conversational Sindhi
- No sarcasm or implicit sentiment handling
Ali Nawaz · 24-BSCS-06 · Shaikh Ayaz University Kaggle · HuggingFace · GitHub