On-demand social media sentiment analysis powered by a custom BiLSTM + Self-Attention deep learning model
📺 Watch Demo · 📊 Model Details · ⚡ Quick Start
Sentiment Radar is an end-to-end NLP system that scrapes comments from multiple social media platforms on any topic, classifies each comment as positive, neutral, or negative using a custom-trained deep learning model, and presents results in an interactive dark-theme dashboard with rich visualisations.
Unlike generic sentiment APIs, this model was trained specifically on social media language — handling slang, negation ("not bad" → positive), sarcasm markers, and informal text that breaks most off-the-shelf classifiers. The system is on-demand — every analysis is triggered manually by searching a topic, fetching a fresh batch of comments across platforms.
- 🔍 Scrapes on demand — enter any topic, fetch up to 50 comments per platform in seconds
- 🧠 Classifies with a custom model — BiLSTM + Self-Attention trained on 2.86M comments from 6 datasets
- 📊 Visualises everything — donut charts, platform radar, sentiment gap bars, confidence histograms
- 💬 Shows raw comments — filterable by platform and sentiment with confidence scores per prediction
- 📥 Exports results — download full analysis as CSV
| Platform | Auth required | Status |
|---|---|---|
| Optional (OAuth improves freshness) | ✅ Active | |
| HackerNews | None | ✅ Active |
| Dev.to | None | ✅ Active |
| YouTube | Free API key | ✅ Optional |
| Mastodon | None | ✅ Optional |
┌─────────────────────────────────────────────────────────────────────┐
│ SENTIMENT RADAR │
├──────────────────────┬──────────────────────┬───────────────────────┤
│ DATA LAYER │ INFERENCE LAYER │ PRESENTATION LAYER │
│ │ │ │
│ scraper.py │ inference.py │ streamlit_app.py │
│ ┌──────────────┐ │ ┌───────────────┐ │ ┌─────────────────┐ │
│ │ Reddit │ │ │ Text Cleaning │ │ │ KPI Cards │ │
│ │ HackerNews │───▶│ │ Negation Tag │───▶│ │ Donut Charts │ │
│ │ Dev.to │ │ │ Tokenize+Pad │ │ │ Radar Chart │ │
│ │ YouTube │ │ │ BiLSTM Model │ │ │ Gap Analysis │ │
│ │ Mastodon │ │ │ Softmax(3) │ │ │ Comment Cards │ │
│ └──────────────┘ │ └───────────────┘ │ └─────────────────┘ │
└──────────────────────┴──────────────────────┴───────────────────────┘
The core model is a Bidirectional LSTM with Self-Attention and dual pooling, trained from scratch on 2.86M social media comments aggregated from 6 datasets.
Input (seq_len=100)
│
▼
Embedding Layer (vocab=20,000 · dim=128) → 2,560,000 params
│
SpatialDropout1D(0.3)
│
▼
BiLSTM(128 units, return_sequences=True) → 263,168 params
│
BiLSTM(64 units, return_sequences=True) → 164,352 params
│
▼
Self-Attention Layer
│
┌────┴────┐
▼ ▼
GlobalAvg GlobalMax
Pool Pool
└────┬────┘
│ Concat
▼
Dense(128, ReLU, L2) → Dropout(0.5)
Dense(64, ReLU) → Dropout(0.3)
│
▼
Dense(3, Softmax)
[negative · neutral · positive]
Total parameters: ~3,000,000
| Metric | Value |
|---|---|
| Test Accuracy | 76.13% |
| Test Loss | 0.5634 |
| Training Time | ~8,244 seconds |
| Epochs | 12 |
Classification Report:
| Class | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Negative | 0.77 | 0.79 | 0.78 | 239,513 |
| Neutral | 0.63 | 0.62 | 0.62 | 83,142 |
| Positive | 0.80 | 0.78 | 0.79 | 249,108 |
| Weighted Avg | 0.76 | 0.76 | 0.76 | 571,763 |
The neutral class underperforms (F1: 0.62) due to class imbalance — neutral samples (~410K) are significantly fewer than positive/negative (~1.2M+ each). This is a known challenge in social media sentiment datasets.
| Dataset | Source | Size |
|---|---|---|
| Sentiment140 | ~1.6M | |
| Reddit Comments | ~400K | |
| IMDB Reviews | IMDB | ~50K |
| Amazon Reviews | Amazon | ~400K |
| YouTube Comments | YouTube | ~250K |
| HackerNews | HackerNews | ~160K |
| Total | 6 sources | ~2.86M |
Class distribution: Positive ~1.25M · Negative ~1.2M · Neutral ~0.41M
A key feature is the negation scope tagging — applied before tokenisation so the model learns good and good_NEG as distinct signals:
Input: "I didn't enjoy this at all."
│
├─ 1. Expand contractions → "I did not enjoy this at all."
├─ 2. Negation scope tag → "I did not enjoy_NEG this_NEG at_NEG all_NEG."
├─ 3. Strip non-alpha → "I did not enjoy_NEG this_NEG at_NEG all_NEG"
└─ 4. Remove stopwords → "not enjoy_NEG all_NEG"
(keeping: not, but, very, can, won't, never, etc.)
Output: "not enjoy_NEG all_NEG" → Predicted: NEGATIVE (98.93%)
Confidence examples from batch testing:
| Input | Prediction | Confidence |
|---|---|---|
| "This product is absolutely amazing!" | Positive | 91.22% |
| "Worst experience I've ever had." | Negative | 98.57% |
| "I never felt so happy in my life!" | Positive | 86.98% |
| "I don't like this at all." | Negative | 87.22% |
Sentiment_Radar/
│
├── streamlit_app.py ← Main dashboard entry point
│
├── app/
│ ├── __init__.py
│ ├── inference.py ← Model loading, preprocessing, batched prediction
│ └── scraper.py ← Multi-platform comment scraper
│
├── model_files/ ← Place your 3 model artifacts here
│ ├── best_sentiment_model.keras (~35 MB)
│ ├── tokenizerB.joblib (~32 MB)
│ └── encoderB.joblib (~0.5 MB)
│
├── nltk_data/ ← Pre-downloaded NLTK corpora (committed)
│
├── .streamlit/
│ ├── config.toml ← Dark theme configuration
│ └── secrets.toml ← API keys (gitignored — never committed)
│
├── requirements.txt ← Linux / Streamlit Cloud dependencies
├── requirements-windows.txt ← Windows local dev dependencies
├── setup.sh ← Pre-start script for Streamlit Cloud
├── download_nltk.py ← One-time NLTK data download
├── .env.example ← Environment variable template
├── .gitattributes ← Line ending configuration
└── .gitignore
- Python 3.10
- The 3 model files (
best_sentiment_model.keras,tokenizerB.joblib,encoderB.joblib)
git clone https://github.com/BhaveshN1015/Sentiment_Radar.git
cd Sentiment_Radar# Windows
python -m venv venv
venv\Scripts\activate
# Mac / Linux
python -m venv venv
source venv/bin/activate# Windows
pip install -r requirements-windows.txt
# Linux / Mac
pip install -r requirements.txtCopy your 3 model artifacts into model_files/:
model_files/
├── best_sentiment_model.keras (~35 MB)
├── tokenizerB.joblib (~32 MB)
└── encoderB.joblib (~0.5 MB)
copy .env.example .env # Windows
cp .env.example .env # Mac / LinuxEdit .env and add your API keys (all optional — app works without them):
YOUTUBE_API_KEY=AIzaSy... # Enables YouTube platform
REDDIT_CLIENT_ID=abc123 # Improves Reddit results
REDDIT_CLIENT_SECRET=xyz789
REDDIT_USER_AGENT=SentimentRadar/1.0streamlit run streamlit_app.pyOpen http://localhost:8501 in your browser.
The dashboard includes 5 analysis views:
| Tab | What it shows |
|---|---|
| 🌐 Overall | Aggregate sentiment donut + platform grouped bar + radar chart |
| 📊 By Platform | Per-platform donut charts with individual breakdown metrics |
| 💬 Comments | Raw comments filterable by platform and sentiment with confidence scores |
| ⚖️ Sentiment Gap | Positivity score (positive% − negative%) comparison across platforms |
| 🎯 Deep Dive | Confidence distribution histogram + volume vs positivity bubble chart + full summary table |
| Key | Platform | Where to get | Cost |
|---|---|---|---|
YOUTUBE_API_KEY |
YouTube | Google Cloud Console → YouTube Data API v3 → Credentials | Free · 10k units/day |
REDDIT_CLIENT_ID + REDDIT_CLIENT_SECRET |
reddit.com/prefs/apps → Create App → script | Free |
No key needed for: HackerNews, Dev.to, Mastodon, and Reddit (falls back to Pullpush.io archive automatically).
ModuleNotFoundError: No module named 'app'
Run from the project root, not from inside app/:
streamlit run streamlit_app.pyOSError: Unable to open best_sentiment_model.keras
Model files are missing from model_files/. See Step 4 above.
ModuleNotFoundError: No module named 'tf_keras'
pip install tf-keras==2.17.0ModuleNotFoundError: No module named 'keras.preprocessing.text'
Pickle compatibility issue with the tokenizer. Ensure you are using inference.py from this repo — it includes a sys.modules shim that resolves this automatically.
Reddit returns 0 comments
The app automatically falls back to the Pullpush.io archive. If both fail, wait 60 seconds and retry. For consistent fresh results, add Reddit OAuth credentials to .env.
YouTube returns 0 / shows quota error
Free quota is 10,000 units/day (~100 searches/day). Quota resets at midnight Pacific Time. The app shows a warning in the sidebar when quota is exceeded.
LF/CRLF line ending warnings on git add
Run the following to apply the .gitattributes line ending rules:
git rm -r --cached .
git add .
git commit -m "Fix line endings"| Layer | Technology |
|---|---|
| Deep Learning | TensorFlow 2.17.1 + tf-keras 2.17.0 (Keras 2 compatibility layer) |
| NLP | Custom negation-scope preprocessing pipeline + NLTK stopwords |
| ML Utilities | scikit-learn (LabelEncoder, classification metrics) |
| Data Serialisation | joblib (tokenizer + encoder) · Keras H5 (model weights) |
| Dashboard | Streamlit 1.32+ |
| Visualisation | Plotly |
| Data Processing | pandas · numpy |
| Scraping | Python stdlib urllib — no external scraping libraries |
This project is licensed under the MIT License — see the LICENSE file for details.
Built with 🧠 + ☕ by BhaveshN1015
⭐ Star this repo if you found it useful