Misinformation spreads 6x faster than the truth. FakeShield uses NLP and machine learning to classify news articles in milliseconds β with a confidence score, fake/real probability breakdown, and the exact words that triggered the verdict.
|
π§ Multi-Model ML
|
π Explainable AI
|
|
π Training Visualisations
|
π REST API
|
git clone https://github.com/aasimansari1/fake-news-detector.git
cd fake-news-detector
pip install -r requirements.txt
# Generate sample dataset + download NLTK data
python dataset/create_dataset.py
# Train all 4 models (best one saved automatically)
python train_model.py
# Launch the web app
python app.py
# β http://localhost:5000Want better accuracy? Replace
dataset/sample_data.csvwith the Kaggle Fake News dataset (~44K articles). Any CSV withtextandlabelcolumns (REAL/FAKE or 0/1) works.
π° Article / Headline
β
βΌ
βββββββββββββββββββββ
β NLP Pipeline β
β lowercase β β
β punctuation β β
β tokenise β β
β stopwords β β
β lemmatise β
ββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββ
β TF-IDF (bigrams) β
β Vectoriser β
ββββββββββ¬βββββββββββ
β
βΌ
βββββββββββββββββββββ
β Best ML Model β β Passive Aggressive (~91%)
ββββββββββ¬βββββββββββ
β
βΌ
β
REAL (confidence: 94%)
β FAKE (confidence: 82%)
+ key words + probabilities
POST /predict
Content-Type: application/json
{"text": "Scientists discover that vaccines cause autism, government hiding truth"}{
"prediction": "FAKE",
"confidence": 89.4,
"fake_probability": 89.4,
"real_probability": 10.6,
"key_words": [
{"word": "hiding", "score": 0.21, "influence": "fake"},
{"word": "truth", "score": 0.18, "influence": "fake"}
],
"model_name": "Passive Aggressive",
"model_accuracy": 0.907
}| Model | Accuracy | Notes |
|---|---|---|
| Passive Aggressive | ~91% | Best overall β auto-selected |
| Naive Bayes | ~88% | Fastest inference |
| Logistic Regression | ~86% | Most interpretable |
| Random Forest | ~84% | Most robust to noise |
Accuracy on sample dataset. Use a larger real-world dataset for production-grade results.
fake-news-detector/
βββ app.py # Flask web app + /predict API
βββ train_model.py # Training pipeline, model comparison
βββ requirements.txt
β
βββ src/ # Reusable data science module
β βββ preprocessing.py # Text cleaning, label normalisation, dataset loader
β βββ features.py # TF-IDF vectoriser builder
β βββ models.py # Train all classifiers, pick best, ROC data
β
βββ notebooks/ # Jupyter data science pipeline
β βββ 01_EDA.ipynb # Exploratory data analysis
β βββ 02_Feature_Engineering.ipynb # Preprocessing + TF-IDF
β βββ 03_Model_Training.ipynb # Train & compare 4 models
β βββ 04_Model_Evaluation.ipynb # Confusion matrix, ROC, feature importance, CV
β
βββ data/
β βββ raw/ # Source datasets (CSV)
β βββ processed/ # Train/test splits + vectoriser (generated)
β
βββ models/ # Saved model + vectoriser (generated)
βββ reports/figures/ # All visualisation outputs (generated)
β
βββ dataset/
β βββ create_dataset.py # Sample dataset generator
βββ static/
β βββ css/style.css
β βββ js/main.js
β βββ images/
βββ templates/
βββ index.html
pip install -r requirements.txt
jupyter notebookRun in order: 01_EDA β 02_Feature_Engineering β 03_Model_Training β 04_Model_Evaluation
| Layer | Technology |
|---|---|
| Backend | Python, Flask, Gunicorn |
| ML / NLP | scikit-learn, NLTK, pandas, numpy |
| Visualisation | matplotlib, seaborn |
| Frontend | Vanilla HTML/CSS/JS, dark mode |
| Server | Nginx + systemd |
Ideas for contributions:
- π Browser extension that checks articles in-page
- π€ Fine-tune a BERT/RoBERTa model for better accuracy
- π± Mobile-friendly UI improvements
- ποΈ Support for more languages
git checkout -b feature/your-feature
git commit -m 'Add your feature'
git push origin feature/your-feature
# Open a Pull RequestMIT Β© Mohd Aasim Ansari
Fighting misinformation one article at a time. If this helped, please β star the repo!