Skip to content

Latest commit

Β 

History

8 Commits

Folders and files

Repository files navigation

CI

πŸ›‘οΈ FakeShield AI β€” Fake News Detector

Paste any headline or article. Get an instant REAL / FAKE verdict with confidence score and linguistic evidence.

Python Flask scikit-learn NLTK License: MIT


Misinformation spreads 6x faster than the truth. FakeShield uses NLP and machine learning to classify news articles in milliseconds β€” with a confidence score, fake/real probability breakdown, and the exact words that triggered the verdict.


✨ Features

🧠 Multi-Model ML

  • 4 models compared: Passive Aggressive, Naive Bayes, Logistic Regression, Random Forest
  • Best model auto-selected at training time
  • ~91% accuracy on Kaggle Fake News dataset

πŸ” Explainable AI

  • Confidence score + fake/real probability
  • Key words that influenced the verdict
  • Per-word influence direction (fake vs real)

πŸ“Š Training Visualisations

  • Accuracy comparison across all 4 models
  • Confusion matrix per model
  • ROC curves
  • Top 20 most predictive words

🌐 REST API

  • POST /predict endpoint
  • JSON in, JSON out
  • Drop-in for any app or browser extension

πŸš€ Quick Start

git clone https://github.com/aasimansari1/fake-news-detector.git
cd fake-news-detector

pip install -r requirements.txt

# Generate sample dataset + download NLTK data
python dataset/create_dataset.py

# Train all 4 models (best one saved automatically)
python train_model.py

# Launch the web app
python app.py
# β†’ http://localhost:5000

Want better accuracy? Replace dataset/sample_data.csv with the Kaggle Fake News dataset (~44K articles). Any CSV with text and label columns (REAL/FAKE or 0/1) works.


πŸ—οΈ How It Works

  πŸ“° Article / Headline
          β”‚
          β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚   NLP Pipeline    β”‚
  β”‚  lowercase β†’      β”‚
  β”‚  punctuation β†’    β”‚
  β”‚  tokenise β†’       β”‚
  β”‚  stopwords β†’      β”‚
  β”‚  lemmatise        β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚  TF-IDF (bigrams) β”‚
  β”‚  Vectoriser       β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚  Best ML Model    β”‚  ← Passive Aggressive (~91%)
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
           β”‚
           β–Ό
  βœ… REAL (confidence: 94%)
  ❌ FAKE (confidence: 82%)
  + key words + probabilities

πŸ“‘ API

POST /predict
Content-Type: application/json

{"text": "Scientists discover that vaccines cause autism, government hiding truth"}
{
  "prediction": "FAKE",
  "confidence": 89.4,
  "fake_probability": 89.4,
  "real_probability": 10.6,
  "key_words": [
    {"word": "hiding", "score": 0.21, "influence": "fake"},
    {"word": "truth", "score": 0.18, "influence": "fake"}
  ],
  "model_name": "Passive Aggressive",
  "model_accuracy": 0.907
}

πŸ“ˆ Model Accuracy

Model Accuracy Notes
Passive Aggressive ~91% Best overall β€” auto-selected
Naive Bayes ~88% Fastest inference
Logistic Regression ~86% Most interpretable
Random Forest ~84% Most robust to noise

Accuracy on sample dataset. Use a larger real-world dataset for production-grade results.


πŸ—‚οΈ Project Structure

fake-news-detector/
β”œβ”€β”€ app.py                        # Flask web app + /predict API
β”œβ”€β”€ train_model.py                # Training pipeline, model comparison
β”œβ”€β”€ requirements.txt
β”‚
β”œβ”€β”€ src/                          # Reusable data science module
β”‚   β”œβ”€β”€ preprocessing.py          # Text cleaning, label normalisation, dataset loader
β”‚   β”œβ”€β”€ features.py               # TF-IDF vectoriser builder
β”‚   └── models.py                 # Train all classifiers, pick best, ROC data
β”‚
β”œβ”€β”€ notebooks/                    # Jupyter data science pipeline
β”‚   β”œβ”€β”€ 01_EDA.ipynb              # Exploratory data analysis
β”‚   β”œβ”€β”€ 02_Feature_Engineering.ipynb  # Preprocessing + TF-IDF
β”‚   β”œβ”€β”€ 03_Model_Training.ipynb   # Train & compare 4 models
β”‚   └── 04_Model_Evaluation.ipynb # Confusion matrix, ROC, feature importance, CV
β”‚
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ raw/                      # Source datasets (CSV)
β”‚   └── processed/                # Train/test splits + vectoriser (generated)
β”‚
β”œβ”€β”€ models/                       # Saved model + vectoriser (generated)
β”œβ”€β”€ reports/figures/              # All visualisation outputs (generated)
β”‚
β”œβ”€β”€ dataset/
β”‚   └── create_dataset.py         # Sample dataset generator
β”œβ”€β”€ static/
β”‚   β”œβ”€β”€ css/style.css
β”‚   β”œβ”€β”€ js/main.js
β”‚   └── images/
└── templates/
    └── index.html

Running the Notebooks

pip install -r requirements.txt
jupyter notebook

Run in order: 01_EDA β†’ 02_Feature_Engineering β†’ 03_Model_Training β†’ 04_Model_Evaluation


πŸ› οΈ Tech Stack

Layer Technology
Backend Python, Flask, Gunicorn
ML / NLP scikit-learn, NLTK, pandas, numpy
Visualisation matplotlib, seaborn
Frontend Vanilla HTML/CSS/JS, dark mode
Server Nginx + systemd

🀝 Contributing

Ideas for contributions:

  • 🌐 Browser extension that checks articles in-page
  • πŸ€— Fine-tune a BERT/RoBERTa model for better accuracy
  • πŸ“± Mobile-friendly UI improvements
  • πŸ—ƒοΈ Support for more languages
git checkout -b feature/your-feature
git commit -m 'Add your feature'
git push origin feature/your-feature
# Open a Pull Request

πŸ“„ License

MIT Β© Mohd Aasim Ansari


Fighting misinformation one article at a time. If this helped, please ⭐ star the repo!

Releases

Packages

Contributors

Languages