Skip to content

Latest commit

ย 

History

62 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

TRUTH-MATRIX ๐Ÿ“ฐ

Truth Matrix is a robust AI-ML model for detecting fake news ๐Ÿ“‰, trained on a comprehensive dataset of global ๐ŸŒ and Indian ๐Ÿ‡ฎ๐Ÿ‡ณ news articles. It leverages a combination of XGBoost ๐Ÿ“ˆ, logistic regression ๐Ÿ”, and Naive Bayes ๐Ÿ“Š, predicting based majority voting. By analyzing text features with a TF-IDF vectorizer ๐Ÿงฉ, it provides reliable classification of news articles to help counteract misinformation ๐Ÿšซ๐Ÿ“ฐ.

Team-Codex

Chetan Sharma : Team Lead and ML developer
Email : chetan.sharma162004@gmail.com

Viraj Singh : ML Developer
Email : anandviraj30@gmail.com

Devansh Tiwari : GitHub Manager and FrontEnd Developer
Email : devanshtiwari2610@gmail.com

Pratham Varma : GitHub Manager and Algoman
Email : prathamvarma178@gmail.com

Model Deployment

Checkout our Website ๐ŸŒ here

-Checkout Google Collab for quick analysis:
-Open file on Colab ๐Ÿ“‚
-Click to Open Statistics of our Model ๐Ÿ“ˆ
-Click to view Performance of different implementations ๐Ÿ“Š

-Checkout the ppt for better understanding : [Click Here]

Dataset

The dataset contains 24,529 news articles with an average word count of 75 (Download the dataset from here)
Word Cloud:
download

Dataset Pre-Processing

1. Import Libraries

You are importing several important libraries at the beginning of the code:

  • pandas, numpy: For data manipulation and numeric operations.
  • re: For regular expressions, used for text cleaning.
  • nltk: Used for natural language processing, especially stopwords removal and stemming.
  • TfidfVectorizer: Converts text to a vector representation based on term frequency-inverse document frequency (TF-IDF).

2. Load Datasets

You are loading two datasets:

  • news_dataset.csv: Contains a labeled dataset of real and fake news (dataset_1).
  • train.csv: Another dataset with text fields (dataset_2).
  • For dataset_1, you replace the labels FAKE with 1 and REAL with 0.
  • For dataset_2, you create a new text column by concatenating the author and title columns and dropping unnecessary columns like id, author, and title.

3. Concatenate Datasets

  • Both datasets are combined into a single DataFrame dataset using pd.concat. This allows you to work with one dataset, combining the real and fake news data.

4. Handle Missing Data

  • isnull().sum() is used to check for missing data in both dataset_1 and dataset_2.
  • You fill any missing values in the combined dataset with a space (' '), ensuring that there are no null values before model training.

5. Stemming Function

  • You define a stemming() function that:
    • Removes non-alphabetical characters.
    • Converts the text to lowercase.
    • Splits the text into words.
    • Removes stopwords (common words like "the", "is").
    • Stems words (e.g., "running" becomes "run") using PorterStemmer.
    • Rejoins the words into a cleaned text string.

6. Apply Stemming

  • You apply the stemming() function to the text column of the dataset to preprocess all the text data.

7. Vectorize Text

  • TfidfVectorizer is used to convert the preprocessed text into a numerical format (TF-IDF vectors) for model training.
  • You fit_transform the vectorizer on the X dataset (the text column), which learns the vocabulary and transforms the text into vectors.

Approach

  1. Data Loading and Preprocessing:

    • Two datasets are loaded: one containing news articles with labels (fake or real) and another with additional text data.
    • Unnecessary columns (like IDs, authors, and titles) are removed from the second dataset.
    • Missing values are filled with empty spaces, and labels are converted to binary (FAKE = 1, REAL = 0).
    • A stemming function is applied to clean and preprocess the text, removing unwanted characters and stopwords, and reducing words to their root form.
    • Both datasets are merged into one and the text is transformed using the TF-IDF vectorizer to convert text into numerical features.
  2. Train-Test Split:

    • The data is split into training (80%) and testing (20%) sets using train_test_split.
  3. Model Training:

    • Three machine learning models are trained on the training set:
      • Logistic Regression
      • XGBoost
      • Naive Bayes
    • Each model is evaluated on the training and testing sets using accuracy_score to calculate prediction accuracy.
  4. Manual Prediction:

    • A function takes a user's inputted article, applies the same preprocessing (stemming and vectorization), and passes it through all three models.
    • The final prediction is based on a majority vote among the three models: if two or more predict "fake news," it is classified as fake, otherwise, it is classified as real.
  5. Output:

    • The program displays the prediction (REAL or FAKE) for the input article and provides the accuracy of the models on the test data.

Workflow

  1. Data Preparation
  2. Pre Processing
  3. Vectorization
  4. Splitting the data
  5. Training The models
  6. Testing with manual Inputs

Modifications

Brief analysis

Train Data accuracy:-
1)Logistics regression : 98.68%
2)XgBoost: 99.25%
3)Naive Bayes: 97.78%

Test Data accuracy:-
1)Logistics regression : 97.92%
2)XgBoost: 98.73%
3)Naive Bayes: 96.35%

Custom Test:- We performed a test consisting of 15 different news articles, and the results for each approach are as follows:
Logistics + XGBoost: 12/15 correct (80% accurate)
Logistics + XGBoost + NB: 14/15 correct (93% accurate)
Logistic Regression: 12/15 correct (80% accurate)

Results

Let's see a demo on our website with a fake sample article:
In June 2020, Baba Ramdev claimed, โ€œWe have prepared the first Ayurvedic-clinically controlled, research, evidence, and trial-based medicine for COVID-19. We conducted a clinical case study and clinical controlled trial, and found that 69% of the patients recovered within three days and 100% recovered within seven days.
image

As predicted we can see that the model correctly gave an output that the news was fake.

Graphs

Confusion matrix:
download
download
download
Model Accuracy:
download
Logistic Regression Learning Curve:
download
Feature Importances of XGBOOST:
image
Real World Performance of our Different Implementations:
newplot

Video Explanation

Check out the Video Here

Disclaimer

This model is trained and tested for recognizing certain words which may indicate that the news is fake and doesn't track the context and the current affairs.

About

A ML model which accurately predicts whether the provided news article is fake or real ๐Ÿ‘‡

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages