Truth Matrix is a robust AI-ML model for detecting fake news ๐, trained on a comprehensive dataset of global ๐ and Indian ๐ฎ๐ณ news articles. It leverages a combination of XGBoost ๐, logistic regression ๐, and Naive Bayes ๐, predicting based majority voting. By analyzing text features with a TF-IDF vectorizer ๐งฉ, it provides reliable classification of news articles to help counteract misinformation ๐ซ๐ฐ.
Chetan Sharma : Team Lead and ML developer
Email : chetan.sharma162004@gmail.com
Viraj Singh : ML Developer
Email : anandviraj30@gmail.com
Devansh Tiwari : GitHub Manager and FrontEnd Developer
Email : devanshtiwari2610@gmail.com
Pratham Varma : GitHub Manager and Algoman
Email : prathamvarma178@gmail.com
Checkout our Website ๐ here
-Checkout Google Collab for quick analysis:
-Open file on Colab ๐
-Click to Open Statistics of our Model ๐
-Click to view Performance of different implementations ๐
-Checkout the ppt for better understanding : [Click Here]
The dataset contains 24,529 news articles with an average word count of 75
(Download the dataset from here)
Word Cloud:

You are importing several important libraries at the beginning of the code:
pandas,numpy: For data manipulation and numeric operations.re: For regular expressions, used for text cleaning.nltk: Used for natural language processing, especially stopwords removal and stemming.TfidfVectorizer: Converts text to a vector representation based on term frequency-inverse document frequency (TF-IDF).
You are loading two datasets:
news_dataset.csv: Contains a labeled dataset of real and fake news (dataset_1).train.csv: Another dataset with text fields (dataset_2).- For
dataset_1, you replace the labelsFAKEwith1andREALwith0. - For
dataset_2, you create a newtextcolumn by concatenating theauthorandtitlecolumns and dropping unnecessary columns likeid,author, andtitle.
- Both datasets are combined into a single DataFrame
datasetusingpd.concat. This allows you to work with one dataset, combining the real and fake news data.
isnull().sum()is used to check for missing data in bothdataset_1anddataset_2.- You fill any missing values in the combined
datasetwith a space (' '), ensuring that there are no null values before model training.
- You define a
stemming()function that:- Removes non-alphabetical characters.
- Converts the text to lowercase.
- Splits the text into words.
- Removes stopwords (common words like "the", "is").
- Stems words (e.g., "running" becomes "run") using
PorterStemmer. - Rejoins the words into a cleaned text string.
- You apply the
stemming()function to thetextcolumn of thedatasetto preprocess all the text data.
TfidfVectorizeris used to convert the preprocessed text into a numerical format (TF-IDF vectors) for model training.- You
fit_transformthe vectorizer on theXdataset (the text column), which learns the vocabulary and transforms the text into vectors.
-
Data Loading and Preprocessing:
- Two datasets are loaded: one containing news articles with labels (fake or real) and another with additional text data.
- Unnecessary columns (like IDs, authors, and titles) are removed from the second dataset.
- Missing values are filled with empty spaces, and labels are converted to binary (FAKE = 1, REAL = 0).
- A stemming function is applied to clean and preprocess the text, removing unwanted characters and stopwords, and reducing words to their root form.
- Both datasets are merged into one and the text is transformed using the TF-IDF vectorizer to convert text into numerical features.
-
Train-Test Split:
- The data is split into training (80%) and testing (20%) sets using train_test_split.
-
Model Training:
- Three machine learning models are trained on the training set:
- Logistic Regression
- XGBoost
- Naive Bayes
- Each model is evaluated on the training and testing sets using accuracy_score to calculate prediction accuracy.
- Three machine learning models are trained on the training set:
-
Manual Prediction:
- A function takes a user's inputted article, applies the same preprocessing (stemming and vectorization), and passes it through all three models.
- The final prediction is based on a majority vote among the three models: if two or more predict "fake news," it is classified as fake, otherwise, it is classified as real.
-
Output:
- The program displays the prediction (REAL or FAKE) for the input article and provides the accuracy of the models on the test data.
- Data Preparation
- Pre Processing
- Vectorization
- Splitting the data
- Training The models
- Testing with manual Inputs
Train Data accuracy:-
1)Logistics regression : 98.68%
2)XgBoost: 99.25%
3)Naive Bayes: 97.78%
Test Data accuracy:-
1)Logistics regression : 97.92%
2)XgBoost: 98.73%
3)Naive Bayes: 96.35%
Custom Test:-
We performed a test consisting of 15 different news articles, and the results for each approach are as follows:
Logistics + XGBoost: 12/15 correct (80% accurate)
Logistics + XGBoost + NB: 14/15 correct (93% accurate)
Logistic Regression: 12/15 correct (80% accurate)
Let's see a demo on our website with a fake sample article:
In June 2020, Baba Ramdev claimed, โWe have prepared the first Ayurvedic-clinically controlled, research, evidence, and trial-based medicine for COVID-19. We conducted a clinical case study and clinical controlled trial, and found that 69% of the patients recovered within three days and 100% recovered within seven days.

As predicted we can see that the model correctly gave an output that the news was fake.
Confusion matrix:



Model Accuracy:

Logistic Regression Learning Curve:

Feature Importances of XGBOOST:

Real World Performance of our Different Implementations:

This model is trained and tested for recognizing certain words which may indicate that the news is fake and doesn't track the context and the current affairs.