Skip to main content

Detecting fake news with the help of the classical machine learning and TF-IDF: Flask Web App

Page 1


International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Detecting fake news with the help of the classical machine learning and TF-IDF: Flask Web App

Nilima Chapke1 , Meenal Gaikwad2

1Assistant Professor, Department of Computer Science, Pune, Maharashtra

2B. Tech Student, Department of Computer Science

Ajeenkya DY Patil University

Lohegaon, Airport Rd, Charholi Budruk, Pune, Maharashtra ***

Abstract - The high rate of development of social media has caused fake news to pose a significant danger to the credibility of the population and democracy. In this paper, a fake news detector with classical machine learning is built on the Kaggle Fake and Real News dataset (44,919 articles) to create an automated detector. The input text is first processed using standard cleaning methods and transformed to TF-IDF features and trained on several different classifiers such as LinearSVC, Logistic Regression and Passive Aggressive. LinearSVC had the highest accuracy of 99.5 percent on the test set and it only took 12 seconds to train on a normal laptop. The entire system comprises of a Flask web application that enables text input, URL scraping, image OCR, and video speech-to-text that is operated using the same training pipeline. This shows real-world implementation other than in experiments. Findings indicate that classical machine learning is still very effective in text classification and it provides a speed and interpretability benefit over deep learning versions. This technique will be immediately applicable to real-world fake news monitoring with the Python implementation being modular and having a live web interface.

Key Words: Fake news, machine learning, text processing, TF-IDF, LinearSVC, Flask app, web deployment.

1. INTRODUCTION

Theaccesstoinformationhaschangedonlinenewsplatformsandsocialmediaplatformswhichhavegiventhebenefitof instantaccessworldwide.Nevertheless,thiseasehascontributedtotherateatwhichfakenewsthatiseitherintentionally falseormisleadingarticlesthatarepackagedasnormalnewshavespread.Thiskindofmisinformationdestroyssocialtrust andposesadangertopoliticalandeconomicstability.

Theproblemwashighlightedintheworldduringthe2016presidentialelectionintheU.S.whensocialmediaexaggeratedfake accountsthatchangedvotingpatterns.AllcottandGentzkowhadrecordedthese,andtheautomateddetectionmethodswere widelystudied[1]

Machinelearning(ML) andnaturallanguageprocessing(NLP)canbeusedtoscalethenewscontentanalysis.Thispaper createsa realistic fakenewsidentifierbasedonclassicalML modelswithTF-IDF features.Inadditiontoanevaluationof classifierperformance,weshowapracticalreliabilitywithaFlaskwebapplication,whichtakestext,URLs,imagesandvideos.

2. RELATED WORK

Variousmethodsofdetectingfakenewshavebeenexploredbyuseoflinguisticfeaturesanalysisuptocomplexdeeplearning scripts. Preliminary research was aimed at deriving deceptive writing patterns based on lexical, syntactic, and semantic featuresobtainedoutofthetext.AstudyfoundbyPerez-Rosasetal.reportedthattextualfeaturesbasedonNLPareapplicable indistinguishingbetweenfakeandrealnewsarticles[2].

Theclassicalmachinelearningtechniqueshavehadalotofcoveragebecauseoftheirefficiencyandinterpretability.Ahmadet al.testedanumberofmachinelearningandensembleclassifierswithfakenewsdetectionandstatedhighperformancewith thetraditionalmodel,includingLogisticRegressionandRandomForest.Themodelsarestillappealingtorealworldsystems duetotheirreducedcostofcomputationandtransparency[3].

Intherecentpast,therehasbeenresearchdoneondeeplearningandtransformer-basedarchitectures,includingBERT,which attaingreataccuracythroughthecontextualinformationthatiscapturedatthetext[5].Yet,theyaremoreexpensiveandless interpretabletocompute.Theclassicalstyleofmachinelearninghasbeenfoundtobecompetitiveinsurveystudies,especially inthosesituationswhenefficiencyandexplainabilityarehighlydemanded[5].Inaddition,explainableAImethodshavebeen offeredtoenhancetrustandtransparencyinfalsifiednewssystemsinvolvingtransformers[4].

3. DATASET DESCRIPTION

TheFakeandRealNewsdataset[6],apubliclyavailablebenchmarkdatasetthatcanbeacquiredatKaggle,isusedinthisstudy. ThedatasetiscomposedofEnglishnewsarticlesthatarelabelledasfakeorrealintheinternationalnewsandisextensively appliedtothestudiesoffakenewsdetection.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

3.1 Dataset Statistics

Thedatasetusedforthisstudycontainsnearly44,000 newsarticlescollectedfrom reliablesources.Outofthese,23,502 articlesarelabeledasfakenews,while21,417arelabeledasrealnews.Thedistributionbetweenthetwoclassesisfairly balanced,whichmakesthedatasetsuitablefortrainingandevaluatingclassificationmodelswithoutsignificantbiasforward onecategory.

3.2 Dataset Features

Each news article in the dataset includes attributes, such as the title, the full news content, the subject category and the publicationdata.Thesefeaturesprovidebothtextualandcontextualinformationthatcansupporteffectiveandclassification. The dataset is large enough and reasonably balanced, which is why it can be used in the supervised machine learning experimentsandcomparativeanalysisofseveralclassifiers.

4. METHODOLOGY

Thefakenewsdetectorthatisproposedisbasedonasystematicandstructuredmachinelearningpipelinewithaspecifictext classificationproblem.Thismethodologyaimsatturningtheunstructuredtextualnewsinformationintoausefulformatand feedmachinelearningmodelswiththisinformationtoclassifynewsarticlesasfakeorreal.

The process of the work is divided into four major steps: text preprocessing, feature extraction, model training, and performanceevaluation.Thismodularityenhancesreproducibilityandenableseachofthestagestobeoptimizedandanalyzed separately.

Figure 1 DistributionofFakeandRealNewsArticlesintheDataset
Figure 2 DistributionofNewsArticleLengths

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

4.1 Text Preprocessing

Preprocessingoftextsisaveryimportantphaseinnaturallanguageprocessingbecauseuntidynewsarticlesusuallyhave noisesandinconsistencythatmayinfluencemodelperformancenegatively.Variouspreprocessingmethodscouldbeusedto cleanandnormalizethetextualdatainthisstudy,andthenfeatureextractionisperformed. Each and every piece of text is converted to lower case so as to maintain a consistent style and not to consider different versionsofthesamewordasdistinctcharacteristics.Punctuations,digits,andspecialcharactersareeliminated,sincetheydo notcarryanysignificantsemanticcontenttoclassifyfakenews. Stopwordsthatarecommonlike‘the’,‘is’,‘and’,‘are ’ are commonly removed to minimize the dimensions and concentrate on informative words. Lastly, additional whitespace is normalized,sothatdocumentsareconsistent.

Thesepreprocessingstepsaimtoeliminateredundantinformationandmaintainsignificantlinguisticpatternsthatdistinguish fakenewsandtherealnews.

4.2 Feature Extraction

The cleansed textual data are then pre-processed and converted to the numerical feature vectors with the help of Term Frequency-Inverse Document Frequency (TF-IDF) method. TF-IDF is a popular feature extraction technique in text classificationbecauseitgivesmoreweighttothewordswhicharecommoninadocumentbutarenotverycommoninthe overallcorpus.

Theschemeofweightedwordtallyingisespeciallyusefulwhenitcomestofakenews detecting,andtherearewordsand phraseswithahigherpropensitytobeusedindishonestorsensationalnews.Usingsuchdiscriminativewords,TF-IDFhelps machinelearningmodelstoimprovetheseparationoffakeandrealnewsarticles.

Togainadditionalinsightsintothetextualpropertiesofthedata,suchexploratoryanalysesasawordcloudareutilized.These visualizationsdemonstratethatitispossibletoidentifysomevisibledifferencesintheprevalenttermsoffakeandrealnews, and the hypothesis that linguistic patterns can be highly relevant when classifying fake news can be proven by these visualizations.

Themodeltrainingandevaluationinvolveareviewoftrainingresults,trainingprocess,andfuturetrainingrequirements.

Figure 3 WordCloudRepresentationofFrequentlyOccurringTermsinFakeNewsArticles
Figure 4 WordCloudRepresentationofFrequentlyOccurringTermsinFake

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

4.3 Model Training and Evaluation

Themodeltrainingandevaluationwillentailreviewoftrainingoutcomes,trainingprocess,andtrainingrequirementsinthe future.

MultipleclassicalmachinelearningclassifiersaretrainedusingtheTF-IDFfeaturevectors.Inordertohaveagoodassessment and maintain the initial distribution of classes, the sample is separated into training and testing sets through a stratified sampling strategy. All models are trained using the training data and tested using the hidden test data. The standard classificationmeasuressuchasaccuracy,precision,recall,F1-score,andReceiverOperatingCharacteristic-AreaUnderthe Curve(ROC-AUC)areusedtomeasureperformance.Thesemeasuresareabletoofferaholisticstudyofthemodelbehavior, especiallyasregardstheabilitytocorrectlydenotefakenewsandreducingfalseidentifications.

5. SYSTEM IMPLEMENTATION

ItiswritteninPythonandstandardlibrariessuchasscikit-learnasamachinelearningsystem,pandasasadataprocessing library,NLTKasatextprocessinglibrary,andFlaskasawebframework.Codeissplitintoindividualfilesthatpreprocessing.py cleansthetext,trainmodel.pytrainsthemosteffectivemodel,evaluatemodel.pyproducesreportsandplotsonitsperformance and app.py is the web server. This arrangement allows me to debug one step at a time and debug other parts without destroyingalltheothers.

Weanalyzedthedatabeforetraining. Figure 1 alsoverifiedthattherewasagoodbalancebetweenfake(23k)andreal(21k) articles and thus it did not require any special balancing tricks. The lengths of the articles were diverse some with short headlines,otherswithlongreportswhichisidealwithTF-IDFbecauseitisabletoworkwithsizes.Wealsogeneratedword cloudswiththefakenewsutilizingwordswithmoredramaticuse,whereastherealnewsutilizestermsofpolicy. Therealhighlightiswebapplication.Theusercandirectlypastesometext,provideanewsURL(andthefeatureautoscrapes), hecanuploadanimage(OCRreadsanytext),andhecanevenavideofile(extractsaudioandconvertsspeechtotext).Allthe stepsarethesameasthoseoftrainingdata,loadsthesavedLinearSVCmodel,anddeliverspredictionswithconfidencescores. Processingis1-4secondsbasedonthetypeofinput,andisfastenoughtooperatenormally. Figure 5,6,7,8 representsthe workinginterfaceandsamplepredictions.

Figure 5 SystemsummarypageexplainingtheAIpipelinefrominputtoLinearSVCclassification

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

RESULTS AND DISCUSSION

Table 1 PerformanceComparisonofMachineLearningModels

Table 2 ClassificationReportoftheFinalModel

Figure 7 Livepredictionresult
Figure 8 Internationalnewssiteslink

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

LinearSVChadthehighestaccuracyof99.5on8,984testsamples.Table1isacomparisonofallmodelsthathavebeentested LinearSVCcomesoutofthegrounddecisivelyoverLogisticRegression(98.7%),LogisticRegression(98.8%),andothermodels intermsofaccuracy,precision,andF1-score.ThisvalidatesthefactthatTF-IDFiseffectivewhendealingwithnewstext. TheconfusionmatrixinFigure9narratestheentiretruth:outof4,696realfakearticles,itidentified4,491(only205were wronglylabeled).Among4,288realarticles,15weremistakenlylabeledasfake.Thetotalerrorwasonly220outofalmost 9,000excellentperformances.Theminorerrorsoccurredwhenactualnewsemployedsensationallanguagethatreadlike clickbait,whichisunderstandable.

Table2breaksitdownbyclass.Thedetectionoffakenewswas99.6%precise(infrequentlyfalselabelsgenuinestories)and 95.6%recall(almosteverythingiscaptured).Precisionofrealnewswasvirtuallyperfectat99.6.Figure10indicatesthatthe ROCcurvehasAUC=0.997,whichimpliesthatitisalsorobusttochangesinthedecisionthreshold. These numbers are superior to Ahmad et al. 96% but these are also much faster to train than other papers. The web deploymentdemonstratesthatitisnotjustlimitedtonotebooksthatanyonecanputittothetest.Lesssignificantconstraints includeuseoftextonlyanduseoftextpatternsonlywithoutimageriesorsocialcues.

Figure 9 ConfusionMatrixoftheFinalClassificationModel
Figure 10 ROCCurveoftheFinalModelShowingAreaUndertheCurve(AUC)

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

7. CONCLUSION AND FUTURE WORK

Accordingtothisproject,afullfakenewsdetectorsystemwasdevelopedbasedondata,toalivewebapplication,withan accuracyof99.5%onreal-worldnewsdataonlybyusingaregularlaptop.TheFlaskinterfacesupportstext,weblinks,images andvideoswhichareratherunfamiliartomostacademicpapers.

Lessonslearned:evenbasicclassicalmachinelearningisnotbadintermsof performance(evenwithoutGPUsorhoursof optimization).Thisisusableinrealpracticeratherthanexperimentsduetothemodularcodeandwebdeployment.

Toimprovefuturework,itwouldbepossibletoaddtoolstoexplainthedatasuchasSHAPthatwoulddemonstratewhat specificwordscontributedtoeachprediction.Itcanbeenhancedbytestingotherlanguagesormixinglikes,sharesorother socialmediaindicators.EasyDockerpackwouldallownewssitestorunitwithoutproblems.Ingeneral,thisdemonstratesthat fakenewsdetectiondoesnotrequiredeeplearningthatisdifficulttoimplement.

REFERENCES

[1]. H.AllcottandM.Gentzkow,"SocialMediaandFakeNewsinthe2016Election,"JournalofEconomicPerspectives,vol.31, no.2,pp.211-236,2017.

[2]. V.Perez-Rosas,B.Kleinberg,A.Lefevre,andR.Mihalcea,"AutomaticDetectionofFakeNews,"inProc.27thInt.Conf. ComputationalLinguistics(COLING),SantaFe,NM,USA,2018,pp.3391-3401.

[3]. Ahmad, M. Yousaf, S. Yousaf, and M. O. Ahmad, "Fake News Detection Using Machine Learning Ensemble Methods," Complexity,vol.2020,Art.no.8885861,2020.

[4]. M.Berrondo-OterminandA.Sarasa-Cabezuelo,"ApplicationofArtificialIntelligenceTechniquestoDetectFakeNews:A Review,"Electronics,vol.12,no.5,Art.no.1154,2023.

[5]. P.Akhtaretal.,"DetectingFakeNewsandDisinformationUsingArtificialIntelligenceandMachineLearning,"Annalsof OperationsResearch,vol.327,pp.633-657,2023.

[6]. M.Szczepanski,M.Pawlicki,R.Kozik,andM.Choras,"NewexplainabilitymethodforBERT-basedmodelinfakenews detection,"ScientificReports,vol.11,Art.no.18262,2021.

[7]. Kaggle,"FakeandRealNewsDataset,"2020.[Online].Available: https://www.kaggle.com/datasets/clmentbisaillon/fake-and-real-news-dataset

Turn static files into dynamic content formats.

Create a flipbook
Detecting fake news with the help of the classical machine learning and TF-IDF: Flask Web App by IRJET Journal - Issuu