
International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
![]()

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
PROF.
Nilima Chapke1 , Meenal Gaikwad2
1Assistant Professor, Department of Computer Science, Pune, Maharashtra
2B. Tech Student, Department of Computer Science
Ajeenkya DY Patil University
Lohegaon, Airport Rd, Charholi Budruk, Pune, Maharashtra ***
Abstract - The high rate of development of social media has caused fake news to pose a significant danger to the credibility of the population and democracy. In this paper, a fake news detector with classical machine learning is built on the Kaggle Fake and Real News dataset (44,919 articles) to create an automated detector. The input text is first processed using standard cleaning methods and transformed to TF-IDF features and trained on several different classifiers such as LinearSVC, Logistic Regression and Passive Aggressive. LinearSVC had the highest accuracy of 99.5 percent on the test set and it only took 12 seconds to train on a normal laptop. The entire system comprises of a Flask web application that enables text input, URL scraping, image OCR, and video speech-to-text that is operated using the same training pipeline. This shows real-world implementation other than in experiments. Findings indicate that classical machine learning is still very effective in text classification and it provides a speed and interpretability benefit over deep learning versions. This technique will be immediately applicable to real-world fake news monitoring with the Python implementation being modular and having a live web interface.
Key Words: Fake news, machine learning, text processing, TF-IDF, LinearSVC, Flask app, web deployment.
Theaccesstoinformationhaschangedonlinenewsplatformsandsocialmediaplatformswhichhavegiventhebenefitof instantaccessworldwide.Nevertheless,thiseasehascontributedtotherateatwhichfakenewsthatiseitherintentionally falseormisleadingarticlesthatarepackagedasnormalnewshavespread.Thiskindofmisinformationdestroyssocialtrust andposesadangertopoliticalandeconomicstability.
Theproblemwashighlightedintheworldduringthe2016presidentialelectionintheU.S.whensocialmediaexaggeratedfake accountsthatchangedvotingpatterns.AllcottandGentzkowhadrecordedthese,andtheautomateddetectionmethodswere widelystudied[1]
Machinelearning(ML) andnaturallanguageprocessing(NLP)canbeusedtoscalethenewscontentanalysis.Thispaper createsa realistic fakenewsidentifierbasedonclassicalML modelswithTF-IDF features.Inadditiontoanevaluationof classifierperformance,weshowapracticalreliabilitywithaFlaskwebapplication,whichtakestext,URLs,imagesandvideos.
Variousmethodsofdetectingfakenewshavebeenexploredbyuseoflinguisticfeaturesanalysisuptocomplexdeeplearning scripts. Preliminary research was aimed at deriving deceptive writing patterns based on lexical, syntactic, and semantic featuresobtainedoutofthetext.AstudyfoundbyPerez-Rosasetal.reportedthattextualfeaturesbasedonNLPareapplicable indistinguishingbetweenfakeandrealnewsarticles[2].
Theclassicalmachinelearningtechniqueshavehadalotofcoveragebecauseoftheirefficiencyandinterpretability.Ahmadet al.testedanumberofmachinelearningandensembleclassifierswithfakenewsdetectionandstatedhighperformancewith thetraditionalmodel,includingLogisticRegressionandRandomForest.Themodelsarestillappealingtorealworldsystems duetotheirreducedcostofcomputationandtransparency[3].
Intherecentpast,therehasbeenresearchdoneondeeplearningandtransformer-basedarchitectures,includingBERT,which attaingreataccuracythroughthecontextualinformationthatiscapturedatthetext[5].Yet,theyaremoreexpensiveandless interpretabletocompute.Theclassicalstyleofmachinelearninghasbeenfoundtobecompetitiveinsurveystudies,especially inthosesituationswhenefficiencyandexplainabilityarehighlydemanded[5].Inaddition,explainableAImethodshavebeen offeredtoenhancetrustandtransparencyinfalsifiednewssystemsinvolvingtransformers[4].
TheFakeandRealNewsdataset[6],apubliclyavailablebenchmarkdatasetthatcanbeacquiredatKaggle,isusedinthisstudy. ThedatasetiscomposedofEnglishnewsarticlesthatarelabelledasfakeorrealintheinternationalnewsandisextensively appliedtothestudiesoffakenewsdetection.

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
3.1 Dataset Statistics

Thedatasetusedforthisstudycontainsnearly44,000 newsarticlescollectedfrom reliablesources.Outofthese,23,502 articlesarelabeledasfakenews,while21,417arelabeledasrealnews.Thedistributionbetweenthetwoclassesisfairly balanced,whichmakesthedatasetsuitablefortrainingandevaluatingclassificationmodelswithoutsignificantbiasforward onecategory.
3.2 Dataset Features
Each news article in the dataset includes attributes, such as the title, the full news content, the subject category and the publicationdata.Thesefeaturesprovidebothtextualandcontextualinformationthatcansupporteffectiveandclassification. The dataset is large enough and reasonably balanced, which is why it can be used in the supervised machine learning experimentsandcomparativeanalysisofseveralclassifiers.

4. METHODOLOGY
Thefakenewsdetectorthatisproposedisbasedonasystematicandstructuredmachinelearningpipelinewithaspecifictext classificationproblem.Thismethodologyaimsatturningtheunstructuredtextualnewsinformationintoausefulformatand feedmachinelearningmodelswiththisinformationtoclassifynewsarticlesasfakeorreal.
The process of the work is divided into four major steps: text preprocessing, feature extraction, model training, and performanceevaluation.Thismodularityenhancesreproducibilityandenableseachofthestagestobeoptimizedandanalyzed separately.

Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072

Preprocessingoftextsisaveryimportantphaseinnaturallanguageprocessingbecauseuntidynewsarticlesusuallyhave noisesandinconsistencythatmayinfluencemodelperformancenegatively.Variouspreprocessingmethodscouldbeusedto cleanandnormalizethetextualdatainthisstudy,andthenfeatureextractionisperformed. Each and every piece of text is converted to lower case so as to maintain a consistent style and not to consider different versionsofthesamewordasdistinctcharacteristics.Punctuations,digits,andspecialcharactersareeliminated,sincetheydo notcarryanysignificantsemanticcontenttoclassifyfakenews. Stopwordsthatarecommonlike‘the’,‘is’,‘and’,‘are ’ are commonly removed to minimize the dimensions and concentrate on informative words. Lastly, additional whitespace is normalized,sothatdocumentsareconsistent.
Thesepreprocessingstepsaimtoeliminateredundantinformationandmaintainsignificantlinguisticpatternsthatdistinguish fakenewsandtherealnews.
The cleansed textual data are then pre-processed and converted to the numerical feature vectors with the help of Term Frequency-Inverse Document Frequency (TF-IDF) method. TF-IDF is a popular feature extraction technique in text classificationbecauseitgivesmoreweighttothewordswhicharecommoninadocumentbutarenotverycommoninthe overallcorpus.
Theschemeofweightedwordtallyingisespeciallyusefulwhenitcomestofakenews detecting,andtherearewordsand phraseswithahigherpropensitytobeusedindishonestorsensationalnews.Usingsuchdiscriminativewords,TF-IDFhelps machinelearningmodelstoimprovetheseparationoffakeandrealnewsarticles.
Togainadditionalinsightsintothetextualpropertiesofthedata,suchexploratoryanalysesasawordcloudareutilized.These visualizationsdemonstratethatitispossibletoidentifysomevisibledifferencesintheprevalenttermsoffakeandrealnews, and the hypothesis that linguistic patterns can be highly relevant when classifying fake news can be proven by these visualizations.
Themodeltrainingandevaluationinvolveareviewoftrainingresults,trainingprocess,andfuturetrainingrequirements.



International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Themodeltrainingandevaluationwillentailreviewoftrainingoutcomes,trainingprocess,andtrainingrequirementsinthe future.
MultipleclassicalmachinelearningclassifiersaretrainedusingtheTF-IDFfeaturevectors.Inordertohaveagoodassessment and maintain the initial distribution of classes, the sample is separated into training and testing sets through a stratified sampling strategy. All models are trained using the training data and tested using the hidden test data. The standard classificationmeasuressuchasaccuracy,precision,recall,F1-score,andReceiverOperatingCharacteristic-AreaUnderthe Curve(ROC-AUC)areusedtomeasureperformance.Thesemeasuresareabletoofferaholisticstudyofthemodelbehavior, especiallyasregardstheabilitytocorrectlydenotefakenewsandreducingfalseidentifications.
ItiswritteninPythonandstandardlibrariessuchasscikit-learnasamachinelearningsystem,pandasasadataprocessing library,NLTKasatextprocessinglibrary,andFlaskasawebframework.Codeissplitintoindividualfilesthatpreprocessing.py cleansthetext,trainmodel.pytrainsthemosteffectivemodel,evaluatemodel.pyproducesreportsandplotsonitsperformance and app.py is the web server. This arrangement allows me to debug one step at a time and debug other parts without destroyingalltheothers.
Weanalyzedthedatabeforetraining. Figure 1 alsoverifiedthattherewasagoodbalancebetweenfake(23k)andreal(21k) articles and thus it did not require any special balancing tricks. The lengths of the articles were diverse some with short headlines,otherswithlongreportswhichisidealwithTF-IDFbecauseitisabletoworkwithsizes.Wealsogeneratedword cloudswiththefakenewsutilizingwordswithmoredramaticuse,whereastherealnewsutilizestermsofpolicy. Therealhighlightiswebapplication.Theusercandirectlypastesometext,provideanewsURL(andthefeatureautoscrapes), hecanuploadanimage(OCRreadsanytext),andhecanevenavideofile(extractsaudioandconvertsspeechtotext).Allthe stepsarethesameasthoseoftrainingdata,loadsthesavedLinearSVCmodel,anddeliverspredictionswithconfidencescores. Processingis1-4secondsbasedonthetypeofinput,andisfastenoughtooperatenormally. Figure 5,6,7,8 representsthe workinginterfaceandsamplepredictions.



International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072


RESULTS AND DISCUSSION
Table 1 PerformanceComparisonofMachineLearningModels
Table 2 ClassificationReportoftheFinalModel

International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
LinearSVChadthehighestaccuracyof99.5on8,984testsamples.Table1isacomparisonofallmodelsthathavebeentested LinearSVCcomesoutofthegrounddecisivelyoverLogisticRegression(98.7%),LogisticRegression(98.8%),andothermodels intermsofaccuracy,precision,andF1-score.ThisvalidatesthefactthatTF-IDFiseffectivewhendealingwithnewstext. TheconfusionmatrixinFigure9narratestheentiretruth:outof4,696realfakearticles,itidentified4,491(only205were wronglylabeled).Among4,288realarticles,15weremistakenlylabeledasfake.Thetotalerrorwasonly220outofalmost 9,000excellentperformances.Theminorerrorsoccurredwhenactualnewsemployedsensationallanguagethatreadlike clickbait,whichisunderstandable.
Table2breaksitdownbyclass.Thedetectionoffakenewswas99.6%precise(infrequentlyfalselabelsgenuinestories)and 95.6%recall(almosteverythingiscaptured).Precisionofrealnewswasvirtuallyperfectat99.6.Figure10indicatesthatthe ROCcurvehasAUC=0.997,whichimpliesthatitisalsorobusttochangesinthedecisionthreshold. These numbers are superior to Ahmad et al. 96% but these are also much faster to train than other papers. The web deploymentdemonstratesthatitisnotjustlimitedtonotebooksthatanyonecanputittothetest.Lesssignificantconstraints includeuseoftextonlyanduseoftextpatternsonlywithoutimageriesorsocialcues.



International Research Journal of Engineering and Technology (IRJET) e-ISSN: 2395-0056
Volume: 13 Issue: 03 | Mar 2026 www.irjet.net p-ISSN: 2395-0072
Accordingtothisproject,afullfakenewsdetectorsystemwasdevelopedbasedondata,toalivewebapplication,withan accuracyof99.5%onreal-worldnewsdataonlybyusingaregularlaptop.TheFlaskinterfacesupportstext,weblinks,images andvideoswhichareratherunfamiliartomostacademicpapers.
Lessonslearned:evenbasicclassicalmachinelearningisnotbadintermsof performance(evenwithoutGPUsorhoursof optimization).Thisisusableinrealpracticeratherthanexperimentsduetothemodularcodeandwebdeployment.
Toimprovefuturework,itwouldbepossibletoaddtoolstoexplainthedatasuchasSHAPthatwoulddemonstratewhat specificwordscontributedtoeachprediction.Itcanbeenhancedbytestingotherlanguagesormixinglikes,sharesorother socialmediaindicators.EasyDockerpackwouldallownewssitestorunitwithoutproblems.Ingeneral,thisdemonstratesthat fakenewsdetectiondoesnotrequiredeeplearningthatisdifficulttoimplement.
[1]. H.AllcottandM.Gentzkow,"SocialMediaandFakeNewsinthe2016Election,"JournalofEconomicPerspectives,vol.31, no.2,pp.211-236,2017.
[2]. V.Perez-Rosas,B.Kleinberg,A.Lefevre,andR.Mihalcea,"AutomaticDetectionofFakeNews,"inProc.27thInt.Conf. ComputationalLinguistics(COLING),SantaFe,NM,USA,2018,pp.3391-3401.
[3]. Ahmad, M. Yousaf, S. Yousaf, and M. O. Ahmad, "Fake News Detection Using Machine Learning Ensemble Methods," Complexity,vol.2020,Art.no.8885861,2020.
[4]. M.Berrondo-OterminandA.Sarasa-Cabezuelo,"ApplicationofArtificialIntelligenceTechniquestoDetectFakeNews:A Review,"Electronics,vol.12,no.5,Art.no.1154,2023.
[5]. P.Akhtaretal.,"DetectingFakeNewsandDisinformationUsingArtificialIntelligenceandMachineLearning,"Annalsof OperationsResearch,vol.327,pp.633-657,2023.
[6]. M.Szczepanski,M.Pawlicki,R.Kozik,andM.Choras,"NewexplainabilitymethodforBERT-basedmodelinfakenews detection,"ScientificReports,vol.11,Art.no.18262,2021.
[7]. Kaggle,"FakeandRealNewsDataset,"2020.[Online].Available: https://www.kaggle.com/datasets/clmentbisaillon/fake-and-real-news-dataset