SMS SPAM CLASSIFIER FOR ROMAN URDU AND ENGLISH SMS Rana Muhammad Ammar, Hamna Shahid, Mudassar Niaz Rana Muhammad Ammar Khan is currently pursuing Bacholer degree program in Software engineering in COMSAT University, Pakistan, PH-+92-308-8741046. Email: ranaammar046@gmail.com Hamna Shahid is currently pursuing Bacholer degree program in Software engineering in COMSAT University, Pakistan, PH-+92-304-0617983. E-mail: hamnashahidkhanewal@gmail.com Mudassar Niaz is currently pursuing Bacholer degree program in Software engineering in COMSAT University, Pakistan, PH-+92-304-6788995. E-mail: niaz.mudasar1122@gmail.com
KeyWords SMS spam, spam classifier, Roman Urdu, English, machine learning, language variations, informal language, slang, limited text length, imbalanced datasets, preprocessing, feature selection, evaluation metrics, n-gram analysis, text normalization, phishing attacks, mobile network security.
ABSTRACT The proliferation of SMS (Short Message Service) spam poses a significant challenge in maintaining a positive user experience and ensuring mobile network security. This research article focuses on developing an SMS spam classifier specifically designed for Roman Urdu and English SMS messages. The classifier employs machine learning techniques to accurately differentiate between spam and legitimate messages. The research explores the challenges associated with language variations, informal language and slang, limited text length, and imbalanced datasets. Preprocessing techniques for both Roman Urdu and English SMS are discussed, along with feature selection strategies. Evaluation metrics, including accuracy, precision, recall, and F1 score, are utilized to assess the classifier's performance. Techniques such as n-gram analysis, text normalization, and handling imbalanced datasets are examined to enhance the accuracy of the SMS spam classifier. The real-world applications of SMS spam classification, including protecting users from phishing attacks and enhancing mobile network security, are highlighted. The findings of this research contribute to a safer and more secure mobile communication environment by effectively classifying SMS spam in both Roman Urdu and English languages.