8
VII
http://doi.org/10.22214/ijraset.2020.7063
July 2020
International Journal for Research in Applied Science & Engineering Technology (IJRASET) ISSN: 2321-9653; IC Value: 45.98; SJ Impact Factor: 7.429 Volume 8 Issue VII July 2020- Available at www.ijraset.com
An Artificial Intelligence based Multi-Dimensional Approach for Identification of Cyber Threats in Web Application Aditya Raj Singh1, Surendra Kumar2 1, 2
School of Computing Science and Engineering, Galgotias University, Greater Noida, India
Abstract: The humongous growth in the number of web applications over the internet has enormously fascinated the crackers, security researchers, penetration testers, etc. as this exponential increase in the total web application count have provided them with the opportunity to infiltrate into more systems than ever before. This increase in the total web application count not only attracted these specific class of individuals but it has also caused the future forecasters to think of a way with the help of which vulnerabilities can be predetermined even before the web application is launched or if they do exist then they could be determined without any human intervention. This paper provides the solution in the form of Yes, as Artificial Intelligence is a field of computing that can be used to determine these vulnerabilities. This paper covers various vulnerabilities that exist in a web application and how artificial intelligence can be used for their identification. Keywords: Artificial Intelligence, Cyber-Security, Security, Penetration Testing, Ethical Hacking I. INTRODUCTION With the progression in the admiration of artificial intelligence it is being embraced in almost all the domains that it can conceivably cover, but still there are many more to be covered yet. Artificial intelligence has entered the domain of healthcare, agriculture and farming, research, etc. and it seems that due to its tendency of self-learning and improvising itself to provide a better solution it is going to reside for a long time[1]. It is forecasted that the upcoming future technologies will widely thrive over artificial intelligence and there is a strong reason for this, as humans dealing with the machines for testing its functionality and other behavioral aspects can sometimes miss out on the fine details that can lead towards drastic losses. These mistakes and weaknesses are termed as vulnerabilities in the cyber security domain[2]. Intruders look for these vulnerabilities that can be caused due to bugs using several tools that are generally scripts written in a scripting language that forces the subject to respond in a manner that it provides all the necessary information required by the intruder to infiltrate into that infrastructure, in a subject these vulnerabilities could exist due to several reasons like not updating the software or using a transmission medium that is not secure enough to hide the data that is being transmitted through it or using a software that has several exploits available for it, etc. and use it for gaining the access of that subject. Once the intruder gets into the system it makes sure that the access is maintained for future use and creates a backdoor in the infrastructure so that if the existing vulnerability gets patched he still has the backdoor through which he can get the access of the compromised subject. Despite implementing all the known possible security-based mitigations there are still tremendous vulnerabilities left in every web application and it takes a skilled intruder to bypass these security measures, as sometimes these vulnerabilities are known vulnerabilities and sometimes these vulnerabilities are new or zero-day vulnerabilities[3]. Cyber-security is another domain in which machine learning and artificial intelligence is being applied at a vast scale. Today nearly all the technical giants’ enterprises are already working on how to improve their security mechanisms using machine learning and artificial intelligence. These enterprises are building applications that are based on artificial intelligence and they tend to improvise themselves and the existing applications like mobile-based applications and web applications are being modified so that artificial intelligence can be used in them. One of the main reasons why artificial intelligence is being implemented in applications like web applications is that the intruders have become smart and are using artificial intelligence to make the malicious software more efficient.[2] The intruders are training the malicious software and scripts on training platforms and are increasing the complexity of the platform on which the attack is being performed in this way they are making the malicious software more and more dangerous and the malicious software is also becoming untraceable. So this becomes a major challenge and a necessity for the IT vendors to adopt artificial intelligence as it is their only hope for survival in the long race.
ŠIJRASET: All Rights are Reserved
388
International Journal for Research in Applied Science & Engineering Technology (IJRASET) ISSN: 2321-9653; IC Value: 45.98; SJ Impact Factor: 7.429 Volume 8 Issue VII July 2020- Available at www.ijraset.com II. ARTIFICIAL INTELLIGENCE Artificial Intelligence is the tendency of a machine to mimic the intelligence behavior of a living human being. Artificial intelligence provides a machine the ability to learn new features, experience different situations and learn using the experiences collected from these situations, adapt, etc. Artificial intelligence has a tendency of decision making that is based entirely on Machine learning as it is responsible for the learning part of the algorithm. Artificial intelligence is based on the following processes: 1) Learning Process: In this process the main aim is to convert the gathered big data into a form so that it can be used for processing by the algorithms. Rules governing the algorithm are also defined briefly in which it is stated that how the steps will be performed for the completion of the given task. 2) Reasoning Process: This is the second and the most important process as in this process the optimal algorithm selection takes place so that the task can be fulfilled without consuming more resources and time. 3) Self-Improvisation Process: This is the third process that aims at refining the algorithm in a manner that produces a more precise result and this process is more of an iterative process. A. Machine Learning Artificial intelligence is composed of several branches, one such branch is machine learning. The main idea of machine learning is that when a large amount of data is given to an algorithm then it can learn several patterns and can perform predictions that are based on the previously given input data to the machine learning algorithm. This data interpretation and pattern recognition ability of machine learning provides the ability of decision making to the artificially intelligent systems without any human intervention[4]. B. Deep Learning Deep Learning is a technique of machine learning that teaches the algorithm about the accuracy of recognizing the patterns in the provided input. The input for deep learning-based algorithms is a huge amount of labelled data. Deep learning-based algorithms require significantly high computational power than machine learning. Deep learning-based algorithms use artificial neural networks hence they are also termed as deep neural networks. One of the most widely used deep learning-based neural networks is a Convolutional Neural Network. Generally neural networks require feature extraction that is performed manually but with convolutional neural networks the need for manual feature extraction is eliminated[5]. Deep learning-based models are comprised of thousands of hidden layers that are responsible for making these algorithms more precise than machine learning algorithms. Deep Learning algorithms can be used for the identification of patterns that reflect vulnerabilities in web applications. It can also be used for the identification of zero-day exploits and can term them as a new vulnerability for the web application[2]. C. Artificial Neural Network Machine learning comprises of tremendous algorithms. One such type of algorithm is ANN or Artificial Neural Network. Artificial Neural Network aims to Copy the learning pattern of Biological Neural Network. It is generally comprised of three components i.e. input layer, hidden layers, and output layer. Artificial Neural Network is composed of several small entities that are commonly known as nodes these nodes mimic the working of Biological neurons and therefore can be termed as artificial neurons[6]. The learning mechanism in Artificial Neural Network is based on the huge number of similar types of examples that are provided to it for training and identification of a particular trait without any prior knowledge of the entity of the example. Artificial Neural Networks tend to derive the characteristics that can be used for the identification of a particular trait. This feature of Artificial intelligence can be used in a very productive manner for the identification of vulnerabilities in the existing systems, as it can produce characteristics that can be used for the derivation of the characteristics that are responsible for creating vulnerabilities in a web application. D. Predictive Analytics When Machine learning is combined with historical input data along with several mathematically composed statistical algorithms then it is termed as predictive analytics. Technologies that are built using Artificial Intelligence or the old ones that are modified for the implementation of Artificial Intelligence are using Predictive for forecasting the future behavior of an application. In similar fashion predictive analytics can prove to be of great help in identification and best assessment of web-based vulnerabilities as a large amount of data related to the web applications is archived and available for educational purposes.
ŠIJRASET: All Rights are Reserved
389
International Journal for Research in Applied Science & Engineering Technology (IJRASET) ISSN: 2321-9653; IC Value: 45.98; SJ Impact Factor: 7.429 Volume 8 Issue VII July 2020- Available at www.ijraset.com III. CYBER THREATS IN WEB APPLICATIONS Any kind of loophole in the security that can exist due to the mistake made by the manufacturer or the developer with the help of which an intruder can get access to the system and exploit it for the fulfilment of his object of desire is termed as a cyber-threat. Cyber threats associated with web applications can be of several types. According to OWASP(Open Web Application Security Project) foundation, a list comprising of ten threats against the web applications are issued on an annual basis that can cause devastating effects and they are considered to be the industry standard as OWASP foundation is dedicated towards the research of various threats associated with the web-based applications and also develops various intentionally vulnerable simulation labs with the help of which exploitation of these vulnerabilities can be performed. Currently, the OWASP top 10 are as follows: 1) Injection 2) Broken Authentication and Session Management 3) XSS (Cross-Site Scripting) 4) IDOR (Insecure Direct Object References) 5) Security Misconfiguration 6) Sensitive Data Exposure 7) Missing Function Level Access Control 8) CSRF (Cross-Site Request Forgery) 9) Using Components with Known Vulnerabilities 10) Un-validated Redirects and Forwards A. Injection According to OWASP an injection is a type of vulnerability in which the web application allows the client to execute a code written in a particular language through the parameters used by the web application. Since the client can execute his code in the web application through the parameters, therefore, it is termed as an injection. The injected code can be written in SQL, PHP, HTML, etc. and is generally named after the scripting language used for the injection. An example of SQL injection can be understood using the given SQL based code: +-Select * from USER_ACCOUNTS_TABLE where CUSTOMER_ID = ‘” + request.getParameter(“CUSTOMER_ID”) + “’”; It appears that the web application which is executing this block of code accepts a parameter i.e. the id of the customer and places it in the CUSTOMER_ID field and sends it back to the server for execution. Injection Code: True‘or’1’=’True If the attacker which is on the client-side passes this injection code as an input to the parameter then it will provide him with all the entries present in the USER_ACCOUNTS_TABLE. B. Broken Authentication and Session Management Before understanding what is broken authentication and session management a question arises that what is a session. As the user logs into the system most of the web application desire of saving the user’s data so that they can generate a state known as session. The main functionality of a session is to let the web application knowledge of the authenticity of the logged-in user so that it could be taken further for other services that are offered by the web application. A session sometimes stores information and is not disposed of properly leading to Broken Authentication and Session Management vulnerability in the web application. Hence proper session management mechanism states that every time the user goes through the authentication phase a unique session should be created and if there is no activity performed on the web application or the browser for a certain period which can vary for different web applications the session should be killed immediately and the user whose session is expired should be sent back to the authentication page from where he should again log in. An example of Broken Authentication and Session Management can be understood through the following link: http://wxyzcart.com/sale/3K9IH8UGDEQZLOPGAPRHOLA/?item=juicer In the above-mentioned URL the web application is storing the credit card details in the session id which can be used by any intruder for making the payment with the same card again and again.
©IJRASET: All Rights are Reserved
390
International Journal for Research in Applied Science & Engineering Technology (IJRASET) ISSN: 2321-9653; IC Value: 45.98; SJ Impact Factor: 7.429 Volume 8 Issue VII July 2020- Available at www.ijraset.com C. XSS (Cross-Site Scripting) Web applications that are built today uses several technologies so that they can provide the user with a very easy to handle and an appealing interface. For providing these features the web application uses technologies like JavaScript, PHP, Flash, etc. Since every coin has two sides in the same manner these technologies provide a ground to the intruder through which it can inject his code after knowing the technology being used by the web application. XSS can be performed after getting the administrative privileges that can be obtained using admin bypass SQL injection and then the attacker can inject the malicious payload into the web application and the victim web application can be used as an infecting vector or could be used as a bot. XSS infected web applications when running on the browser along with other applications on the browser are capable of stealing informative data used by the subject very easily. Example of the injection payload that can be used for performing the cross site scripting attack is as follows:"/>jaVasCript:/-/`/\`/'/"//(/ */oNcliCk=prompt() )//%0D%0A%0d%0a//</stYle/</titLe/</teXtarEa/</scRipt/-!>\x3csVg/<sVg/oNloAd=prompt(123)//>\x3e D. IDOR (Insecure Direct Object References) Insecure Direct Object References is a type of vulnerability that exists in the web applications due to the inefficiency of the access control mechanism used by that web application. In an Insecure Direct Object References vulnerability the intruder can perform tampering of the URL in a manner that it could provide him the access of the unauthorized features of that web application or it could also result in the exposure of the sensitive data of other users. An example for Insecure Direct Object References vulnerability could be understood using the below-stated URL: http://abc.xyz/user=Aditya/home If this web application is not using proper access control mechanism and there exists an Insecure Direct Object References vulnerability then the attacker can try to tamper this URL and can change the name of the user from Aditya to root or administrator which could result in providing him the access of root user account. E. Security Misconfiguration Security Misconfiguration is among the most disastrous vulnerabilities that are encountered to date, the reason being that this vulnerability is capable of compromising the whole system. Some examples can describe Security Misconfiguration vulnerability in the easiest of ways like if a default account settings like credentials of the default account of a server are not modified then an intruder can get access to the server using the default account and can cause sincere damage to the web application. Another example could be a non-disposal of login pages that are associated with a root user account or administrator account. Other examples can be improper management of directories and directory listings along with the read-write-execute privileges for the users. F. Sensitive Data Exposure Any type of data that contains information related to an individual or an organization that should not be available for the public as unnecessary advantages can be taken with the help of that data is considered to be sensitive data. Sensitive data exposure can be understood using some examples. Like if the data that is being transmitted over a media and the media channel is not able to provide confidentiality to the transmitting data then it can lead to Sensitive data exposure if the channel is under a man in the middle attack. Another example of Sensitive data exposure can be storing the credentials of the users in the database without using any strong encryption mechanism. If the database gets compromised then all the credentials will get exposed to the attacker and he can take undue advantage of those users accounts. G. Missing Function Level Access Control Functional Level Access is performed on the web application and the server during the staging phase and if the functional level access control is not done efficiently then the intruder can get access to the web application or the server. In this vulnerability tampering of parameters or the functions available in the URL is done and then the intruder gets access as if he is a privileged user. Generally access control in the web application is performed using Settings that are available in the configuration section of the web application and no code or script is used by the administrator to manage this issue and hence sometimes attackers take advantage of these misconfigurations.
ŠIJRASET: All Rights are Reserved
391
International Journal for Research in Applied Science & Engineering Technology (IJRASET) ISSN: 2321-9653; IC Value: 45.98; SJ Impact Factor: 7.429 Volume 8 Issue VII July 2020- Available at www.ijraset.com H. CSRF (Cross-Site Request Forgery) When it comes to some of the most common and most effective vulnerabilities Cross-Site Request Forgery always makes up to the list as it is one of the favourites of the attackers. In CSRF or XSRF attack an intruder uses a file as a bait. The file is not a normal file but is a tampered one which is composed of a hidden link in it. The attacker sends this file and wait for it to be clicked. As soon as this file is clicked a request is generated from the victim's system and then the user can take advantage of the victim's system. But there is a constraint for Cross-Site Request Forgery to work efficiently and it is that the user should be an authorized user. I. Using Components with Known Vulnerabilities Whenever a project is given to a developer it always has a deadline within which the developer has to submit the finished project with all the desired features. Sometimes it happens that the development team or the developer uses pre-existing code or functions or libraries or GitHub based codes for the completion of the project in the given period. But due to the use of these codes the development team forgets that the code or the libraries or the functions that they have used might be vulnerable to attacks and that they can cause vulnerabilities in the system in which it is used. J. Un-validated Redirects and Forwards Redirection is a property of a web application in which the web application takes the user to the pages of the web application that are hosted in some external domain. The process of redirection is a client-side process which means that it takes place in the user’s browser and not on the server-side. Whereas forwarding process is a server-side process which means that it is not handled by the user's browser but by the server on which the web application is running. In the process of forwarding the browser of the user sends a request to the server on which the web application is hosted and request for a resource, the server then validates this request and if it appears to be genuine then the browser takes the user to the page he wanted to visit. If proper management of redirection and Forwarding is not done along with proper validation methodology for these processes then the request could be tampered by the attacker and it can change the redirect from authentic web application page to some infected web application that is capable of compromising the system of the user. IV. PROPOSED MODEL The proposed model states that Artificial intelligence based concepts like Machine learning, deep learning, Convolutional Neural Network, and predictive analytics along with several other security-based concepts can be combined to develop a system that is capable of identification of existing cyber threats along with zero-day cyber threats for web applications. The proposed system is supposed to work in stages. During the training stage which is the first stage of this system information collected from several vulnerable lab environments that are developed so that penetration testing can be performed on them like BWAPP, Google Gruyere, HackThis!!, Mutillidae, Peruggia, Root Me, WebGoat, etc. will be used along with automation scripts which will comprise of the information gathering methodologies, scanning methodologies and exploitation techniques for identification of cyber threats in web application and how to perform them effectively without any human intervention. Convolutional Neural Networks will be used for eliminating the manual feature extraction. The testing stage which is the next step in the first stage will comprise of some of the vulnerable lab environments with known vulnerabilities along with some of the vulnerable lab environments with unknown vulnerabilities will be tested and predictive analytics will be used for prediction of vulnerabilities that the system has learned during the training phase. After the system performance is measured by the means of accuracy then the categorization of vulnerabilities will be done based on their impact and the threat landscape. This phase which is the next phase is known as the validation phase The next stage which is the tuning stage will focus on increasing the accuracy of this system by increasing the number of vulnerable lab environments along with newly discovered exploits from exploits databases and web vulnerability search engines like Shodan, etc.
ŠIJRASET: All Rights are Reserved
392
International Journal for Research in Applied Science & Engineering Technology (IJRASET) ISSN: 2321-9653; IC Value: 45.98; SJ Impact Factor: 7.429 Volume 8 Issue VII July 2020- Available at www.ijraset.com
Data Collection (Intentionally Vulnerable lab environments and archived historical web applications)
Data Understanding Feature Engineering Modelling Validation Tuning Deployment Monitoring This model is then deployed and further monitoring takes place for future references. V. FUTURE SCOPE The idea proposed in this paper has a very bright future as the technology is evolving every second at a very fast pace and so is artificial intelligence. The concept of using artificial intelligence in the cyber security domain is the future of both the technologies as with the growth in internet-based systems and other applications that are based on artificial intelligence cyber security is becoming a must for these types of technologies. Since these technologies are based entirely on internet and networks therefore there exist a point of risk that the security of these type of systems should be considered to be the topmost priority or else if this is not done then it could lead towards devastating circumstances. Modifications can be made in this idea and a few more Artificial Intelligence based algorithms can be added in this research so that this concept can be used for building a system that can implement this concept. VI. CONCLUSION This paper proposes an Artificial Intelligence based model with the help of which vulnerabilities in the web application can be determined using Artificial intelligence and Machine Learning Techniques. Currently vulnerability assessment and penetration testing are being performed either manually using commands, scripts, etc. or with the help of vulnerability scanners. These vulnerability scanners and manual vulnerability assessment sometimes prove to be inefficient in determining some of the known web-based attacks and zero-day attacks for web applications. So this paper present forward an approach with the help of which this issue of vulnerability detection can be resolved very easily with the help of Artificial Intelligence. REFERENCES [1] [2] [3] [4] [5] [6]
R. Talwar and A. Koury, “Artificial intelligence – the next frontier in IT security?,” Netw. Secur., vol. 2017, no. 4, pp. 14–17, 2017, doi: 10.1016/S13534858(17)30039-9. D. Arivudainambi, V. K. Varun, S. C. S., and P. Visu, “Malware traffic classification using principal component analysis and artificial neural network for extreme surveillance,” Comput. Commun., vol. 147, no. August, pp. 50–57, 2019, doi: 10.1016/j.comcom.2019.08.003. V. Kanimozhi and T. P. Jacob, “Artificial Intelligence based Network Intrusion Detection with hyper-parameter optimization tuning on the realistic cyber dataset CSE-CIC-IDS2018 using cloud computing,” ICT Express, vol. 5, no. 3, pp. 211–214, 2019, doi: 10.1016/j.icte.2019.03.003. U. Noor, Z. Anwar, A. W. Malik, S. Khan, and S. Saleem, “A machine learning framework for investigating data breaches based on semantic analysis of adversary’s attack patterns in threat intelligence repositories,” Futur. Gener. Comput. Syst., vol. 95, pp. 467–487, 2019, doi: 10.1016/j.future.2019.01.022. N. Balakrishnan, A. Rajendran, D. Pelusi, and V. Ponnusamy, “Deep Belief Network enhanced intrusion detection system to prevent security breach in the Internet of Things,” Internet of Things, pp. 100–112, 2019, doi: 10.1016/j.iot.2019.100112. T. Semwal and S. B. Nair, “A decentralized Artificial Immune System for solution selection in Cyber–Physical Systems,” Appl. Soft Comput. J., vol. 86, p. 105920, 2020, doi: 10.1016/j.asoc.2019.105920.
©IJRASET: All Rights are Reserved
393