Regulatory & Compliance
Generating Powerful New Pharma Insights from Graphs
Neo4j’s Alex Jarasch explains how knowledge graphs can be a valuable asset for the pharmaceutical industry to extract insights from complex data.
An interesting possibility is emerging for the life sciences industry – can we use a new approach to organising complex data to gain a more holistic view of the underlying problems we want to address? The new approach is knowledge graph, which is defined by the AI research group The Alan Turing Institute as capturing information about entities of interest in a given domain or task (like people, places or events) and finding the connections between them. Tackling the Hard Bio Problems A knowledge graph is a type of data structure that represents information as a set of entities and the relationships between them. It is a way to organise and link data together in a way that is more easily understandable and accessible. The main feature of knowledge graphs is that they allow for complex connections between different data sources, which can reveal new insights and relationships that would be difficult or impossible to uncover using traditional SQL databases. Knowledge graphs are multidimensional and can represent data in a variety of forms, including text, images, and structured data. This allows them to store and link together different types of information, such as scientific research, clinical trial data, and market data. The ability to make connections between different data sources is what makes knowledge graphs so powerful for the life sciences industry. The first real-world use of graph technology that entered the public's consciousness was the Panama Papers scandal, where a network of journalists used graphs to link together millions of documents and uncover illicit financial activities. That seems at a remove from life sciences in one way, but in another it isn’t – both areas involve seeing useful patterns in a lot of what seems like “noise” to the average person. In the field of biological science, for example, understanding the complex interrelationships between different factors, such as genes, environment, diet, and behaviour, is essential for gaining a deeper understanding of diseases and developing new treatments but is notoriously hard. Knowledge graphs could play a crucial role in aiding us here, by empowering analysis of these interrelationships and correlations on a large scale. Modern native graph databases are particularly well-suited for this type of analysis. They are optimised for handling large amounts of interconnected data and so can store and link together billions of connections. That makes it possible to 12 INTERNATIONAL BIOPHARMACEUTICAL INDUSTRY
analyse and make sense of massive amounts of data – important in fields such as medicine, where the amount of data being generated is growing rapidly, and traditional methods of data analysis are no longer sufficient. The potential for using knowledge graphs in the life sciences industry is significant, but it’s still early days. While many pharmaceutical companies are already using knowledge graphs, the majority of these companies are only using them for a small portion of their database work. This means that there is a lot of untapped potential for using knowledge graphs to gain new insights and improve decision making in the industry. The use of knowledge graphs in the life sciences industry has been primarily focused on three problems: identifying novel drug targets for new therapies, transforming clinical trials, and managing supply chains in a more dynamic fashion. As companies begin to use knowledge graphs to connect and analyse their data, they are hoping to find new ways to supplement these use cases and approach high throughput screening and compound registrations in more efficient ways. Neo4j recently hosted an event to explore the possibilities of knowledge graphs in the life sciences industry. Participants shared a wide range of emerging use cases. An example was AstraZeneca's application of knowledge graph technology with chemical reactions to predict new reactions and molecular synthesis. This was to aid that AstraZeneca to discover new compounds and to improve the efficiency of the drug discovery process. Additionally, knowledge graphs are predicted to help AstraZeneca circumvent or attack existing patents by finding new ways to synthesize the same compound. Graph Queries Execute 3,000 Times Faster than SQL Another example that’s emerging as best practice is at GlaxoSmithKline (GSK), which is exploring the use of knowledge graphs to improve clinical reporting workflows. The traditional clinical reporting process can be labour-intensive, involving multiple handoffs and data transformations. GSK researchers found that by applying knowledge graph technology, they were able to overcome these challenges by creating a single, contextualised knowledge graph that can be used to streamline the process and make it more efficient. It was found to help reduce the time and resources required to complete clinical reports, and boost the overall accuracy and completeness of the data. In addition, by using a Graph Data Science library and out-of-the-box machine learning tools, GSK was able to process large amounts of data and extract valuable insights that can be used to improve the drug development process. The use of knowledge graphs helps make sense of complex data and uncovers hidden patterns and relationships that would be difficult or impossible to find using traditional methods. One of the key advantages of knowledge graphs is that they are not dependent on data sources being prepared or formatted Spring 2023 Volume 6 Issue 1