Vir Mehta Applying Biomedical Knowledge Graphs for Precision Medicine
Senior Thesis | 2026

Introduction
Inrecentyears,thedevelopmentoflarge-scale,data-intensive biomedicalresearchhasrequiredincreasinglysystematicmethods tointegrateandmakesenseofheterogeneousinformation(Zhang etal.).Becauseofthis,knowledgegraphshaveemergedasa popularandwell-establishedmeanstoorganizedata.Inthese graphs,biologicalandclinicalentities,likegenes,proteins, diseases,ordrugs,arerepresentedasnodes,whileexperimentally validatedorcomputationallyinferredassociationsandrelations betweenthemareencodedasedges.Incontrasttoother representations,suchastablesorpathway-centricviews,theycan encodearangeofrelationshiptypesanddifferentontologieswhich canbehelpfultorepresentcomplexdatasets.Althoughknowledge graphscantaketheformofhomogeneousgraphs,inwhichall nodesandedgesareofthesametype(forexamplealldrugnodes, andedgessimplyshowingwhentwodrugsarerelated),theyare morecommonlymulti-relationalwithheterogeneousgraph topologies.Thisisbecausemulti-relationalstructuresareableto depictmuchmorecomplexdatasets.Furthermore,theyareuseful tostudybiologicalsystemsinwhichinteractionsaretypically network-structuredanddistributed,suchasprotein-protein interactionnetworks.

Figure1:Heterogeneousbiomedicalknowledgegraphwith multipleedgetypesandnodetypescenteredaroundanacidnode (TomazBratanic).
Figure1showsanexampleofasimplebiomedicalknowledge graphwithonly3differentnodetypes,asdepictedabovebythe greencolors,bluecolorsandonepinkcolornodetype.Figure1 alsoshows2differentedgetypes,the“inhibit”relation,asseen connectingBCL2geneandascorbicacid,andthe“treats”relation, asseenconnectingascorbicacidwithhemorrhagiccystitis. Knowledgegraphsthatareusedforrealresearchandanalysis oftencontainthousandsofnodesandmillionsofedges,meaning thatvisualizationssimilartoFigure1becomeextremelydifficult tocreate.Instead,visualizationsofthesetypesoflargegraphsare
oftencreatedusingclusteringofnodestosimplifythelarger,more complexgraph.
Knowledge-graphframeworksarealsoextremelyapplicableand well-suitedtoavarietyofdownstreammachinelearningtasks, includingnodeclassification,edgeprediction,linkprediction,and clustering,eachofwhicharebasedoncapturingstructural regularitiesofmultiscalebiologicalnetworks.Thesetasksalso havealotofapplicationsacrossbiomedicalresearch.Forexample, edge-predictiontasksonadrug-targetgraphmightidentify candidatedrugmechanismsofaction,whiledisease-associated nodeclassificationcouldbeusedtoidentifygenesinvolvedin poorlycharacterizedpathologies.Furthermore,conditionswith complexetiologies,forwhichnosingledatamodalitycanprovide completeinformation,haveespeciallybenefitedfromthisdata form.Thisisbecauseknowledgegraphs,unlikenormaldatatables, arebettersuitedtoperformjointreasoningacrossmolecular, clinical,andpharmacologicalinteractionsbecauseoftheirnode andedgestructure.Theprimarymotivationofthisthesisistouse knowledgegraph-basedinferencebyapplyingalgorithmsonthe knowledgegraphtotryandidentifygeneticfactorsassociatedwith diseasessuchasParkinson'sDiseaseandBreastCancerand identifydrugswhichcouldbeusedeffectivelyin drug-repurposing.
Bycombiningdataondrug-targetinteractions,molecular pathways,adverseeventprofiles,andknowndiseaseassociations, graph-basedapproachesmakethispossibleandmayevenuncover therapeuticrelationshipsthatwouldotherwisegooverlookedin dataformslikeatable.Inaddition,becauseParkinson'sDiseaseis multifactorial,applyingalgorithmscenteredaroundParkinson’sto theknowledgegraphmakesaninterestingtestcasefor graph-basedbiologicalinference.Althoughsomemonogenic formsofthediseasehavebeenidentified,mostcasescannotbe
ascribedtosingle-genemutations,andaheterogeneousgraph providesanorganizedwaytointegrategenomic,transcriptomic, proteomic,andclinicaldatainordertoidentifythegenesthatmay confersusceptibilityoraffectthecourseofthedisease.
Onedownsideoftheknowledgegraphisthatitishardtoencode intoanumberformatformachinelearningalgorithmstorunon. Oneoptionistocreateanadjacencymatrixcontainingbinary representationsforconnectionsforthegraph(anentrywillbe1 wherenodesareconnected,otherwiseitwillbe0),withthe dimensionnbynwherenrepresentsthenumberofnodesinthe graph.However,becausebiomedicalknowledgegraphsareso large,thisformatbecomesextremelytedioustocreateandwork with,andtakeslongerthanadaytorunasimplemachinelearning algorithm. Instead,embeddingtechniquesareemployedtomapgraphnodes intoacontinuousvectorspaceforpredictivemodelinginthese intricatenetworks.Differentneuralnetworkslearnthetopologyof eachnodebylookingattheneighbouringnodesandedges,and thenforeachnode,outputacorrespondingvectorwithrealnumber values(BusraOguzoglu).Usingvectorsallowsmachinelearning modelstofunctiononfixed-dimensionalnumericalrepresentations whilemaintainingcrucialstructuralandsemanticrelationships. Theseembeddingsserveasabridgeandaccuraterepresentationof contemporarycomputationalmodelsandsymbolicbiological knowledge.Bycombiningdatafromanode'slocaltopologyand identifyingnonlinearinteractionssuggestiveofunderlyingbiology, neuralnetworks(andgraphneuralnetworksinparticular)represent apowerfulmechanismforlearningfromtheseembeddings(Z.Ye, etal.).Theyareidealfortaskslikepredictingputativetherapeutic linksbetweendrugsanddiseasephenotypesorinferringnovel disease-geneassociationsbecauseoftheirabilitytogeneralize acrossavarietyofgraphstructures.
Theseelementstogetherprovideacoherentcomputationalstrategy forintegratingbiomedicaldata,generatingbiologically informativeembeddings,andapplyingpredictivemodelingto questionsofdrugrediscoveryandParkinson'sDiseasebiology. Thisframeworkillustratesthebroaderpotentialof knowledge-graphguidedinferenceandscientificdiscoveryina biomedicalsetting.
Methods
Inordertoconstructthebiomedicalknowledgegraph,astructured data-engineeringworkflowdesignedtointegrateheterogeneous biologicalandclinicalresourcesintoaconsistentrepresentation wasfollowed.Rawdatasetswerefirstparsedfromoriginaloutputs suchasJSON(JavaScriptObjectNotation),XML(eXtensible MarkupLanguage)andTSV(Tab-SeperatedValues).Dueto datasetsoftenbeingextremelydifferentformats,multiplestepsand scriptswererequiredtoensurethattheultimategraph representationwouldreflecteachdatasetsimilarly.Therewerealso somedatasetssuchasMONDOwhichwerealreadyformattedina graphicalrepresentation,unlikeDrugBankorCTDdata (Drugbank).
Eachdatasetrequiredaregularizationstage(thoughmultiple scripts)inwhichdifferentidentifiersandstringsweremappedto standardontologies,drug/diseasenameswereharmonized (mappedtogetherandsortedbycase),andanyredundantor conflictingrecordswereremoved. Ifthereweremultipledatasets withconflictingdataregardingaparticularedgeorpairofnodes, thenthatspecificrelationwassimplyomitted.Thisstepensured thatentitiessuchasdiseases,drugs,andgenescouldbealigned unambiguouslyacrosssources.Oncenormalized,alldatasetswere
integratedbymappingtheirentitiesandrelationshipsintoasingle unifiedschemarepresentativeofthegraph'snodeandedgetypes.
Fromthere,modularscriptswerewritteninPythonthatwere developedtoincorporatenewdatasetswithaminimumof additionalwork,allowingforlong-termextensibility,providinga viablefutureuseforthegraph.Thescriptsallowforeasier applicationinthefutureandadaptingthegraphwelltoaddingnew datasetsorupdatesdatasets.Alargenegativeofknowledgegraph usetodayisthatlargegraphsareoftenextremelydifficulttokeep updated,socreatingthisfunctionheavilyincreasestheuseand applicabilityofthisgraph.Eachscriptperformsschemavalidation, identifierreconciliation,andformattingchecksbeforeaddingnew nodesoredges.Thefinalintegratedgraphwasthenexportedin both.csvand.parquetformats.Theseformatswerechosenbecause theyensurethewidestcompatibilitywithanalyticalworkflowsbut alsoallowthedatasettobeefficientlyhostedonHarvard Dataverseforpublicaccessandreproducibility.

Figure2:Overviewofthestructuresandvisualizationoflarge clusterswithinPrimeKG(knowledgegraph).
Figure2showsdifferentvisualizationsoftheoverallknowledge graph.Theultimategraphconsistsofapproximately160,000nodes andover21millionedges.Thus,itisextremelyhardtovisualize theentiregraph,resultinginvisualizationswhichrelyonclustering ofnodetypestoseealargerscalepicture.Panelbshowsall diseasenodesinPrimeKGvisualizedinacircularlayouttogether withdisease-associatedinformation.Itshowstherelationships betweendiseasenodesandanyothernodetype,demonstrating howdiseasenodesaredenselyconnectedtofourothernodetypes intheknowledgegraphthroughseventypesofrelations.This highlightshowthegraphisabletoencapsulatelargeamountsof dataandrepresentcomplexrelationshipsrelatedtodiseases.
Oncethegraphwasbuilt,algorithmstoanalyzethegraphandfind nodesimilaritywereimplemented.Theseconsistedofalgorithms
basedonrandomwalks.Akeyexampleofthisismetapathbased randomwalks,whichutilizespecificnodetypestomakedecisions whentraversingarandomwalk.Thistypeofalgorithmwould simplytakealistargument,describinginorderthenodetypesthat arebeingsoughtafter.Forexample,whenanalyzingdisease-gene relationshipswithinthegraph,the[‘disease’,‘drug,’ ‘gene/protein’]metapathisused,constrainingtherandomwalkto onlylookforpathswhichstartedonadiseasenode,traversedtoa drugnodeandthenfinishedatagene/proteinnode(Noorietal.). Ultimately,thisalgorithmisutilizedinmultipleways,most commonlytotryandfindnodesimilaritybycountingthetotal numberofmetapathsbetweentwonodes,andalsotolookfor potentiallysignificantproteinsandgenesthatwerehighlylinkedto adiseasenodesuchasParkinson'sDisease,butnotheavily researchedinthepastdecade.
OtheralgorithmsusedtoanalyzethegraphconsistedofthePage Rankalgorithmwhichisarandomwalk–basedmethod.Itstartsat anodeXandcalculatestheprobabilitymassthatreachesanodeY througharandomwalk.Inthisapplication,nodeXisspecifiedas theParkinson'sDiseasenodewiththehighestdegree(i.e.,themost connectededges).NodeYisspecifiedasaparticulardrugnode, andscoresarecalculatedbyiteratingthroughalistofthemost commonParkinson’sdrugswithintheknowledgegraph.The bi-directionalscoreisthencomputedbyswappingnodesXandY, recalculatingthePageRankscore,andaveragingthetwovalues. Thisapproachhelpsminimizelossofaccuracywhileensuringa largersamplesizeofwalks.Finally,allvaluesarenormalized usingz-scoreandmin–maxnormalizationtoensurethatthefinal heatmapvisualizationiseffective.Conceptually,thisalgorithm answersthequestion:“Howeasyisittoreachthisdrugfromthis disease,andviceversa,throughbiologicallymeaningfulpathsin theknowledgegraph?”
Unlikethemetapath-basedalgorithm,PageRankconsidersall possiblepathsregardlessoftheintermediatenodetypes.Asa result,itplacesgreaterweightonpathlengthandnodedegree: shorterpathsareweightedmoreheavilyandinterpretedasstronger connections,whilepathsthatpassthroughhigh-degreenodes dilutetheprobabilitymassandyieldlowerconnectionscores.This canbeunderstoodaswhenafixedresourceisdistributedequally acrossmoreconnections,eachconnectionreceivesasmallershare, sohigher-degreenodesexertweakerinfluenceonanyindividual neighbor.Overall,thePageRankalgorithmenablesaclearer understandingoftheglobaltopologyandsharedneighborhood structureoftheknowledgegraph(Gallo,Denis,etal.).
Therealsoexistsakeyparameter alpha inthePageRank algorithm,whichencodestheprobabilitythatthenextstepinthe randomwalksimplyjumpstoarandomnode.Thispreventsthe walkfromgoingonirrelevantdetoursordeadendsthroughoutthe graph,helpsthewalkdiscovermorepartsofthegraph,andensures thatthewalkwillneverget“lost.”Inthiscasebecausethe algorithmispersonalized(PersonalizedPageRank),insteadof jumpingtoanyrandomnode,thealgorithmwilljumptothe originalstartingnode,inthiscasetheParkison’sdiseasenode.
Graph-basedneuralmodelswerealsousedtogeneratevector representationsofnodesfordownstreammachinelearningtasksin ordertobetterunderstandandanalyzethegraph.Thisisextremely importantbecausecomputeralgorithmsallworkbasedon numbers,andvectorrepresentationscaneasilyallowforsuch algorithmstoworkefficientlyandmeasuremetricssuchas similarity. ThecodewasimplementedthroughlibrariessuchasPyTorchand PyTorchGeometric.GraphSAGEandgraphattentionnetworksare preferredduetotheirsupportforheterogeneousneighborhood
aggregationandeffectivescalingtolargegraphs.Typical hyperparametersconsistedoflearningratesbetween1e-3and1e-4, hiddendimensionsbetween64and256,batchsizesofseveral hundrednodes,anddropout(settoaround0.2).Multipletraining phaseswererequiredduetothepresenceofskewednodedegrees, noisybiologicalassociations,andtheriskofover-smoothingupon stackingmultiplemessage-passinglayers.Earlystopping, neighborhoodsampling,andregularizationwerealsoappliedas mechanismsforensuringstableconvergence.Thesemethodswere enforcedmoreheretotryandreducetheoverallcompute,asa knowledgegraphofthissizeinvolvedalargenumberofcompute nodesandRAMtotrainon.
Thegeneratedembeddingswerethenusedasinputforvarious downstreampredictivetasks,suchaslinkpredictionbetweendrugs anddiseasesandsimilarity-basedanalysesforgeneprioritization. Allembeddings,modelconfigurations,andevaluationoutputs werestoredinstandardformatstosupportreproducibilityand re-analysis,suchas.parquetfiles,.graphand.csv.Thisentire pipelineallowsforeasyreproducibilityofgraphsaswellas applicationtoothertypesofknowledgegraphssuchasNeuroKG andPrimeKG.
Results
First,non-embedding-basedalgorithmswereappliedtothe biologicalknowledgegraphtoassesswhethermeaningfulresults couldbeobtainedwithoutextensivecomputationalresources.A metapath-basedrandomwalkalgorithmwasused,withthe underlyingmetapathdefinedas[‘disease’,‘drug’,‘gene/protein’] pathway.ThediseasenodewasinitializedastheParkinson's Diseasenodewiththehighestdegreeinthegraph,i.e.,the Parkinson'sDiseasenodewiththegreatestnumberofconnections. ThisselectionwasnecessarybecausemultipleParkinson'sDisease
nodesexistedinthegraph,correspondingtodifferentvariantsof thedisease.
Afterconductingtherandomwalk1,000timesandaggregatingthe terminalnodesofeachwalk,thefollowingfivegenesemergedas thetopresults,asdescribedinTable1:
Table1:Summarizingtop5genesfromametapath-basedrandom walkalgorithmwithmetapath[‘disease’,‘drug’,‘gene/protein’].
Thesegeneswerefurtherevaluatedusingabiologicallarge languagemodelfromHuggingFace(Labraketal.),aplatform whereopenmodelsareavailabletodownload,toassessthelink score(howlikelytheyaretobelinkedtoParkinson'sDisease)and thebiologyscore,representingbiologicallyspeaking,whichgenes aremostlikelytobeusefulwhenidentifyingpotentialtherapeutic targetsforPD).Evaluatingthesetwometricsisimportantbecause itnotonlyhighlightssignificantrelationshipsbetweenthegene anddisease,butalsodemonstratesthepotentialofeachgenetobe useful.Ultimately,thesegenescouldbepassedontoabiologistto conductexperimentsormoredeeplyinvestigatethegenomicsto evaluatewhetheragivengeneorproteinistrulyuseful.Inorderto isolategenesthathavealreadybeenstudied,aseparatelinkscore wasgeneratedtoquantifyhowmuchresearchhasbeenconducted connectingeachgenetoParkinson'sDiseaseoverthepast20 years.Thismetricwasprimarilymeasuredbythenumberof scientificarticlespublishedwithinthelasttwodecadesthat examinedtheconnectionbetweenParkinson'sDiseaseandthe gene.UsingtheBioMistrallargelanguagemodel(Labraketal.) withthespecifiedhyperparameters,thecorrespondingscoresfor eachgeneweregeneratedandhavebeensummarizedbelowin Table2:
Table2:SummarizingBioMistralLLMscoresforeachgenebeing relatedtoParkinson'sDiseaseandhavingpotentialtobe biologicallyrelevantwhentryingtodevelopacure.
Afteranalyzingthescores,bothCYP3A4andHTR1Astandoutas promisingcandidategenesforfurtherinvestigation.CYP3A4and CYP2D6arecytochromeP450enzymesthatplaycentralrolesin drugmetabolism,particularlyinthebreakdownofcompounds commonlyusedinneurologicaltreatments,makingtheir appearancebiologicallyplausibledespitelowerlinkscores.
HTR1A,aserotoninreceptor,isdirectlyinvolvedin neurotransmittersignalingpathwaysrelevanttomotorcontroland
dopaminergicregulation,whichareheavilyimplicatedin Parkinson'sDiseasepathology.Furthermore,HTR1Acanalsobe potentiallyleveragedasatherapeuticadjuncttargetforParkinson’s becauseofitsroleinmodulatinglevodopasideeffects,thus provingtheefficacyofthealgorithmtofindrealisticgeneswhich havepotentialtobeusedintryingtobetterunderstandanddevelop atherapyforParkinson'sDisease.
Otherrandomwalk–basedalgorithmsnotreliantonmetapaths werealsoconsidered.Afterexperimentingwithdifferent algorithms,abi-directionalpersonalizedPageRank(PPR) algorithmwasselectedtomeasurenodesimilarity.Asinthe previousmetapath-basedapproach,bi-directionalPPRallowsfor measuringsimilaritybetweenParkinson'sDiseaseandotherrelated nodes.However,inthiscase,thefocuswasshiftedtorelateddrugs ratherthangenesorproteinsinordertoisolatedrugsthatcouldbe usefulfortreatingParkinson'sDisease(e.g.,drugrediscoveryor drugrepurposing).
Whenexperimentingwith alpha (theparameterusedinPPRas explainedinthemethods), theoptimalvaluewasfoundtobe0.1. ThisvaluewasoptimizedbasedonpreviousiterationswithPage Rankinotherknowledgegraphs,aswellasevaluatingtheheatmap toseeifmultiplesimilaritylevelsarebeingreflected(wide spectrumofcolorsinheatmap).Afterrunningthisalgorithmand creatingaheatmapbasedonthescores,thefollowingheatmap diagramwasgeneratedforthetop51Parkinson'sDiseasenodes and13Parkinson'sDiseasedrugsintheknowledgegraph:

Figure3:Disease-DrugSimilarityHeatmapbasedonvariantsof Parkinson'sDiseaseviaNormalizedBi-DirectionalPersonalized PageRankScores(rightaxisshowssimilarity,mostsimilarbeing darkredandleastsimilarbeingdarkblue).
Figure3demonstrateshowtheknowledgegraphisbothaccurate. First,theheatmapshowstheaccuracyoftheknowledgegraphby theverticalstripswithintheheatmap.Thesestripscorrespondto somedrugbeingwidelyapplicabletomanydifferentvariantsof Parkinson'sDisease.AccordingtoFigure3,selegiline, pramipexoleandamantadineallshowsignsofbeinghighuse-case drugs,andthisisfurtherbackedupbymultiplearticles(Zahooret al.),(Li,Wentingetal.)anddatasourcespointingtoahighutility ofthesedrugswhendealingwithParkinson’s.Furthermore,the persistenceofsignalacrossgeneticallydis-similarvariantsof Parkinson’ssupportsthegraphusinganetwork-basedapproach
(withrandomwalks)totryandencapsulatesimilaritybetween nodes.
Thedisease-druganalysiscanalsobeappliedtootherdiseases withmorevariability,suchascancer.TheanalysisofFigure4 belowshowstheresultsofthesamealgorithmrunonthemost commoncancerdiseasesanddrugs.

Figure4:Disease-DrugSimilarityHeatmapbasedonvariantsof cancerviaNormalizedBi-DirectionalPersonalizedPageRank Scores
Similartothepreviousheatmap,Figure4demonstratesbothsigns ofvalidation(i.e.,evidencethattheheatmapcontainsaccurate information)andclearareasofinterestareobservable,wherea drugshowsasurprisinglyhighdegreeofcorrelationwithacancer, signalingapotentialopportunityfordrugrepurposing.The accuracyoftheheatmapisdemonstratedbythecenter-leftcluster
linkinggastrointestinalcancerswithclassicalcytotoxicagents. Thisclustershowsahighcorrelationbetweendrugssuchas docetaxelanddoxorubicinandgastricandsmallintestinecancers, relationshipsthathavebeenverifiedbynumerousscholarlyarticles overthelastdecade(Bekaii-SaabandVillalona-Calero).
Usefulembeddingswerealsogeneratedfromtheknowledgegraph tocalculatenodesimilarityandvisualizethegraph.Togenerate theseembeddings,arelationalgraphconvolutionalnetworkwas developed(Schlichtkrulletal.).Theuseofaconvolutional networkisimportantforenhancingthemodel’saccuracyof predictingneighboringnodes,therebycreatingmore comprehensiverepresentationsofeachnode.Additionally,the modelwasrelational,unliketraditionalgraphconvolutional networks,itaccountedformultipleedgetypesinthegraphby assigningeachedgetypeadistinctweightmatrix.Thisapproach enabledmorenuancedvectorrepresentationsbyincorporating variationacrossedgetypes,allowingthemodeltodistinguish betweenrelationshipssuchasadrugtreatingadiseaseversusthe samedrugcausingasideeffect.

Figure5:UMAPProjectionofPrimeKG(knowledgegraph) embeddingscreatedbyarelationalgraphconvolutionalnetwork.
Figure5showsaprojectionoftheembeddingsoftheknowledge graph.Whenthemodelembedsthegraph,theoutputconsistsof 64-dimensionalvectors,whicharedifficulttovisualizedirectly.As aresult,thesevectorsareprojectedintoatwo-dimensionalspace tofacilitatevisualization.Thefigureabovealsorepresentsthe projectedembeddingsoftheentireknowledgegraph,witheach distinctclustercorrespondingtoaspecificnodetype.Forexample, mostdiseasenodesinthegraphareclusteredtogetherinthe bottom-mostcluster.Thisfigurealsoshowsmanyisolatedpoints andshortlinesegments,sometimesappearingbetweenclusters. Thesecorrespondtogroupsofnodesintheknowledgegraphthat arerelativelyisolated,meaningtheyhavelowdegreeandfew connectionstoothernodes.Consequently,theirembeddedvectors alsoappearisolatedfromthelargerclusters.Pointsorline
segmentslocatedbetweentwoclustersrepresentnodesthatare similarlyconnectedtobothclusters.Forinstance,adrugthatlinks adiseasewithmanyassociatednodestoasideeffectwithmany associatednodesmayappearpositionedbetweenthelargerdisease andside-effectclusters.
Discussion
Figure3showstheusefulnessoftheknowledgegraphonmultiple levels,primarilygivingmotivationbehinddrug-repurposing.In contrasttoprimaryneurodegenerativeParkinson'sDisease, secondaryparkinsoniansyndromes,includingcarbon monoxide–inducedparkinsonism,cyanide-inducedparkinsonism, postencephaliticParkinsondisease,andsecondaryParkinson disease,showextremelysimilarityprofiles(heatmapcoloring). Thissimilarityisdemonstratedbyeachoftheirsimilarlightred colorsintheheatmap.Thus,thispointstoeachsyndromereacting verysimilarlytodifferentdrugs.Thisresultisusefulbecauseit meansthatifthereexistsausefulandimpactfuldrugfortackling carbonmonoxide–inducedparkinsonism,thenthereisahigh probabilitythatthesamedrug,orsomevariantofthatdrugcanbe usedeffectivelyonothersyndromeswithsimilarprofiles,like postencephaliticParkinsondisease.Thisremainsextremely importantbecausecreatingandvalidatingadrugforlegaluseis oftenanextremelyexpensive,laboriousandtime-consuming process.Tobeabletouseanexistingdrugandslightlymodifyit savesmuchmoneyandtimeintheentireprocess.
Figure3alsoshowspotentialfordrugstobeusedtotackleother variantsofParkinson’sthroughcolumnanalysisaswell.For example,thereexistsingularextremelyredrectangles,indicatinga highsimilaritybetweendruganddiseasenodes.Thishappens4 timeswithintheheatmap,twicefortheAmantadinedrug.This findingsuggestsinvestigatingthechemicalmakeupofthedrug,
andseeinghowitcanpotentiallybeusedfordrug-repurposing. Specifically,thedrugexhibitsastrongsimilaritysignalfor postencephaliticParkinsondisease,exceedingthatobservedfor severalotherPDsubtypes.Thisfindingisnotablegiven amantadine’sNMDAreceptorantagonism,antiviralhistory,and anti-dyskineticeffects,whichalignwithinflammatoryand excitotoxiccomponentsimplicatedinpostencephaliticsyndromes. Theselectiveelevationofthissignalsuggeststhatthenetwork capturesnon-dopaminergicmechanismsthatdifferentiate postinfectiousparkinsonismfromidiopathicPD.Overall,this showstheeffectivenessofthegraphtobeabletofinddrugswith clearpotentialfordrugrepurposing.
Asdescribedintheresultssection,theheatmapinFigure4clearly reflectswell-establishedrelationshipsbetweendiseasesanddrugs. However,theheatmapalsorevealsahighcorrelationbetween peripheralnervoussystemcancerandcarmustine,arelationship thathasnotbeenextensivelystudiedinthepast15years.This novelfindinghasahighpotentialtobeextremelyusefulfordrug repurposingandshowstheapplicabilityandpowerof encapsulatingdatainaknowledgegraph.
Theembeddingsoftheknowledgegraph,asvisualizedinFigure5, canalsobeusedeffectivelytocalculatenodesimilarity(Giannis Nikolentzosetal.).Whencreatingembeddings,theneuralnetwork takesinaccounttheneighborhoodofeachnode,andthusoutputs nodeswithsimilarneighborhoods(eg.hasasimilarprofileofside effectrelationships,oralsohasaconnectiontoParkinson's Disease)tosimilaroutputvectors.BoththeEuclideandistanceand cosinesimilaritymetricremainaccuratewaystocalculatehow similartovectorsare.
TheEuclideandistanceisbetweentwovectors,
, isdefinedasequation1below:
Inthe2dimensionalplace,thisisjustequivalenttotakingthe shorteststraightlinebetweentwopoints.Thisformulaabove essentiallymeasuresthedistancebetweentwovectorsthrough Pythagoras’Theorem,calculatingthesquarerootofthedifference ofeachvectorcoordinatesquared.


Thecosinesimilaritymetricbetweentwovectors and is simplydefinedinequation2below:


where denotestheanglebetweenthetwovectors.Thismatrix givesasimilarityfrom0to1bytakingaccounttheanglebetween twovectors,where1representsthevectorsbeingthesameand0 representsthevectorsbeingorthogonal.
Boththesemetricswereusedtofindsimilaritybetweentwonodes bysimplycalculatingthedistancebetweentheircorresponding vectors.Ultimately,whenappliedtothesamelistofParkinson's DiseasesasseeninTable1,thisalgorithmalsodemonstrateda strongsimilaritybetweenHTR1AandParkinson'sDisease, supportingthepreviousclaimforHTR1Atobesignificantly Parkinson'sDiseaseandusefulwhentryingtodevelopatherapy.
Conclusion
Inconclusion,large-scalebiomedicalknowledgegraphscanserve asareliableandbeneficialfoundationforbiologicaldiscoveryand drugrepurposing.Theknowledgegraphisabletoconsistently uncoverknownrelationshipsbetweendiseasesanddrugsaswellas betweendiseasesandgenesthroughvariousanalyticalparadigms, suchasmetapath-basedrandomwalks,bidirectionalPersonalized PageRank,andrelationalgraphneuralnetworks.Thefactthatthe high-scoringresultsbasedonalgorithmsareinlinewithclinical usageandbiologicalmechanismsshowsthatthegraphisableto capturesignificantbiomedicalstructureandnotjusterroneous correlations.Evenso,thediscoveryoftenablebutunderstudied correlationsuncoversitsstrengthsasageneratorofhypotheses ratherthanaconfirmatorysystemthatisstrictlyso.
Theapplicabilityofthisframeworkalsoextendsbeyonddrug repurposingforParkinson'sDiseaseandcancer.Thesame graph-basedinferencetechniquescanbereadilyadaptedtoother biomedicaltasks,includinggeneprioritizationforrarediseases, adverseeventprediction,andtheintegrationofclinicaldatasuch aselectronichealthrecordsorpatientstratificationreports.By incorporatinglongitudinalclinicaloutcomes,population-level phenotypes,orreal-worldevidence,futureextensionsofthegraph couldsupporttranslationalresearchthatbridgesmolecularbiology andclinicaldecision-making.
Someofthelimitationswhencreatingthisgraphincludethe variabilityindata;thisincludednumerousinstanceswhendata fromdifferentsourceswouldcontradicteachother,aswellas whereportionsofthedatawouldbefullyremovedfromthegraph. Furthermore,itisalsoimportanttoconsiderthebiasinthe knowledgegraph;theentiregraphisbuiltfrommanyofthemost populardatasources,thusresultinginmoreinformationinthe
graphbeingfocusedonmorepopulardrugsthanlesspopular drugs.Thisresultsintherebeingsomebiastowardsmorepopular useddrugswhenrunningalgorithmssuchasrandomwalks.
However,theknowledgegraphstillremainsapivotalmethodof representationdata,andthisthesiscontributestoshowingthat network-basedrepresentationsareessentialforunderstanding complexbiologicalsystems.Asbiomedicaldatacontinueto increaseinscaleandheterogeneity,knowledgegraphsoffera principledwaytointegratediversemodalitiesandenablemachine learningmodelstoreasonoverstructuredbiologicalknowledge. Theresultspresentedheresuggestthatsuchgraphswillplayan increasinglyimportantroleinacceleratingtherapeuticdiscovery, reducingdevelopmentcosts,andguidingexperimentalvalidation. Inthissense,biomedicalknowledgegraphsarenotmerely analyticaltools,butfoundationalinfrastructureforthefutureof data-drivenmedicine.
Bibliography
BusraOguzoglu.“CreatingEmbeddingsfromKnowledge GraphsUsingGraphNeuralNetworks.” Medium,18Nov. 2024, medium.com/@busra.oguzoglu/creating-embeddings-fromknowledge-graphs-using-graph-neural-networks-ffc6cc622 75c.
Bekaii-Saab,TaniosS.,andMiguelA.Villalona-Calero. “PreclinicalExperiencewithDocetaxelinGastrointestinal Cancers.” Seminars in Oncology,vol.32,1Oct.2005,pp. 3–9, www.sciencedirect.com/science/article/pii/S009377540500 1466,https://doi.org/10.1053/j.seminoncol.2005.04.002.
Drugbank.“DrugBank.” Go.drugbank.com,2023, go.drugbank.com.
Gallo,Denis,etal.“AdvancesinDatabaseTechnology.” EDBT2020,volume2020-March,pp.447-450.
GiannisNikolentzos,etal. Matching Node Embeddings for Graph Similarity.Vol.31,no.1,13Feb.2017, https://doi.org/10.1609/aaai.v31i1.10839.Accessed28May 2023.
Labrak,Yanis,etal.“BioMistral:ACollectionof Open-SourcePretrainedLargeLanguageModelsfor MedicalDomains.” ArXiv.org,2024, arxiv.org/abs/2402.10373.
Li,Wentingetal.“Comparisonoftheeffectiveness,safety, andcostsofanti-Parkinsondrugs:Amultiple-center retrospectivestudy.” CNS neuroscience & therapeutics vol. 30,4(2024):e14531.doi:10.1111/cns.14531.
Noori,Ayush,etal.“Metapaths:SimilaritySearchin HeterogeneousKnowledgeGraphsviaMeta-Paths.” Bioinformatics,vol.39,no.5,1May2023,
https://doi.org/10.1093/bioinformatics/btad297.Accessed6 Mar.2025.
TomazBratanic.“ConstructaBiomedicalKnowledge GraphwithNLP.” Medium,TDSArchive,25Oct.2021, medium.com/data-science/construct-a-biomedical-knowled ge-graph-with-nlp-1f25eddc54a0.Accessed20Jan.2026.
Schlichtkrull,Michael,etal.“ModelingRelationalData withGraphConvolutionalNetworks.” ArXiv:1703.06103 [Cs, Stat],26Oct.2017,arxiv.org/abs/1703.06103.
Zahoor, Insha, et al. “Pharmacological Treatment of Parkinson's Disease.” Parkinson's Disease: Pathogenesis and Clinical Aspects, vol. Chapter 7, no. 1, 21 Dec. 2018, pp. 129–144, www.ncbi.nlm.nih.gov/books/NBK536726/, https://doi.org/10.15586/codonpublications.parkinsonsdisea se.2018.ch7.
Z.Ye,etal."AComprehensiveSurveyofGraphNeural NetworksforKnowledgeGraphs,"inIEEEAccess,vol. 10,pp.75729-75741,2022,doi: 10.1109/ACCESS.2022.3191784.
Zhang,Yuan,etal.“BioKG:AComprehensive, High-QualityBiomedicalKnowledgeGraphfor AI-Powered,Data-DrivenBiomedicalResearch.” BioRxiv (Cold Spring Harbor Laboratory),17Oct.2023, https://doi.org/10.1101/2023.10.13.562216.
