BROAD STREET SCIENTIFIC
VOLUME 11 | 2021-22
The North Carolina School of Science and Mathematics Journal of Student STEM Research
Front Cover The Aizawa attractor derives from a set of equations that evolve to map out a threedimensional shape when repeatedly applied on three-dimensional coordinates. The attractor forms this beautiful sphere with a tube-like structure running down the y-axis. Credit: Cedrick Argueta (https://www.cedrick. ai/posts/attractors.html)
Biology Section Proposed by Rene Thomas, the equations of motion that lead to this attractor are cyclically symmetric in the x, y and z variables. These equations have also been used to model the trajectory of a frictionally damped particle moving in a 3D lattice of forces. Credit: Cedrick Argueta (https://www.cedrick. ai/posts/attractors.html)
Chemistry Section This image traces the path of a double pendulum, a simple dynamical system that exhibits complex, chaotic behavior that is extremely sensitive to initial conditions. A double pendulum consists of two point masses at the end of light rods that are joined together. The system then oscillates freely, creating patterns that are nearly impossible to predict. Image credit: MOD (https://mod.org.au/exhibits/chaos-machine/)
Engineering Section The Clifford Attractor is also known as the fractal dream attractor. Upon closer inspection, one can see that the lines that make up the attractor are actually a bunch of separate lines. No matter how far you zoom in, the lines keep separating and there are more to count. Credit: Cedrick Argueta (https://www.cedrick. ai/posts/attractors.html)
Mathematics and Computer Science Section A Lorenz Attractor is a set of chaotic solutions to the fluid dynamics equations that represent rolling fluid convection. The complex behavior of this phenomenon became the first known example of a strange attractor. Image credit: Wikimol, via Wikimedia Commons (https://en.wikipedia.org/wiki/Chaos_ theory#/media/File:Lorenz_attractor_yb.svg)
Physics Section This is another example of a pattern created by a double pendulum, illustrated using a light source attached at the free end of the pendulum. This sensitive dependence on initial conditions in which a small change results in large differences later on is representative of the butterfly effect, such that each pattern created is unique. Image credit: Cristian V., via Wikimedia Commons (https://commons.wikimedia.org/ wiki/File:Chaos_Theory_%26_Double_Pendulum_-_4.jpg)
TABLE of CONTENTS 4
Letter from the Chancellor
5
Words from the Editors
6
Broad Street Scientific Staff
7
Essay: The Beauty of Nature Inspired Computation MELODY LEE, 2023
10
Photography: Flora, Fauna, Phenomena GRACE AMANTEA, 2023
Biology 11
Signaling Pathways During Wound Closure in Planarian Epithelium LEVI CRUZ, 2022
17
Cheminformatics Analysis of the N-methyl-D-aspartate (NMDA) Receptor Neurotoxicity SAMANYU DIXIT, 2022
23
Investigation on Joint Usage of GSK1016790A and OPC-31260 on Restoring Mating Ability and Raising Intracellular Calcium Levels in lov-1 and pkd-2 Single and Double Mutant Caenorhabditis elegans GRACIE LIN, 2022
Chemistry 32
Selective and Cost-Effective Depolymerization of Linear Polyethylene via Tandem Dedrogenation and Metathesis Dual Catalyst System MEGHANA CHAMARTY, 2022
39
Applying Quantum Computational Methods to Analyze the Reaction Progression of an Sn2 Functional Group Transformation Reaction SOPHIE SCHERER, 2022 ONLINE
44
Rational Design and Synthesis of a Novel Class of Boronic-Acid Containing Tubulin Inhibitors as Tumor Vascular Disrupting and Antiproliferative Agents WINNIE WANG, 2022
Engineering 57 Development of a Bioactive, Biodegradable, and Variable-Density 3D Printer Filament for Patient-Specific Bone Reconstructive Implants JACOB ROSE, 2022
63
Ziegler-Nichols Tuning Implementation on Arduino-Based PID Controller for DC Motor Rotation PRACHEETI SHIKARKHANE, 2022 ONLINE
Mathematics and Computer Science 71
Minimum Number of Triangles with Integer Side Lengths in a Saturated Arrangement TAKUMI FUJITA AND SEAN KIM, 2022
76
Mapping the Landscape of RNA-Sequencing Tools SHERRY LIU, 2022
Physics 83
Simulating Quantum Key Distribution in Three Polarization Bases KATHERINE PANEBIANCO, 2022
90
Effects of Underlayment Roughness and Angle on Granular Chute Flows NOAH SIEKIERSKI, 2022
Featured Article 94
An Interview with Dr. Amay Bandodkar
LETTER from the CHANCELLOR “Education is the foundation for all we do in life. It shapes who we are and what we aspire to be. Creativity fuels innovation, and it's what all states should strive to instill in the next generations.” ~The Honorable James B. Hunt, Jr. I am proud to introduce the eleventh edition of the North Carolina School of Science and Mathematics’ (NCSSM) scientific journal, Broad Street Scientific. Each year students at NCSSM conduct significant scientific research, and Broad Street Scientific is a student-led and student-produced showcase of some of the impressive research being done by our students. As we prepare to open a second NCSSM campus in Morganton this spring, I continue to be grateful for the opportunities for innovation and exploration that we have provided to students across North Carolina for four decades now. Each year we learn of the outstanding, impactful research and innovation of our alumni, while our current students themselves are pushing into more complex areas of inquiry. Though the second year of the global pandemic continued to present challenges to all educational institutions, it in no way dampened the spirit of exploration that draws bright, young minds to our programs year after year. These student researchers have remained unfailingly focused and positive, and will undoubtedly change our state and the world for the better in the years to come. Indeed, that work has already begun. Opened in 1980, NCSSM was the nation’s first public residential high school where students study a specialized curriculum emphasizing science and mathematics. Teaching students to do research and providing them with opportunities to conduct high-level research in biology, chemistry, physics, computational science, engineering and computer science, math, humanities, and the social sciences are critical components of NCSSM’s mission to
4 | 2021-2022 | Broad Street Scientific
educate academically talented students to become state, national, and global leaders in science, technology, engineering, and mathematics. I am thrilled that each year we continue to increase the outstanding opportunities NCSSM students have to participate in research. Each edition of Broad Street Scientific features some of the best research students conduct at NCSSM under the guidance of our outstanding faculty and in collaboration with researchers at major universities and other research institutions. For thirty-seven years, NCSSM has showcased student research through our annual Research Symposium each spring and at major research competitions such as the Regeneron Science Talent Search and the International Science and Engineering Fair. Even during the past two years, we were able to provide these opportunities to our students in a virtual format. The publication of this journal provides another opportunity to share with the broader community the outstanding research being conducted by NCSSM students. I would like to thank all of the students and faculty involved in producing Broad Street Scientific, particularly faculty sponsor Dr. Jonathan Bennett, and senior editors Vish Ravichandran, Hrishika Roychoudhury, and Lucia Wang. Explore and enjoy! Dr. Todd Roberts Chancellor
WORDS from the EDITORS Welcome to the Broad Street Scientific, NCSSM’s official journal of student research in science, technology, engineering, and mathematics. In this eleventh edition of Broad Street Scientific, we aim to showcase the breadth and depth of research conducted by our students and encourage readers to be scientifically literate individuals. This year, we are especially proud of how our student researchers pressed on with their quest for knowledge with unwavering dedication and enthusiasm, despite the challenges that COVID-19 has continued to pose around the world. During these times, it is essential to take the initiative to keep up with new advancements so that we can hold informed perspectives and make informed decisions. That is exactly what these student researchers have done, and we are excited to showcase their hard work in this volume. We hope you enjoy this year’s edition! This year’s theme is chaos theory: a field of mathematics and the sciences studying disordered systems. The butterfly effect is a layman’s term for describing chaos—a butterfly flapping its wings could cause a hurricane thousands of miles away—small changes in a chaotic system can cause wildly varying outcomes. In a chaotic system, outcomes are unpredictable. This happens not because the system doesn't follow deterministic laws but because the initial conditions are near impossible to determine with complete accuracy and because the long term behavior of the system is extremely sensitive to the initial conditions. However, with advances in computation, chaotic systems can more easily be simulated and visualized. When graphing these chaotic systems, they often converge towards a set of states, creating the visually-striking attractors showcased in this edition’s themed images. This idea of creating beauty from chaos teaches us to embrace the unexpected and take the spirit of exploration with us wherever we go. In a year as unpredictable and unprecedented as this one has been, it is more important than ever that we strive to find beauty within the chaos around us. Despite our world still dominated by an ever-evolving pandem-
ic, scientific research has continued to develop in large strides, demonstrating how well humanity can come together to achieve common goals and support one another through trials and tribulations. We hope this journal sheds light on some of the amazing research conducted by the students of NCSSM as they take it upon themselves to accept the greater challenge and expand the frontiers of their knowledge. As NCSSM prepares to open its second campus in Morganton, we must keep this love of learning and drive for progress in mind as we continue to forge the way ahead for student research in STEM. We would like to thank the faculty, staff, and administration of NCSSM for continuing to support and build the scientific community that NCSSM represents. Now in its 42nd year, NCSSM continues to nurture a stimulating academic environment that encourages motivated students to apply their interests towards solving real-world problems. To many, NCSSM serves as a symbol of passion and determination in the next generation of young people who will no doubt change the world. We give special thanks to Dr. Jonathan Bennett for his invaluable support and guidance throughout the publication process. We would also like to thank Chancellor Dr. Todd Roberts, Dean of Science Dr. Amy Sheck, and Director of Mentorship and Research Dr. Sarah Shoemaker. Lastly, we would like to acknowledge Dr. Amay Bandodkar, Assistant Professor at NC State’s Electrical and Computer Engineering Department, for speaking with us about his unique career path and perspectives in STEM and imparting important advice to young scientists so that they, too, may tread fearlessly and find their own path forward. Lucia Wang, Hrishika Roychoudhury, and Vish Ravichandran Editors-in-Chief
Broad Street Scientific | 2021-2022 | 5
BROAD STREET SCIENTIFIC STAFF Editors-in-Chief
Vish Ravichandran, 2022 Hrishika Roychoudhury, 2022 Lucia Wang, 2022
Publication Editors
Sreenidhi Elayaperumal, 2022 Spriha Manjigani, 2023 Shweta Shah, 2023 Esha Singaraju, 2023 Sara Talekar, 2022 Allison Zhang, 2023
Biology Editors
Chemistry Editors
Engineering Editors
Mathematics and Computer Science Editors
Physics Editors
Faculty Advisors
6 | 2021-2022 | Broad Street Scientific
Antonio Alonso-Stepanov, 2023 Linden James, 2023 Suhani Ramchandra, 2022
Aislin Eberhardt, 2023 Eliot Ferriol, 2023 Mihika Gottimukkala, 2023
David Kim, 2023 Mariah Snuggs, 2022
Melody Lee, 2023 Vishakh Sandwar, 2022 Helen Wu, 2023
Aryan Goyal, 2023 Sahil Patel, 2023
Dr. Jonathan Bennett
THE BEAUTY OF NATURE INSPIRED COMPUTATION Melody Lee Melody Lee was selected as the winner of the 2022 Broad Street Scientific Essay Contest. Her award included the opportunity to interview Dr. Amay Bandodkar, Assistant Professor in the Electrical & Computer Engineering Department at North Carolina State University. Nature is an inescapable construct. A careful investigation of what nature has inspired reveals a complex network of intertwined stories between mankind and nature. According to an overview of The History of Mankind, by Professor Friedrich Ratzel, such tales are said to consist of “the rise of civilisation, and the development of language, religion, science and art, and family and social customs” (Ratzel, 1899). While this description skims over the precise role nature has played, it is clear humans have evolved with the times under the hand of environmental circumstances over the course of thousands of years (Seymour, 2016). Technology, on the other hand, is a more recent invention of civilized society – a manmade advancement seemingly far removed from nature. In the last century or so, high-level calculations using computers have become possible through the leaps and bounds made in technological advancements. Computations have transitioned from use of the abacus, to Babbage’s Analytical Engine, and then to modern computers (Encyclopædia Britannica, n.d.). Since then, processing capacities have increased exponentially, a pattern noted in the 1960’s by Gordon E. Moore, alongside his namesake law (Moore, 1965). Despite these advancements, the solutions to certain problems we as a part of the human population seek to solve continue to elude researchers. The continual evasion of a definite solution (in the form of algorithms) is due in part to several reasons. On the one hand, it is rather difficult to solve certain problems without a brute-force algorithm of some sort. As a result, the estimated time required for computation grows to unmanageable lengths, often exceeding several hundred thousand human lifetimes. Such problems are classified as NP-complete or NP-hard problems, in reference to the amount of time required relative to the size of the input data (Hochba, 1997). Many of the problems that fall under this class are also a type of optimization problem, meant to maximize or minimize some set of results. Take, for example, a set of points where the distances between each point are defined as some positive, nonzero values. The dilemma is as follows: how could one best visit every single point in this set, such that the total distance traveled is minimized? Known as the Traveling Salesman problem, iterations
of the dilemma are applicable to logistical operations that fall under many human industries (including door-to-door marketing and communications). The ideal method for finding the perfect solution generally involves attempting every possible travel route, which is nearly impossible to finish in a logical interval of time, even with the help of supercomputers. Consequently, others have sought to utilize approximations to get close, but generally inexact, solutions (Karlin, Klein, & Gharan, 2021). This is where nature steps in. It is clear that nature has been a driving force in human society. So, naturally, what follows is this question: what precisely has nature inspired in this world of computational science, specifically? The response is a series of beautiful algorithms aimed at mimicking patterns seen in the natural world (Fister, et al., 2013). To briefly illustrate its significance as a source of inspiration, one could argue that nature is inclined to optimize the situation regardless. After all, following billions of years of evolution and subjugation to the laws of nature, the most efficient or least energy consuming processes are more likely to prevail (Smith, 1978). In the case of biological constructs, ineffective solutions to problems encountered on the regular (including those relating to motor skills and instincts) could potentially be a cause for die-offs in the face of hardship (Smith, 1978). As for examples of nature enacting an optimal result, simply consider a ball of water in zero gravity, where the water tends to minimize the surface area. Consider, now, the paths of ants on the lookout for food. The world is quite broad, even for humans whose strides are at least several times the length of that of these small insects. Consequently, ants must travel from destination to destination in search for resources without a significant loss in time (which corresponds to minimizing the distance traveled). In order to remedy the difficulties that may go along with this, ants will lay down pheromones as natural traffic cones to direct other ants and, more importantly, leave a record of places they have visited (Chalissery, et al., 2019). In the case one ant encounters a pleasant surprise somewhere along their path, they may use these pheromones to indicate to other ants to help them bring back what they have found to the colony (Chalissery, et al., 2019). This simple, yet elegant, method of “remembering” Broad Street Scientific | 2021-2022 | 7
the paths proves to be quite effective in illustrating travel. We thus arrive back again at the Traveling Salesman problem. A brute force solution for this problem would be the equivalent of an ant leaving the colony, choosing a path to walk, visiting all destinations, and then walking all the way back, retrying again and again with a different path each time in order to find the best path. Even for an entire colony of ants, it would not make any sense to request every individual ant go out in an attempt to test a series of paths over and over. The resultant solution to cut down on the work lies in the use of pheromones, the computational equivalent of which involves the storage of the “result” for some subsection of the path in the memory. This dynamic storage of what information has already been gathered (and what paths have been traversed) cuts down significantly on the running time of the algorithm and, therefore, presents a much better solution than sheer brute force (Sudholt & Thyssen, 2011). This solution, therefore, is beautiful in the sense that it takes from nature and converts these elements into points of data – mere numerical entities with a sudden flair of character. Perhaps, then, it follows that other algorithms, too, have sought to use data to mimic the mannerisms of natural phenomena, a tactic not confined to merely the biological world. Namely, in recent decades, the spotlight on the quantum community has grown. On the one hand, researchers have proposed the plausibility of utilizing quantized particles – called qubits– and their probability-based characteristics to perform calculations (Katwala, 2020). The advantage these types of quantum computers have has been coined “quantum speedup”, in reference to the ability to compute a similar result to a problem at a rate several times that of typical, classical computers (Rønnow, et al., 2014). However, one may argue this is the equivalent of hiring an army of ants to walk every path known to man – a plausible prospect, but preferably avoided if possible. Likewise, the maintenance of a significant number of qubits in low temperature, isolated conditions so as to avoid decoherence (loss of information) is consuming on all fronts (Pakin, 2019). Furthermore, it is yet unproven as to whether or not quantum supremacy is definite, and therefore it follows that alternatives are being sought after. The typical routes for computation (using the binary-based computers) have taken these quantum characteristics as inspiration, and nature strikes yet again. Indeed, quantum-inspired classical algorithms made waves in the computing field, as they challenged what were previously believed to be “untouchable” capabilities of quantum computers (Tang, 2019). Take, for instance, the quantum-inspired algorithm developed by Ewin Tang; unlike the ant-based algorithm, it was centered around the development of a recommendation system. However, unlike other existing 8 | 2021-2022 | Broad Street Scientific
algorithms, it managed to dramatically cut down on the running time needed to formulate a result by mimicking the properties of qubits (Tang, 2019). The applications of these nature-inspired systems, in both optimization and general large-scale data processing, makes these algorithms valuable in the context of an ever expanding world. Traffic control, infrastructural development, or even entertainment recommendation systems all pose as areas with undiscovered solutions nature has the potential to infiltrate. Of course, these are only several of the many possibilities out there. As no one yet knows the secrets of the universe, further investigation into the patterns and phenomena that manifest themselves in nature is a worthy investment. Perhaps the areas of computation and natural science are not quite as far apart as believed. Who knows? Perhaps their pairing may one day go as far as save the world from strife. References Chalissery, J. M., Renyard, A., Gries, R., Hoefele, D., Alamsetti, S. K., and Gries, G. (2019, November 1). Ants sense, and follow, trail pheromones of ant community members. Insects. Retrieved January 11, 2022, from https://www. ncbi.nlm.nih.gov/pmc/articles/PMC6921000/ Encyclopædia Britannica, inc. (n.d.). History of Computing. Encyclopædia Britannica. Retrieved January 11, 2022, from https://www.britannica.com/technology/computer/History-of-computing Fister Jr., I., Yang, X.-S., Fister, I., Brest, J., and Fister, D. (2013, July 16). A brief review of nature-inspired algorithms for optimization. arXiv.org. Retrieved January 11, 2022, from https://arxiv.org/abs/1307.4186 Hochba, D. S. (Ed.). (1997). Approximation algorithms for NP-Hard problems. ACM SIGACT News, 28(2), 40–52. https://doi.org/10.1145/261342.571216 Karlin, A. R., Klein, N., and Gharan, S. O. (2021, May 8). A (slightly) improved approximation algorithm for Metric TSP. arXiv.org. Retrieved January 11, 2022, from https:// arxiv.org/abs/2007.01409 Katwala, A. (2020, March 5). Quantum computing and quantum supremacy, explained. WIRED UK. Retrieved January 11, 2022, from https://www.wired.co.uk/article/ quantum-computing-explained Moore, G. E. (1965). Cramming More Components Onto Integrated Circuits. Electronics, 38(8). https://doi.org/ https://hasler.ece.gatech.edu/Published_papers/Technology_overview/gordon_moore_1965_article.pdf
Pakin, S. (2019, June 10). The Problem with Quantum Computers. Scientific American Blog Network. Retrieved January 11, 2022, from https://blogs.scientificamerican. com/observations/the-problem-with-quantum-computers/ Ratzel, F. (1899, July 20). The History of Mankind. Nature News. Retrieved January 11, 2022, from https://www.nature.com/articles/060269a0 Rønnow, T. F., Wang, Z., Job, J., Boixo, S., Isakov, S. V., Wecker, D., Martinis, J. M., Lidar, D. A., and Troyer, M. (2014, June 19). Defining and detecting quantum speedup. Science. Retrieved January 11, 2022, from https://www. science.org/doi/10.1126/science.1252319 Seymour, V. (2016, November 18). The human-nature relationship and its impact on health: A Critical Review. Frontiers in public health. Retrieved January 11, 2022, from https://www.ncbi.nlm.nih.gov/pmc/articles/ PMC5114301/ Smith, J. M. (1978). Optimization theory in evolution. Annual Reviews. Retrieved January 11, 2022, from https:// www.annualreviews.org/doi/abs/10.1146/annurev.es .09.110178.000335?journalCode=ecolsys.1 Sudholt, D. & Thyssen, C. (2011, June 12). Running time analysis of ant colony optimization for Shortest Path Problems. Journal of Discrete Algorithms. Retrieved January 11, 2022, from https://www.sciencedirect.com/science/ article/pii/S1570866711000682 Tang, E. (2019, May 9). A quantum-inspired classical algorithm for recommendation systems. arXiv.org. Retrieved January 11, 2022, from https://arxiv.org/abs/1807.04271
Broad Street Scientific | 2021-2022 | 9
FLORA, FAUNA, PHENOMENA Grace Amantea Grace Amantea was selected as the winner of the 2022 Broad Street Scientific Photo Contest. Her award included the opportunity to have some of her photographs featured in the 2022 volume of the Broad Street Scientific. Grace's photographs are shown below.
An Elk-cellent View This piece features a cow (female elk) and her calf in their natural habitat--the grassy fields of Great Smoky Mountains National Park, where they feed-in Cataloochee, North Carolina.
Little Things This piece features moss hidden in the sunny crook of a tree branch on Mount Mitchell, encouraging all to take the time to observe and appreciate the "little things" the natural world has to offer.
10 | 2021-2022 | Broad Street Scientific
SIGNALING PATHWAYS DURING WOUND CLOSURE IN PLANARIAN EPITHELIUM Levi Cruz Abstract Tissue regeneration is a biological process that reverses otherwise permanent wounds. However, true regenerative capacity is relatively limited among species. In order to investigate molecular pathways important for tissue regeneration, planarians were modeled for regeneration. We measured the effect of certain signaling pathways in regeneration by treating planarians that received an anterior-posterior cut with spring water, dimethyl sulfoxide (DMSO), BMP inhibitor, Mdm2 inhibitor, or Pan-Akt inhibitor solutions for 24 hours. During a seven-day time course, the blastemas were measured at twenty-four-hour intervals. The average rate of growth for the anterior blastema was 0.071 mm/hr for spring water, 0.064 mm/hr for DMSO, 0.038 mm/hr for BMP inhibitor, 0.076 mm/hr for Mdm2 inhibitor, and 0.050 mm/hr for Pan-Akt inhibitor. The average rate of growth for the posterior blastema was 0.079 mm/hr for spring water, 0.067 mm/hr for DMSO, 0.040 mm/hr for BMP inhibitor, 0.078 mm/hr for Mdm2 inhibitor, and 0.047 mm/hr for Pan-Akt inhibitor. There was a significant difference in both anterior and posterior segments for BMP inhibition versus DMSO. There was a significant decrease in the posterior segment for Pan-Akt versus DMSO. Here, it was shown that the Akt pathway is involved in regeneration on the posterior side and not the anterior side of regeneration of the blastema in planarians. Amplification of the Akt pathway at the site of the wound could facilitate a quicker onset of regeneration. 1. Introduction 1.1 Tissue Regeneration Neoblasts, a type of stem cell, are responsible for the extent of regeneration in planarians; however, understanding the path from a stem cell to a differentiated cell to a new tissue demonstrates the whole organism’s response to a wound. The process is initiated with the infliction of a wound that is later covered with old epidermis tissue [3]. This first step is similar to what humans experience as the formation of a scab or scar. Yet, the process is vastly different from humans when neoblasts begin to migrate to the wound site and occasionally differentiate during their journey [3]. The migration of neoblasts highlights that the process is not solely a localized event and instead requires an effort throughout the organism. The accumulation of neoblasts at the wound site converts into an unpigmented blastema that will eventually differentiate into the required tissue [3]. Yet, planarians are susceptible to environmental factors, and when introduced to enough stress, planarians are known to spontaneously fission [1]. The entire regeneration process is a body-wide effort that requires the location and information from the wound site, the migration of neoblasts, and the formation of a blastema that will eventually differentiate.
Akt pathway can lead to cancers and diabetes mellitus [5]. In planarians, the Akt pathway regulates stem cell proliferation, cell death, and maintains differentiated tissues [9]. Therefore, it is a valid candidate for a signaling pathway involved in the initiation of tissue regeneration. Its direct involvement with differentiated tissues and stem cells means that it is possible for it to be a vital aspect of the regeneration process. p53, a downstream target of Akt, is responsible for regulating proliferation, and without p53, planarians demonstrate hyper-proliferation [8]. Mdm2 is a downstream protein of the Akt pathway and regulates p53 through inhibition, which provides an option in controlling tissue proliferation (Figure 1, Figure 2) [2].
Figure 1. Timeline of regeneration in a Pan-Akt inhibited sample.
1.2 Akt Signaling Pathway In mammals, the Akt signaling pathway plays a role in cell survival and development, proliferation, migration, and stem cell development [5]. Overamplification of the BIOLOGY
Broad Street Scientific | 2021-2022 | 11
ular tissue [10]. With the introduction of stem cells into the wound site and the amplification of the Akt pathway, quick and complete recovery from injuries could be possible. Amplifying the Akt pathway would possibly allow for the stem cells to migrate to the site of the wound at a quicker rate to initiate the overall process of regeneration more quickly. This treatment would prevent unnecessary scar tissue from forming.
Figure 2. Representation of the Akt pathway and expected results from inhibition. 1.3 Planarian Model Unlike humans, planarians exemplify tissue regeneration capabilities and display a strong ability to do so throughout all their tissue [11]. Planarians utilize their pluripotent adult stem cells called neoblasts to regenerate the required tissue at the wound site in a short amount of time [11]. However, the regeneration process is not solely dependent on stem cells because they require a stimulus to begin their migration to the wound site [3]. At the wound site, the differentiated epithelial cells are responsible for initiating the process through cellular communication and providing positional cues for the neoblasts [4]. Additionally, the bone morphogenetic protein (BMP) pathway plays a critical role in the initiation of tissue regeneration. For instance, this pathway establishes and maintains the dorsal-ventral axis during regeneration [6]. Without the BMP pathway, planarians develop fewer dorsal markers and present them on ventral surfaces instead [6]. The lack of the BMP pathway also leads to smaller blastemas [6]. The BMP pathway also serves as a signal to neoblasts to arrive at the wound that comes from the interaction of dorsal and ventral epithelium [7]. 1.4 Regeneration Post-Injury Application Myocardial infarctions, commonly known as heart attacks, are detrimental disorders that damage myocardial cells and cause the human body to respond with collagenous scar tissue that hinders heart function [10]. Over 1 million Americans are affected by this disorder and live with the risk of heart failure; therefore, current treatments focus on controlling the remodeling of the heart to prevent dangerous new circuits or dilation of ventricles [10]. The scar tissue formation is due to the human’s inability to regenerate tissue adequately; therefore, the heart cannot replicate its pre-heart attack efficiency in electrophysiology [10]. This specific disorder highlights humans’ overall inability to regenerate tissue and instead deal with trauma through wound healing that focuses on rehabilitating the structural integrity of the partic12 | 2021-2022 | Broad Street Scientific
1.5 Hypothesis We hypothesize that the Akt pathway is similar to the BMP pathway in that it is essential in the initiation of regeneration through the interaction of epithelial tissue. When the Akt pathway is inhibited, blastema formation will decrease, as shown through a decreased average length and growth per hour. 2. Materials and Methods Planarian Maintenance Dugesia dorotocephala (Brown planaria) were bought from Carolina Biological and kept in non-sealed plastic containers. The cultures were maintained at room temperature and kept in a dark environment to accommodate their photonegative behavior [1]. The planarians were fed small pieces of Lumbriculus variegatus (Californian blackworm) weekly for 30 minutes to prevent excess waste in the water. Partial spring water changes occurred every three days or when the water became visibly dirty. Induction of Renegeration Planarians were pipetted from their cultures into a petri dish with 8mL of either spring water, Dorsomorphin dihydrochloride, Dimethyl sulfoxide (DMSO), Nutlin-3, or MK-2206 2HCl. The sample size for each treatment was ten planarians. The organisms were treated in their selected solution for 24 hours. After the treatment time, the petri dishes with planarians were placed under a dissecting scope. Then, a scalpel was used to apply an anterior-posterior cut to each planarian to cut them approximately in half. The two planarian segments were maintained in the same dish, and the petri dishes were stored in a dark container until they were taken out for imaging (Figure 3).
BIOLOGY
Figure 3. Treatment time and location in which the wound site was induced. Treatments All treatments were chosen based on the region they targeted, safety, and availability. The BMP inhibitor used was Dorsomorphin dihydrochloride, a BMP type I receptor inhibitor. A 10 mM stock solution of the BMP inhibitor was created from 10 mg of the drug with 2.12 mL of DMSO. The concentration used for the experiment was 2 μM. The selective Mdm2 antagonist used was Nutlin-3. A 10 mM stock solution of Nutlin-3 was created from 5 mg of the drug with 0.8598 mL of DMSO. The concentration used for the experiment was 10 μM. The Pan-Akt inhibitor used was MK-2206 2HCl, a highly selective inhibitor of Akt 1/2/3. A 10 mM stock solution of MK-2206 2HCl was created from 5 mg of the drug with 1.0408 mL of DMSO. The concentration used for the experiment was 10 μM. A 0.02% DMSO treatment was used. Measurement of blastema Pictures were taken of the planarians at 24-hour intervals for seven days (Figure 4). The petri dishes containing the specimens were placed under a dissecting microscope with a transparent ruler underneath. The ruler and planarian section were kept in focus during the pictures. Pictures were taken when the planarian was in full extension. These pictures were analyzed using ImageJ. The blastema was measured from the center of the wound site using the measurement tool and was differentiated from the rest of the organism’s body through its white color. If a planarian fissioned, the original segment was measured normally, and the fission was disregarded.
Figure 4. Pictures of the final stages of blastema growth for Spring Water, DMSO, BMP inhibitor, PanAkt Inhibitor, and Mdm2 inhibitor on day 7. Statistical Analysis Blastema lengths were averaged to find the average length per hour for both anterior and posterior segments for each treatment. The trendlines from the growth of each sample were averaged to find the average growth rate for a particular treatment. A one-way ANOVA test was used to determine the significance between treatments. Error bars were provided with graphs using plus and minus standard error of the mean (SEM). 3. Results Anterior Blastema Formation Growth Rate The anterior blastema growth rate was not significantly affected by the vehicle DMSO compared to spring water (Fig. 5). Therefore, DMSO was compared to each treatment to determine the significance of any changes in blastema growth. The positive control of BMP had a growth rate of 0.038 mm/hr, significantly less than the DMSO growth rate of 0.064 mm/hr. The Mdm2 treatment had an insignificant growth rate of 0.076 mm/hr, and Pan-Akt also had an insignificant growth rate of 0.050 mm/hr.
Figure 5. Average anterior blastema growth rate based on sample trendlines; error bars are +/- SEM. Growth was measured for Spring Water, DMSO, BMP inhibitor, Mdm2 inhibitor, and Pan-Akt inhibitor for a seven-day period. BIOLOGY
Broad Street Scientific | 2021-2022 | 13
Posterior Blastema Formation Growth Rate The posterior blastemas were investigated for changes in growth rate. The vehicle DMSO demonstrated a significant decrease in growth of 0.067 mm/hr compared to spring water’s 0.079 mm/hr (Fig. 6). Therefore, DMSO was compared to the rest of the treatments to highlight any significant changes in growth. BMP continued to show a significant decrease in growth rate of 0.040 mm/ hr. The Mdm2 treatment produced a significant increase in growth rate of 0.078 mm/hr. However, the Pan-Akt treatment showed a significant decrease in growth rate of 0.047 mm/hr.
Figure 6. Average posterior blastema growth rate based on sample trendlines; error bars are +/- SEM. Growth was measured for Spring Water, DMSO, BMP inhibitor, Mdm2 inhibitor, and Pan-Akt inhibitor for a seven-day period.
All treatments presented an upward trend in their growth over the seven-day period of the experiment (Figs. 7 and 8). The DMSO average posterior length was only significantly less than spring water on day 2 (the DMSO average length was 0.142 mm, the spring water length was 0.200 mm). To analyze the inhibitor treatments, each one was compared to DMSO for each day in the seven-day experiment. All inhibitor treatments were insignificant when compared to the DMSO treatment on the first day. On day 2, Mdm2 was significantly higher than DMSO. Mdm2 led to a blastema length of 0.204 mm, while DMSO led to a blastema length of 0.142 mm. On day 3, the BMP treatment produced a length significantly less than DMSO. The BMP treatment had a length of 0.161 mm, and DMSO had a length of 0.263 mm. On day 4, the Pan-Akt and BMP treatment were both significantly lower than the DMSO treatment. The Pan-Akt treatment had a length of 0.238 mm. The BMP treatment had a length of 0.200 mm. The DMSO treatment had a length of 0.375 mm. On day 5, this trend continued as both the Pan-Akt treatment and BMP treatment were significantly lower than the DMSO treatment. The Pan-Akt treatment had a length of 0.281 mm, and the BMP treatment had a length of 0.261 mm. The DMSO treatment had a length of 0.471. 14 | 2021-2022 | Broad Street Scientific
On day 6, both the Pan-Akt and BMP treatment were significantly lower than the length of the DMSO treatment. The Pan-Akt treatment’s length was 0.284 mm, and the BMP treatment’s length was 0.259 mm. The DMSO treatment’s length was 0.480 mm. On day 7, the BMP treatment was the only inhibitor to have a significantly lower length than DMSO. The BMP treatment had a length of 0.316 mm versus the DMSO length of 0.489 mm. Average Length for Anterior Blastema Formation
Figure 7. Average anterior blastema length per hour for the 5 treatments; error bars are +/- SEM. Growth was measured for Spring Water, DMSO, BMP inhibitor, Mdm2 inhibitor, and Pan-Akt inhibitor for a seven-day period. Double asterisks represent p-values less than .01, and single asterisks represent p-values less than .05 in comparison to DMSO. Average Length for Posterior Blastema Formation
Figure 8. Average posterior blastema length per hour for the 5 treatments; error bars are +/- SEM. Growth was measured for Spring Water, DMSO, BMP inhibitor, Mdm2 inhibitor, and Pan-Akt inhibitor for a seven-day period. Double asterisks represent p-values less than .01, and single asterisks represent p-values less than .05 in comparison to DMSO.
BIOLOGY
4. Discussion 4.1 Anterior versus Posterior The anterior and posterior segments of the planarians were measured separately because of the different anatomical structures within each region. The expectation was that a drug might have had a more substantial presence in a particular region than another, which was validated through the measurements. The measurements of blastema formation for each region had different values of significance and for different treatments because the organism region influenced the drug’s effect. 4.2 BMP as a Positive Control
effects that were expected. For the anterior segments, its growth rate was insignificant compared to DMSO and had a higher value. For the posterior segments, the growth rate was significantly higher than DMSO and insignificant compared to spring water. For the average length measurements, this treatment was insignificant compared to DMSO except for one day on the posterior segments when it was significantly higher. Thus, the Mdm2 protein did not affect blastema formation for the anterior segments and increased blastema formation for the posterior segments when inhibited. Since only one of the roles of this protein is to inhibit p53, it is reasonable for it to have affected other aspects of the planarians that could contrast its expected activation of p53 when Mdm2 is inhibited.
The BMP pathway’s inhibition consistently led to it being a positive control, thus validating its involvement in the initiation of tissue regeneration. In both the anterior and posterior segments, it was significantly lower than the vehicle of DMSO for the average blastema growth rate. The anterior average length per day had the trend of smaller blastemas throughout most of the seven days and was significantly smaller than DMSO at two periods. The posterior average length per day also demonstrated the trend of having smaller blastemas and was significantly smaller than DMSO on five days. For all measurements, this inhibitor had the trend of being the lowest value. These values validate the work of past literature in this experiment and allow it to be used as a benchmark for what the Akt pathway’s inhibition should replicate. Yet, the BMP inhibitor led to the deformation of the head region for the planarians and led to the death of two of the anterior segments and one posterior segment.
4.4 Limitations
4.3 Akt Inhibition
Signaling pathways are a critical aspect of the regeneration process because of their role in alerting neoblasts to form blastemas at the wound site. Specifying which pathways are involved allows for the opportunity to expedite the completion of regeneration. By measuring blastema formation, the BMP pathway has been shown to be the most critical aspect of this initiation so far in all segments of the planarian. The Akt pathway as a whole is involved in this process as well, although the Akt pathway has only been shown to significantly influence blastema formation for the posterior region of the planarian. On the other hand, the Mdm2 pathway does not have a role in this process. To continue exploring the influence of these pathways, it would be beneficial to experiment with their amplification. This experimentation would allow for the increase of blastema formation and a method of control to be investigated. Additionally, identifying the specific portions of the Akt pathway that are more critical in initiat-
Although the Pan-Akt inhibitor led to a lower trend in blastema formation for all measurements, it never was significant for the anterior portion. This outcome may have resulted from a weaker presence in anterior structures. However, the Pan-Akt inhibitor did lead to decreased blastema formation in the posterior segments. For instance, it had a significantly reduced growth rate when compared to DMSO. Its blastema length measurements had three days in which it was significantly lower than DMSO and had a consistent trend of lower blastema lengths after the initial three days. These results support the hypothesis that the Akt pathway is involved in the initiation of tissue regeneration, yet this experiment suggests that it is exclusive to the posterior region. The Mdm2 treatment was chosen because of its regulation on p53 and its position as a downstream protein of the Akt pathway. However, it presented the opposite BIOLOGY
A limiting factor to this experiment is the inability to anesthetize the planarians for imaging. Using an anesthetic to prevent movement would allow planarians to be positioned in the same format and extension to produce more consistent images. Instead, the planarians were pipetted to induce locomotion, which would cause them to extend fully. This technique could have caused unnecessary stress and could have led to some of the fissions. Another limitation was the stress that the planarians faced when placed under the dissecting scope. The planarians were photographed under the dissecting scope every 24 hours, which stressed them because of their photonegative behavior [1]. This factor is an additional variable that could have caused fissions. 5. Conclusion and Future Work
Broad Street Scientific | 2021-2022 | 15
ing regeneration is valuable. The Akt pathway is known to cause cancer in humans when amplified, and the goal of this study is to find a pathway that could be amplified in humans. Therefore, identifying a specific downstream portion of the pathway that is safe to amplify would help transfer this capability to humans. Another consideration is finding a drug that also works on the anterior segments of planarians. Discovering why the Akt pathway only worked on the posterior region could help identify a drug or pathway effective on all regions. Lastly, quantifying these pathways’ effect on other aspects of the planarians should be examined. The BMP pathway’s inhibition led to some planarians’ death and the deformation of their eyespots and head. Quantifying this damage could aid in the investigation of the Akt pathway as a safe alternative. Furthermore, utilizing methods that identify where the pathway is activated would help determine specific regions for amplification. With this information and further experimentation, post-injury regeneration in humans is a closer possibility. Amplifying the pathways involved in the initiation of regeneration would lead to quicker recovery times for patients with time-sensitive injuries and help people ease back into society more comfortably after severe injuries.
Cells. Developmental Biology, 233(1), 109–121. https://doi. org/10.1006/dbio.2001.0226
6. Acknowledgements I would like to thank Dr. Kimberly Monahan for her support and guidance throughout this research project. I would also like to thank the Research in Biology class of 2022 for their encouragement. Thank you to Dr. Heather Mallory, Benet Ge, and Lixin Yang for their assistance during the summer. I would like to thank the North Carolina School of Science and Mathematics, Burroughs Wellcome Fund (BWF), NCSSM Foundation, and Glaxo Foundation for making this opportunity possible.
9. Peiris, T. H., Ramirez, D., Barghouth, P. G., & Oviedo, N. J. (2016). The akt signaling pathway is required for tissue maintenance and regeneration in planarians. BMC Developmental Biology, 16 doi:http://dx.doi.org/10.1186/ s12861-016-0107-z
7. References 1. Cao, Z., Liu, H., Zhao, B., Pang, Q., & Zhang, X. (2020). Extreme Environmental Stress-Induced Biological Responses in the Planarian. BioMed Research International, 2020, 7164230. https://doi.org/10.1155/2020/7164230
11. Wagner, D. E., Wang, I. E., & Reddien, P. W. (2011). Clonogenic neoblasts are pluripotent adult stem cells that underlie planarian regeneration. Science (New York, N.Y.), 332(6031), 811–816. https://doi.org/10.1126/science.1203983
5. Liao, Y., & Hung, M.-C. (2010). Physiological regulation of Akt activity and stability. American Journal of Translational Research, 2(1), 19–42. 6. Molina, M. D., Saló, E., & Cebrià, F. (2007). The BMP pathway is essential for re-specification and maintenance of the dorsoventral axis in regenerating and intact planarians. Developmental Biology, 311(1), 79–94. https:// doi.org/10.1016/j.ydbio.2007.08.019 7. Ogawa, K., Ishihara, S., Saito, Y., Mineta, K., Nakazawa, M., Ikeo, K., Gojobori, T., Watanabe, K., & Agata, K. (2002). Induction of a noggin-Like Gene by Ectopic DV Interaction during Planarian Regeneration. Developmental Biology, 250(1), 59–70. https://doi.org/10.1006/ dbio.2002.0790 8. Pearson, B. J., & Alvarado, A. S. (2010). A planarian p53 homolog regulates proliferation and self-renewal in adult stem cell lineages. Development, 137(2), 213–221. https:// doi.org/10.1242/dev.044297
10. Richardson, W., Clarke, S., Quinn, T., & Holmes, J. (2015). Physiological Implications of Myocardial Scar Structure. Comprehensive Physiology, 5(4), 1877–1909. https://doi.org/10.1002/cphy.c140067
2. Freedman, D. A., Wu, L., & Levine, A. J. (1999). Functions of the MDM2 oncoprotein. Cellular and Molecular Life Sciences: CMLS, 55(1), 96–107. https://doi.org/10.1007/ s000180050273 3. Kato, K., Orii, H., Watanabe, K., & Agata, K. (1999). The role of dorsoventral interaction in the onset of planarian regeneration. Development, 126(5), 1031–1040. 4. Kato, K., Orii, H., Watanabe, K., & Agata, K. (2001). Dorsal and Ventral Positional Cues Required for the Onset of Planarian Regeneration May Reside in Differentiated 16 | 2021-2022 | Broad Street Scientific
BIOLOGY
CHEMINFORMATICS ANALYSIS OF THE N-METHYL-DASPARTATE (NMDA) RECEPTOR NEUROTOXICITY Samanyu Dixit Abstract Neurotoxicity refers to the disruption and damage that xenobiotic chemicals may cause to neurons. As new compounds are developed, it is essential to understand the potential neurotoxic effects these compounds may have. Quantitative structure-activity relationship (QSAR) models are statistical/machine learning methods that can learn from previous data and predict the toxicological effects of compounds lacking experimental data. The main goal of this study is to develop models of the N-methyl-D-aspartate (NMDA) receptor, an important protein responsible for multiple physiological functions, such as memory. When unintentionally overstimulated, this target can trigger neurotoxicity. In this study, we have collected, curated, and integrated the largest publicly available data on the NMDA receptor from the EPA and ChEMBL databases. This work was executed in KNIME, a data analytics tool that allows the user to create comprehensive workflows with chains of commands that manipulate the data. The result of this project can be conceptualized as a machine learning model that can be employed to predict the safety of new chemicals related to neurotoxicity in the NMDA receptor. The future steps of this project include developing QSAR models for multiple curated datasets to develop a comprehensive neurotoxicity platform. 1. Introduction 1.1 - Background Neurotoxicity is the damage that occurs to either the central or peripheral nervous system when the body is exposed to natural or man-made toxic substances (Fig. 1). Those toxins can affect neural communication by disrupting or even killing nerves that are essential for processing information throughout the body. In addition, neurons, which are fundamental units of the nervous system, have a high metabolic rate and therefore are at further risk of damage from toxins. Untested chemicals leached in the environment or industrial chemicals requiring special protection or restricted usage instructions are examples of substances that can exhibit neurotoxic behavior.
Figure 1. Depiction of the central and peripheral components of the human nervous system. Figure retrieved from [Queensland Brain Institute]. In order to understand the nature of drugs and other compounds that can potentially induce neurotoxic effects in the body, it is important to conduct risk assessment. Risk refers to the balance between the exposure and severity of a compound’s effects on people. If a compound is highly toxic but it is unlikely that a person will get exposed to it, then the risk is low. However, as the exposure increases, then the risk is higher for that same compound. BIOLOGY
Broad Street Scientific | 2021-2022 | 17
In most cases, compounds that are extremely toxic are used in a controlled manner, such as a chemistry laboratory, where the exposure and thus risk are low. Nonetheless, this is not the case for all compounds that could exhibit neurotoxic effects. Pesticides, for example, are not strictly regulated and can result in farmers and other people on farms coming into contact with these compounds. Subsequently, the risk from pesticides is higher as people can be directly exposed to toxins or indirectly through produce. Toxicity tests can be carried out to estimate the toxicity of the compound (Bal-Price et al., 2018; Linne, 2018). The results from those tests can then be used to conduct a risk assessment. For example, a compound that is both catastrophic and frequent poses a much higher risk while the same catastrophic compound that is extremely improbable only poses a mild risk. Typically, the compound that is catastrophic can not be a drug as it will not be approved by the FDA, but in other situations it is important to be aware of how likely a factory worker may be exposed to a chemical or what types of household products contain dangerous chemicals. In these instances, there are varied levels of exposure to a possible neurotoxic substance thus reinforcing the significance of exposure in a risk assessment (Assessment, 1998). The N-methyl-D-aspartate (NMDA) receptor is one of the three ionotropic glutamate receptors; the other two are AMPA and kainate (Fig. 2).
resulting in neurological disorders (Zambre et al., 2019). 1.2 - QSAR Modeling Quantitative structure activity relationship (QSAR) modeling is a powerful technique that uses statistical and machine learning methods to generate models that can predict biological properties of different chemical compounds based on the molecular structure. QSAR modeling helps to explain the trend between molecular substructures and biological activity. The molecular substructures are represented as descriptors, a computed number that describes the chemical, such as lipophilicity, hydrogen bond donors/acceptors, and topological polar surface area. QSAR has served as the foundation for experimental medicinal chemistry due to the similarity approach that is implied by this statistical technique: similar molecular structures will result in similar biological activities (Tropsha, 2010). Historically, QSAR was only applied to the field of computer aided drug design (CADD). This field has evolved and grown which has resulted in the application of QSAR to the prediction of pharmacokinetic properties and toxicological behavior in addition to its usage for relaying molecules with corresponding biological activity (Muratov et al., 2020). Similar to other in silico research methods, QSAR is a more cost-effective way of being able to predict toxicity profiles of chemicals. This approach is more ethical as well, since in silico methods prevent the need of large amounts of laboratory tests on animals in order to determine biological properties. Decreasing the amount of animal testing, cost, and greater efficiency are therefore reasons that make QSAR modeling a useful tool for prediction of neurotoxicity (Alves et al., 2018; Worth et al., 2011). 2. Materials and Methods 2.1 - Data Collection
Figure 2. Diagram of the NMDA binding site. Figure retrieved from [CureGRIN]. NMDA is important for the control of synaptic plasticity, which refers to the modification of synaptic transmission that occurs at pre-existing synapses. This process is thought to play an important role in the development of neural circuitry. Evidence suggests that disfigurement of the mechanisms that control synaptic plasticity can contribute to psychiatric disorders such as epilepsy, Parkinson’s, Alzheimer’s and Huntington’s. It is therefore reasonable to infer that neurotoxicity induced by the NMDA receptor can cause faulty synaptic plasticity mechanisms 18 | 2021-2022 | Broad Street Scientific
In order to determine a set of targets that can exhibit neurotoxic behavior when bound to different compounds (drugs, pesticides, cosmetics, etc.), we first analyzed datasets from the Integrated Chemical Environment (ICE), a database developed by the US Department of Health and Human Services’ National Toxicology Program (NTP). ICE contains datasets that were curated by the NTP Interagency Center for the Evaluation of Alternative Toxicological Methods (NICEATM) and the Interagency Coordinating Committee on the Validation of Alternative Methods (ICCVAM) and contain in vitro and in vivo test data, computational toxicity predictions and other in silico tools for chemical characterization and predicting toxicity. ICE contains a set of 92 assays that are associated BIOLOGY
with the transmission of signals through neurons. These assays are a combination of smaller subsets of assays that are related to different receptor and signaling pathways. Some of the neuronal transmission pathways include the dopamine receptor signaling pathway, histamine receptor activity and glutamate receptor signaling pathway, etc. After compiling a list of various targets associated with neural behavior, we conducted a literature review in order to identify whether the targets were associated with neurotoxic behavior, neuroprotective behavior, or neither. This would allow us to isolate the specific targets and/or their corresponding subunits to ensure that we are building predictive neurotoxicity models for targets that are known to exhibit this behavior when bound to certain compounds (U.S. Department of Health and Human Services). Our research on the glutamate NMDA receptor revealed that two receptor subunits were shown to exhibit neurotoxic behavior based on a previous study. In that study, the expression of the NMDA receptor subunits was examined in relation to NMDA-mediated neurotoxicity. Murine cerebral cortices were investigated using reverse transcription polymerase chain reaction and Western blotting. The authors concluded that the epsilon-2 and zeta-1 NMDA receptor subunits were primarily mediating glutamate neurotoxicity in the cultures as those proteins were clearly expressed both when they were exposed to glutamate and when they were not (Mizuta et al., 1998). Using this information, we were able to isolate the subunits of the target that exhibited neurotoxic behavior from those that were not shown to exhibit any neurological behavior in the literature. Once the specific targets were known, we utilized ChEMBL, a database that contains small molecules and their associated properties and bioactivites. By searching for different types of drug targets, ChEMBL can filter out compounds that interact or bind to those targets in any way. After doing a filtered search for specific targets of the NMDA pathway, we examined the Target Report Cards for NMDA receptor subunits epsilon-2 and zeta-1 in Homo sapiens (as ChEMBL also contains lists of compounds that target the same receptor subunits in other organisms like Rattus norvegicus and Mus musculus). Since better QSAR models are typically generated from larger sets of data, we chose to only focus on Ki and IC50 bioactivity types for both receptor subunits. However, this parameter was not satisfied for the zeta-1 subunit. The Ki dataset in ChEMBL only contained 23 compounds for zeta-1 and therefore we had to use a small dataset to generate the model. This may have contributed to the fact that some of the model’s numerical values were not predictive. Besides the Ki data for zeta-1 however, we had around 100 or more compounds to use for the other 3 datasets (Ki and IC50 for epsilon-2 and IC50 for zeta-1). BIOLOGY
In addition to ChEMBL data, we also looked for NMDA data on the EPA’s CompTox Chemicals Dashboard, a compilation of databases that prioritizes chemicals based on health risks. Since the EPA would be concerned with the effects of pesticides and drugs as they relate to humans and the environment, the data available through CompTox would also serve as a good predictive measurement for neurotoxicity for the various targets. By conducting an assay search of NMDA, we were able to find two datasets of approximately 1000 compounds: rGluNMDA_Agonist and rGLUNMDA_MK801_Agonist. We curated these datasets along with the ChEMBL data to generate our models. 2.2 - Data Curation After the datasets are collected, they must be cleaned and tidied up following specific protocols for chemical/ biological curation. In cheminformatics, curation involves five main steps. The first step is the removal of inorganics and mixtures. Since current cheminformatics tools are only able to calculate molecular descriptors for organic compounds, it is challenging to develop descriptors that can be used for inorganic compounds and mixtures, and hence those compounds are filtered out of the dataset. The next step of curation involves the structural conversion and cleaning of the dataset where SMILES strings are typically translated into 2D molecular graphs. Additionally, salt, charged organic molecules, and hydrogen atoms have to be evaluated for the dataset. Salts are often removed as they generate errors when molecular descriptors are calculated. Molecules need to be standardized, so charged organic molecules are neutralized. The presence of explicit hydrogen atoms in some cases can lead to higher predictiveness but in some cases can cause problems with the descriptor matrix, resulting in less predictive models. Normalizing specific chemotypes is the third step of the curation process. Oftentimes, the datasets contain the same functional group represented by a different structural pattern. Molecular descriptors calculated for the different types of representations of the same group yield different values, however, and hence this must be considered. The fourth step of the process is the removal of duplicates. The presence of duplicates is a problem for QSAR models because structural duplicates of a compound sourced from two different places or the observed frequency of chemotypes can skew the predictivity of the model. Lastly, a final manual check is generally performed to inspect each molecular structure. This process can be carried out in less time consuming ways by only examining compounds with complex structures or those with many atoms. Some errors that are often identified during this final curation step include an incorrect structure, incomplete bond normalization, or the presence of duplicates despite software being used to remove Broad Street Scientific | 2021-2022 | 19
them (Fourches et al., 2010; Fourches et al., 2016). In order to visualize and carry out the curation process in an organized and comprehensive manner, we utilized KNIME, a software that can be used to build workflows that consist of a sequence of nodes that manipulate the dataset in certain ways. The workflows that were created for the purpose of curating the datasets we had collected from ChEMBL and CompTox contained some basic organizational nodes that helped to filter or rename different columns and entries. In addition, they contained a series of metanodes, which itself is a combination of nodes, that carried out processes such as duplicate removal, cleaning of the molecules and general optimization of the dataset that best allows for the molecular descriptor to describe the data. The resulting output of these workflows were six curated datasets from which the models could be generated.
Table 1. Statistical characteristics for continuous QSAR models for NMDA.
2.3 - Model Building Similarly to curation, the model building process also involved the development of a workflow that would allow us to input our curated datasets and output numerical values that are indicative of model predictivity. The models were generated using a machine learning algorithm known as Random Forest. This system consists of thousands of decision trees that each contain a sequence of nodes that can be altered in each decision tree. For example, after you reach a certain node, there are two or more possible outcomes or successive nodes that you could go to next. This idea continues for multiple iterations until an entire decision tree is generated. Multiple decision trees are then generated, establishing a mathematical relationship between the chemical structure and the biological activity. An analysis of these values can then be completed in order to determine the general predictivity of that particular QSAR model (Cherkasov et al., 2014; Dearden et al., 2009). 3. Results
A similar principle applies to the binary models. We are looking for CCR, sensitivity, PPV, specificity and NPV values as close to 1.0 as possible but at the very least greater than 0.6. As we can see in Table 2, for the epsilon-2 receptor subunit IC50 dataset, the model is most predictive as the majority of the values are significantly greater than 0.5. For the remaining ChEMBL datasets, there are values that are less than 0.5, deeming these models ineffective at predicting neurotoxicity in the epsilon-2 Ki, zeta-1 IC50 and zeta-1 Ki datasets (Table 2). The EPA dataset did not reveal conclusive results when it was run through the workflow and therefore we cannot comment on the model’s ability to predict neurotoxicity in rGluNMDA_Agonist and rGluNMDA_MK801_Agonist.
3.1 - QSAR Modeling The data did not allow us to build predictive models. In order to be predictive, the Q2, i.e., the coefficient of determination, should be above 0.5, preferably 0.6, with a lower error. The results from the continuous models indicate that they are not predictive for neurotoxicity in the epsilon-2 and zeta-1 receptors as the Q2 values in both the training and test sets are less than 0.5 (Table 1).
20 | 2021-2022 | Broad Street Scientific
BIOLOGY
Table 2. Statistical characteristics of binary QSAR models for NMDA.
were to bind to a target in those pathways. 6. Acknowledgements The author thanks Dr. Alexander Tropsha and Dr. Vinicius Alves at UNC Chapel Hill, as well as Mr. Robert Gotwals at the North Carolina School of Science and Mathematics, for their mentorship, guidance, and support on this project. 7. References Alves, V. M., Borba, J., Capuzzi, S. J., Muratov, E., Andradi, C. H., Rusyn, I., & Tropsha, A. (2018). Oy Vey! A comment on “Machine Learning of Toxicological Big Data Enables Read-Across Structure Activity Relationships (RASAR) Outperforming Animal Test Reproducibility”. Toxicological Sciences.
4. Discussion The results of the study indicate that our models can accurately predict the neurotoxicity a compound will induce if it binds to the NMDA epsilon-2 receptor unit with the most certainty. It also indicates that we can predict, with some certainty, the neurotoxicity of a compound if it binds the NMDA zeta-1 receptor. With additional refinement, these models can be applied as an alternative to animal testing when screening compounds lacking experimental activity. Additionally, the models are expected to bring safer chemicals faster to the market and will generate less synthetic residue. Compounds that are harmful to humans will be screened out with the usage of our web tool that can predict the neurotoxicity that any given compound will have if it were to bind with the targets we investigated in the study. 5. Conclusion In this study, we developed binary and continuous quantitative structure activity relationship models employing machine learning algorithms for the NMDA receptor, an important target associated with neurotoxicity. Here we collected, curated, and integrated data from ChEMBL and EPA. Continuous models for NMDA failed, but we succeeded to develop binary models (to predict toxic vs. non-toxic) for two units of the NMDA receptor, namely the epsilon-2 and zeta-1 receptor. In the future, we hope to expand on this study by generating QSAR models for other targets such as those in the histamine receptor pathway, adrenergic receptor pathway and dopamine signaling pathway. The end goal is to then integrate all of these models into a comprehensive web tool that could predict the neurotoxicity of any compound if it BIOLOGY
Assessment, N. R. (1998). Guidelines for Neurotoxicity Risk Assessment. Bal-Price, A., Pistollato, F., Sachana, M., Bopp, S. K., Munn, S., & Worth, A. (2018). Strategies to improve the regulatory assessment of developmental neurotoxicity (DNT) using in vitro methods. Toxicology and applied pharmacology, 354, 7-18. Cherkasov, A., Muratov, E. N. et al. (2014). QSAR modeling: where have you been? Where are you going to?. Journal of medicinal chemistry, 57(12), 4977-5010. CureGRIN. https://curegrin.org/understanding-nmda-receptors-a-guide-for-grin-disorder-families/ Dearden, J. C., Cronin, M. T., & Kaiser, K. L. (2009). How not to develop a quantitative structure–activity or structure–property relationship (QSAR/QSPR). SAR and QSAR in Environmental Research, 20(3-4), 241-266. Fourches, D., Muratov, E., & Tropsha, A. (2010). Trust, but verify: on the importance of chemical structure curation in cheminformatics and QSAR modeling research. Journal of chemical information and modeling, 50(7), 1189. Fourches, D., Muratov, E., & Tropsha, A. (2016). Trust, but verify II: a practical guide to chemogenomics data curation. Journal of chemical information and modeling, 56(7), 1243-1252. Linne, M. L. (2018). Neuroinformatics and computational modeling as complementary tools for neurotoxicology studies. Basic & clinical pharmacology & toxicology, 123, 56-61. Broad Street Scientific | 2021-2022 | 21
Mizuta I, Katayama M, Watanabe M, Mishina M, Ishii K. Developmental expression of NMDA receptor subunits and the emergence of glutamate neurotoxicity in primary cultures of murine cerebral cortical neurons. Cell Mol Life Sci. 1998 Jul;54(7):721-5. doi: 10.1007/s000180050199. PMID: 9711238. Muratov, E. N., Bajorath, J. et al.(2020). QSAR without borders. Chemical Society Reviews, 49(11), 3525-3564. Queensland Brain Institute. https://qbi.uq.edu.au/brain/ brain-anatomy/central-nervous-system-brain-and-spinal-cord Tropsha, A. (2010). Best practices for QSAR model development, validation, and exploitation. Molecular informatics, 29(6‐7), 476-488. U.S. Department of Health and Human Services. (n.d.). National Institute of Environmental Health Sciences. Retrieved December 9, 2021, from https://ice.ntp.niehs.nih. gov/Search. Worth, A., Fuart‐Gatnik, M., Lapenna, S., & Serafimova, R. (2011). Applicability of QSAR analysis in the evaluation of developmental and neurotoxicity effects for the assessment of the toxicological relevance of metabolites and degradates of pesticide active substances for dietary risk assessment. EFSA Supporting Publications, 8(6), 169E. Zambre, V. P., Patil, R. B., Sangshetti, J. N., & Sawant, S. D. (2019). Comprehensive QSAR studies reveal structural insights into the NR2B subtype selective benzazepine derivatives as N-Methyl-d-Aspartate receptor antagonists. Journal of Molecular Structure, 1197, 617-627.
22 | 2021-2022 | Broad Street Scientific
BIOLOGY
INVESTIGATION ON JOINT USAGE OF GSK1016790A AND OPC-31260 ON RESTORING MATING ABILITY AND RAISING INTRACELLULAR CALCIUM LEVELS IN LOV-1 AND PKD-2 SINGLE AND DOUBLE MUTANT CAENORHABDITIS ELEGANS Gracie Lin Abstract Polycystic Kidney Disease (PKD) is a disease that is characterized by the accumulation of cysts on the kidneys, which progressively increase in size due to proliferation. Approximately 90% of all PKD cases are Autosomal Dominant Polycystic Kidney Disease (ADPKD). Multiple treatments are undergoing clinical trials, and GSK1016790A and OPC31260 have emerged as effective treatments. However, either drug alone has limitations in efficacy and off-target effects. In order to study these drugs further, the model organism C. elegans can be employed to study ADPKD due to mating behaviors in C. elegans that are dependent on the same genes associated with ADPKD. Utilizing this model, the efficacy of a combination of GSK1016790A and OPC-31260 was quantified by crossing male lov-1 and pkd-2 single and double mutant C. elegans with hermaphrodite C. elegans of the same mutation and treating them with a 2nM GSK1016790A concentration and 0.05% OPC-31260 concentration. If the treatments worked, mating would be restored in these crosses and males would be observed in the F2 population. To support the role of intracellular calcium in the restoration of mating behaviors, worms were also stained using fura-2/AM. A combination of GSK1016790A/OPC-31260 and GSK1016790A individually were not found to be more effective than each treatment individually in restoring mating ability in lov-1, pkd-2 and lov-1/pkd-2 mutants. However, the joint treatment and single GSK1016790A did demonstrate a statistically significant increase in Ca2+ fluorescence levels when compared to each treatment individually. Thus, the combination of GSK1016790A and OPC-31260 effectively increased intracellular Ca2+ levels, but did not increase mating ability, indicating that GSK1016790A/OPC-31260 likely has an inexplicable off-target effect. Further trials investigating the combinatorial effects of GSK1016790A and OPC-31260 and its specific effect on ADPKD cells should be conducted.
1 Introduction 1.1 Polycystic Kidney Disease Polycystic Kidney Disease (PKD) is the primary disease responsible for chronic kidney disease and chronic renal failure. PKD is characterized by the accumulation of cysts on the renal parenchyma, which progressively increases in size due to proliferation [8]. Over time, the cysts compress the surrounding nephrons, leading to a decline in renal function [8]. Symptoms of PKD typically include hypertension and growth of cysts on the surrounding organs. There are two main types, differing in inheritance: Autosomal Dominant Polycystic Kidney Disease (ADPKD) and Autosomal Recessive Polycystic Kidney Disease (ARPKD). ADPKD is caused by mutations in the PKD1 and PKD-2 genes, primarily affecting adults around 30-60 years of age and consisting of 90% of all PKD cases [2]. Although the disease is present at birth, it gradually presents symptoms, such as hypertension and enlarged kidneys, as cyst proliferation progresses. BIOLOGY
ARPKD is observed in infants and children and is caused by a mutation in the PKHD1 gene. The genes mutated in ADPKD encode proteins polycystin-1 (PC1) and polycystin-2 (PC2). PC1 functions as a large integral membrane receptor, while PC2 functions as a non-selective cation channel in the TRP family which transports calcium [3]. Together PC1 and PC2 regulate calcium channel activity in the cilia and other parts of the cell, namely intracellular calcium release in the endoplasmic reticulum [2]. Loss of PC1 and PC2 function leads to dysregulation of calcium signaling. Disruptions in calcium signaling have been linked to a cascade of pathways eventually resulting in activation of pathways associated with abnormal cell proliferation, therefore leading to the development and growth of the kidney cysts [6]. ADPKD is a multisystem disease with its influences on varying signaling pathways not well understood [13]. Therefore, current treatments aim to delay the growth of cysts or manage symptoms, while novel Broad Street Scientific | 2021-2022 | 23
therapies targeting genetic mechanisms and signaling pathways are under development. Disruption in calcium signaling caused by PC1 and PC2 deficiency has been identified as the potential cause behind rapid cell growth observed in ADPKD. This is supported by evidence demonstrating that calcium restriction in ADPKD cells causes cAMP-dependent B-Raf/ERK pathway activation, leading to increased cell growth [16]. Furthermore, Yamaguchi et al. demonstrated that controlled calcium addition, in the form of a calcium channel activator and a calcium ionophore, raised intracellular calcium levels and restored normal growth. Thus, intracellular calcium has been identified as a highly viable target for developing therapies [16]. 1.2 GSK1016790A and OPC-31260 Having identified calcium as a promising target, several drugs influencing intracellular calcium levels in ADPKD models are in development. At the forefront of these treatments are GSK1016790A and OPC-31260, two novel treatments currently undergoing research. GSK1016790A is a highly potent TRPV4 agonist that selectively activates TRPV4 and induces calcium influx, as Figure 1 shows. Tomilin et al. demonstrated the central role TRPV4 plays in calcium homeostasis by administering a 4 µM concentration of HC-067047, a selective TRPV4 antagonist, to NHK cells [12]. Decreased calcium levels and reduced flow-induced responses were observed after prolonged acute inhibition of TRPV4, comparable to levels of calcium observed in ADPKD cells [12]. The drastic reduction of calcium in TRPV4 blocked NHK cells demonstrates the integral role TRPV4 maintains in controlling calcium homeostasis.
Figure 1. Schematic demonstrating the impact GSK1016790A has on TRPV4 channels in ADPKD cells. Without addition of GSK1016790A, TRPV4 channels remain closed and Ca2+ is unable to pass through. Addition of GSK1016790A opens up the TRPV4 and Ca2+ can enter the cell OPC-31260 is a Vasopressin V2 Antagonist and functions to lower renal accumulation of cyclic 24 | 2021-2022 | Broad Street Scientific
adenosine monophosphate (cAMP). Figure 2 demonstrates how OPC-31260 lowers cAMP levels through inhibition. In renal tubular epithelial cells, intracellular calcium functions to limit cAMP accumulation and cell proliferation. However, in cases of calcium deprivation environments, such as PKD, it causes renal accumulation of cAMP, stimulating cell proliferation. Furthermore, cAMP activates the cAMP-dependent B-Raf/ERK pathway, leading to cell growth. Gattone et al. demonstrated that OPC-31260 administration reduced renal accumulation of cAMP and inhibited disease progression, supported by lower kidney weights, renal cyst volumes, mitotic indices, and systolic blood pressures [14].
Figure 2. Schematic demonstrating OPC-31260 acting as a vasopressin V2 receptor antagonist, binding to the V2 receptor, preventing renal accumulation of cAMP, and lowering levels of cAMP in renal epithelial cells 1.3 C. elegans Model PKD-1 and PKD-2 homologs in Caenorhabditis elegans (C. elegans) are lov-1 and pkd-2, genes expressed in male-specific neurons [15]. The homologs of PC1 and PC2, LOV-1 and PKD-2 (LOV-2) are necessary for efficient male mating, and affect mating behaviors in three aspects: sex drive (the willingness to leave food to mate), response (typical circling behavior after locating a potential mate), and location of the vulva [15]. The vulva is a hermaphrodite-specific ectodermal organ required for mating, as the males fertilize hermaphrodites through the vulva. [10] Those three behaviors have been found to be defective in lov-1 and pkd-2 deficient worms. Utilizing lov-1 and pkd-2 dependent mating behaviors presents us with a robust and tractable system for understanding polycystic kidney disease and testing the efficacy of varying treatments [15]. The mating behaviors of C. elegans are simple to BIOLOGY
observe, and effects on mating can be easily visualized through specific crossing [15]. Based on this research, the C. elegans model could be used to investigate the combinatorial effects of GSK1016790A and OPC-31260 on raising intracellular calcium levels. 1.4 Applications and Significance ADPKD is the most frequent form of PKD, and is the most common hereditary renal disease [9]. Prevalence worldwide is estimated to be about 1:400 to 1:1000, however, a significant fraction of those suffering from ADPKD are undiagnosed, making incidence difficult to assess [4]. The high volume of those with ADPKD demands the development of effective treatments, but therapies targeting the genetic mechanisms of ADPKD have inherent limitations [4]. Therefore, current therapies focus on slowing disease progression or managing symptoms. The development of a new treatment to better reverse the widespread effects of ADPKD on bodily systems is imperative to help millions of those suffering from ADPKD. 1.5 Hypothesis If lov-1 and pkd-2 mutant C. elegans are treated with a combination of GSK1016790A and OPC-31260, they will exhibit higher levels of intracellular calcium than lov1 and pkd-2 C. elegans treated with only GSK1016790A or OPC-31260 because the respective drugs have been shown to raise intracellular calcium levels individually. 2 Methods This study consisted of two preliminary experiments and two main experiments. The two preliminary experiments established the baseline calcium levels in untreated N2 worms and expected results from N2/ him-5, lov-1, pkd-2, and lov-1/pkd-2 crosses. Figure 3 demonstrates the structure of the crosses; the first cross is a male crossed with a hermaphrodite, and the second is a hermaphrodite progeny from the first cross left to self-fertilize. To encourage mating in the first cross, two hermaphrodites were placed in a spotted plate with 10-12 males. For the main experiment, the same preliminary crosses were treated with GSK1016790A, OPC-31260, or both drugs. The negative control was treating the N2/him-5 cross with the drugs.
Figure 3. Experimental design for observing the impacts of single and joint drug treatment on mating ability in N2 C. elegans. This experimental design was followed for every worm strain and drug treatment. Agar Plate and Spotted Plate Preparation Agar plates were used to maintain the C. elegans and to perform crosses. To prepare agar plates, in a 500 mL erlenmeyer flask, 1.5 grams of NaCl, 8.5 grams of agar and 1.25 grams of bactopeptone were added to 500 mL of distilled water. The mixture was swirled to dissolve the reagents, then autoclaved for 20 minutes for sterilization. The sterilized mixture was allowed to cool for 5 minutes, after which 500 μL of 1M CaCl2, 500 μL of 5mg/mL cholesterol in ethanol, 500 μL 1M MgSO4 and 12.5 mL of KPO4 buffer were added. The final mixture was then poured into 5.5 cm x 1.25 cm petri dishes, at around a height of 0.5-0.75 cm per dish. Spotted plates were prepared by spotting 85 μL into the center of each agar plate, and allowing it to completely dry before usage. C. elegans Diet The primary source of food for C. elegans is OP50, an Escherichia coli strain conventionally used as a bacterial food source for C. elegans. To prepare OP50, a LB broth was first created. In a 500 mL glass bottle, 2.5 grams of tryptone, 2.5 grams of NaCl, and 1.25 grams of yeast extract were added to 250 mL of distilled water. Slightly shaking the bottle, the reagents were dissolved in the distilled water. The LB broth was then autoclaved for 20 minutes to sterilize the solution. After autoclaving, an inoculation loop was used to take a single streak of an E. coli lawn, and was swirled in the LB broth to release the E. coli. Then, it was placed overnight with a loose lid in an incubator at 37 °C. The incubated OP50 was placed in a refrigerator to store. C. Elegans Strains and Maintenance All strains of C. elegans were purchased from the
BIOLOGY
Broad Street Scientific | 2021-2022 | 25
University of Minnesota Caenorhabditis Genetics Center. Upon arrival, the worms were chunked and placed onto agar plates containing 85 μL of OP50. The worms were maintained by chunking onto a fresh spotted plate every 4-7 days, or in the presence of contamination. Drug Concentration A 2nM concentration of GSK1016790A and a 0.05% concentration of OPC-31260 were used for the main experiment. To create a 2nM concentration, 6.55 mg of GSK1016790A was added to 1 mL of DMSO to make a 10nM stock concentration. From the stock concentration, 20 μL of 10 nM concentration of GSK1016790A was added to 80 μL of DMSO. For the 0.05% concentration of OPC-31260, a 0.1% OPC-31260 stock solution was created by adding 100 mg of OPC31260 to 1000 mL of DMSO. 50 μL of the 0.1% OPC31260 stock solution was added to 50 μL of DMSO to create a 0.05% concentration of OPC-31260. Drug Treatment For the single drug treatment, 50 μL of 2nM GSK1016790A or 0.05% OPC-31260 was placed on an agar plate and spread around the plate. After drying, 85 μL of OP50 was spotted into the center of the plate. For the combination drug treatment, 25 μL of 2nM GSK1016790A and 25μL of 0.05% OPC-31260 were placed and spread around the plate, left to dry and spotted with 85 μL of OP50. Ca2+ Staining and Fluorescence To stain Ca2+ levels, first a fura-2/AM stock solution was created. 1 mg of fura-2/AM was dissolved into 1 mL of DMSO. 50 μL samples were dispensed into Eppendorf Safe Lock tubes and stored at -15 degrees °C. One 50 μL sample was used per plate. Prior to adding Fura-2/AM, the specific strain of C. elegans to be stained was added to an agar plate containing 85 μL of OP50 and 50 μL of the specific treatment. They were allowed to rest in these conditions for four days before being stained and imaged. To stain the worms, 25 μL of Kolliphor EL was diluted in 75 μL DMSO. One sample of fura-2/am was added to the Kolliphor dilution. The Kolliphor dilution was then diluted to 9 mL with M9 buffer. It was then bath sonicated at 60 sonics/minute for 15 minutes at 0 °C. The worms to be stained were picked off of their original plate and onto a fresh agar plate. After bath sonication, 1 mL of M9 buffer was added to the fresh plate of agar containing worms to wash them off. The 1 mL of M9 buffer containing worms was added to the 26 | 2021-2022 | Broad Street Scientific
bath sonicated solution and left to incubate for 6 hours at 10-12 rpm. After 6 hours, the solution was centrifuged two times at 2000 rpm for 2-3 minutes. Almost all the M9 buffer was removed, leaving approximately 100150 μL of M9 buffer and the pelleted worms. It was left to rest for 20 minutes at 20 °C before being fixed onto a microscope cover slip for imaging (Figure 4, Table 1).
Figure 4. Schematic depicting the simplified methodology of staining and imaging Ca2+ fluorescence in C. elegans Table 1. Images of Ca2+ fluorescence in each worm strain (N2, lov-1, pkd-2, lov-1/pkd-2) and drug treatment (No Treatment, 50 μL of DMSO, 50 μL of 2nM GSK1016790A, 50 μL of 0.05% OPC31260, and 25 μL of 2nM GSK1016790A + 25 μL of 0.05% OPC-31260)
Corrected Total Cell Fluorescence Value Calculations In order to compare Ca2+ fluorescence levels between C. elegans, the Corrected Total Cell Fluorescence (CTCF) value was calculated for each worm within a treatment group and averaged when performing statistical analysis. To obtain the CTCF value, the C. elegans of interest were selected by outlining the worm using the freehand selections tool. Using the Measure function in the Analyze group, integrated density and cell area values were obtained. Next, an area in the background that did not contain fluorescence was selected and analyzed for the mean grey value. The mean grey value of the background was taken three times to ensure accuracy, and the average was used in calculations. BIOLOGY
Using the CTCF value to compare fluorescence was utilized in order to mitigate differences in background color, allowing fluorescence values to be standardized. The equation below was used to calculate the CTCF value: CTCF = Integrated Density - (Mean of background fluorescence * Cell Area) Imaging and Microscopy A Fura-2/AM stain was administered to the C. elegans, and was imaged using a fluorescence microscope. To score F2 populations for males, a light dissecting microscope was utilized. Statistical Measurements The number of plates containing males in F2 C. elegans populations was recorded four days after F1 hermaphrodite self-fertilization was recorded. The scoring data among each treatment and strain pair was analyzed using a Fisher’s Exact Test at p < 0.05. The CTCF values were compared amongst varying treatments and strains using a One-Way ANOVA test, through which p-values were obtained. 3 Results Mating ability quantification To quantify the restoration of performance of mating ability, percentages of males scored for each C. elegans strain and treatment were collected. The parental cross was treated with 85 μL of either no treatment, DMSO, single drug or joint treatment, and consisted of 2 hermaphrodite C. elegans with 10 male C. elegans. After four days, 12 hermaphrodites from the F1 progeny were individually picked into plates containing no treatment, and allowed to self-fertilize for 4 days. Among the F2 progeny, plates were scored for males, and the percentage was determined by dividing the number of plates containing males over the total number of plates. Figure 5 represents the percentage of males present for each C. elegans strain and treatment.
BIOLOGY
Figure 5. Percentage of males scored in F2 C. elegans populations for different strains and treatments. Five different treatments to 4 strains of C. elegans (N2, lov-1, pkd-2, lov-1/pkd-2) were administered: No Treatment (85 μL OP50 for 4 days), DMSO (85 μL OP50 + 50 μL DMSO for 4 days), GSK1016790A (85 μL OP50 + 50 μL 2nM GSK1016790A for 4 days), OPC-31260 (85 μL OP50 + 50 μL 0.05% OPC-31260 for 4 days) and GSK1016790A/OPC-31260 (85 μL OP50 + 25 μL 2nM GSK1016790A + 25 μL 0.05% OPC-31260 for 4 days). For lov-1 C. elegans treated with 2nM GSK1016790A, 0.05% OPC-31260 and 2nM GSK1016790A/0.05% OPC31260, the percentage of males increased from 0% to 64%, 30% and 100%, respectively. pkd-2 and lov-1/ pkd-2 mutant C. elegans displayed similar increases, with the percentages going from 14% and 0% to 100%, 89%, 100% and 67%, 30%, 83% for 2nM GSK1016790A, 0.05% OPC-31260 and 2nM GSK1016790A/0.05% OPC31260 treatments, respectively. However, between single treatments vs joint treatment, the increase in percentage of males varied widely with the worm strain, although there is a general increase in the percentage of males.
Broad Street Scientific | 2021-2022 | 27
Table 2. p-Values for each pair of treatments (No Treatment, 50 μL of DMSO, 50 μL of 2nM GSK1016790A, 50 μL of 0.05% OPC-31260, and 25 μL of 2nM GSK1016790A + 25 μL of 0.05% OPC-31260) and C. elegans strains (N2, lov-1, pkd-2, lov-1/pkd-2) determined using a Fisher’s Exact Test. Cells highlighted in green indicate statistical significance (p < 0.05)
Table 2 shows the p-Values that were obtained through a Fisher’s Exact Test at p < 0.05. Statistically significant values were found to be scattered amongst the varying comparisons. Statistically significant pairs were clustered within the 2nM GSK1016790A vs No Treatment and 2nM GSK1016790A/0.05% OPC-31260 vs No Treatment groups, demonstrating p-values below 0.05 in 75% of the strains. This indicates that 2nM GSK1016790A and 2nM GSK1016790A/0.05% OPC-31260 treatments in lov-1, pkd-2 and lov-1/pkd-2 mutant C. elegans, respectively, are able to significantly raise the percentage of males present compared to worms of the same strain without treatments. On the other hand, pairs comparing either 2nM GSK1016790A or 0.05% OPC-31260 to the 2nM GSK1016790A/0.05% OPC-31260 joint treatment were all statistically insignificant, aside from 2nM GSK1016790A/0.05% OPC-31260 vs 0.05% OPC-31260 in lov-1 mutant C. elegans. The statistical insignificance indicates that the combination of 2nM GSK1016790A and 0.05% OPC31260 was unable to statistically significantly raise the percentage of males.
Figure 6. CTCF values for N2, lov-1, pkd-2, lov-1/pkd-2 mutant C. elegans across various treatments. Five different treatments were administered: No Treatment (85 μL OP50 for 4 days), DMSO (85 μL OP50 + 50 μL DMSO for 4 days), 2nM GSK1016790A (85 μL OP50 + 50 μL 2nM GSK1016790A for 4 days), 0.05% OPC-31260 (85 μL OP50 + 50 μL 0.05% OPC-31260 for 4 days) and 2nM GSK1016790A/0.05% OPC-31260 (85 μL OP50 + 25 μL 2nM GSK1016790A + 25 μL 0.05% OPC-31260 for 4 days). For the 2nM GSK1016790A vs No Treatment and 0.05% OPC-31260 vs No Treatment CTCF values, the fluorescence experiences little change across all four C. elegans strains. However, the 2nM GSK1016790A/0.05% OPC-131260 vs No Treatment, 2nM GSK1016790A/0.05% OPC-131260 vs 2nM GSK1016790A, and 2nM GSK1016790A/0.05% OPC131260 vs 0.05% OPC-31260 demonstrated consistent increases in fluorescence among all four C. elegans strains.
Corrected Total Cell Fluorescence Values As another method of demonstrating the potential efficacy of the joint treatment, Ca2+ fluorescence levels were taken separate from the crosses performed to assess mating ability. C. elegans were picked onto a plate containing 85 μL of either no treatment, DMSO, 2nM GSK1016790A, 0.05% OPC-31260 or both treatments, and left to rest for 4 days. After 4 days, the C. elegans were stained and analyzed for CTCF values, which are represented below in Figure 6.
28 | 2021-2022 | Broad Street Scientific
BIOLOGY
Table 3. p-Values for CTCF values for varying treatments and C. elegans strains obtained through an One-Way ANOVA test. Five different treatments were administered to N2, lov-1, pkd-2 and lov-1/pkd-2 mutant C. elegans: No Treatment (85 μL OP50 for 4 days), DMSO (85 μL OP50 + 50 μL DMSO for 4 days), GSK1016790A (85 μL OP50 + 50 μL 2nM GSK1016790A for 4 days), OPC31260 (85 μL OP50 + 50 μL 0.05% OPC-31260 for 4 days) and GSK1016790A/OPC-31260 (85 μL OP50 + 25 μL 2nM GSK1016790A + 25 μL 0.05% OPC-31260 for 4 days). Cells highlighted in green indicate statistical significance (p < 0.05)
Table 3 demonstrates the p-Values obtained using a One-Way ANOVA test, and like results in Table 2, statistical significance varies between C. elegans strains and treatments. There was no statistical significance between DMSO and No Treatment CTCF values for all strains, demonstrating that DMSO has no statistically significant influence on Ca2+ fluorescence levels in C. elegans. CTCF values for either single treatment compared to no treatment are statistically insignificant, aside from No Treatment vs 2nM GSK1016790A in pkd2 mutant C. elegans, indicating that 2nM GSK1016790A and 0.05% OPC-31260 were not able to increase Ca2+ fluorescence levels significantly. However, the CTCF value for nearly every strain treated with the joint treatment vs either single treatment and joint treatment vs no treatment was statistically significant. This indicates that the combination of 2nM GSK1016790A and 0.05% OPC-31260 was able to increase Ca2+ fluorescence levels more than either treatment alone or no treatment. 4 Discussion 4.1 GSK1016790A and OPC-31260’s Individual and Joint Effect on Mating Ability Table 2 establishes that 2nM GSK1016790A and 2nM GSK1016790A/0.05% OPC-31260 have a statistically significant effect on mating in lov-1, pkd-2, and lov1/pkd-2 mutant C. elegans, whereas OPC-31260 does not have a statistically significant impact. It can BIOLOGY
be concluded from this that GSK1016790A plays a significant role in restoring mating ability in lov-1, pkd2 and lov-1/pkd-2 mutant worms, and could lead to the conclusion that TRPV4 plays a role in mating ability and proper performance of mating behaviors. Liedtke et al. expressed mammalian TRPV4 in ASH sensory neurons in osm-9 mutant C. elegans and found avoidance responses to osmotic and mechanical stimuli restored [5]. In addition, it was found that TRPV4 is integrated into normal ASH sensory neurons, as TRPV4 function in ASH requires endogenous C. elegans avoidance genes [5]. The increase in percentage of males after treatment with 2nM GSK1016790A, a TRPV4 agonist, could implicate a secondary function TRPV4 plays within C. elegans, in addition to generating avoidance behaviors. OPC-31260 did not display consistent statistically significant results when compared to the percentage of males in no treatment, 2nM GSK1016790A, 2nM GSK1016790A/0.05% OPC-31260 treated crosses. This could indicate that OPC-31260 is not as favorable as GSK1016790A as a treatment for ADPKD, or that the vasopressin V2 receptor does not play a role in mating, thus not restoring mating ability to lov-1 or pkd-2 mutants. Furthermore, interactions between GSK1016790A and OPC-31260 may lead to a decrease in efficacy of the joint treatment, as percent of males between 2nM GSK1016790A/0.05% OPC-31260 treated mutant C. elegans and either 2nM GSK1016790A or 0.05% OPC-31260 individually were found to be statistically insignificant. 4.2 GSK1016790A and OPC-31260’s Individual and Joint Effect on Ca2+ Levels Table 3 describes the statistical significance between Ca2+ fluorescence levels for different treatments for N2, lov-1, pkd-2 and lov-1/pkd-2 mutant C. elegans. For lov-1 mutant C. elegans, there was no statistical significance between any of the treatments, which aligns with prior research which discovered that lov-1 lacks the Ca2+-binding EF-hand found in polycystin 2, the gene product of pkd-2 [1]. Because lov-1 does not interact with Ca2+, the lack of statistical significance between treatments and no treatment is understandable. However, there is statistical significance between 2nM GSK1016790A treatment and 2nM GSK1016790A/0.05% OPC-31260 joint treatment and No Treatment in pkd-2 mutant worms. These data demonstrate that pkd-2’s function involves Ca2+ while lov-1’s function does not, potentially defining the relationship between lov-1 and pkd-2. Furthermore, in specific comparisons, there are statistically significant increases in either percentage Broad Street Scientific | 2021-2022 | 29
of males or Ca2+ fluorescence levels, while in the other, the data are statistically insignificant. This is found in the 2nM GSK1016790A vs No Treatment comparison, where it was statistically significant in lov-1, pkd-2 and lov-1/pkd-2 mutant C. elegans. However, when looking at the Ca2+ fluorescence levels for the same comparisons and strains, the data are statistically insignificant, aside from pkd-2. Thus, there appears to be an effect on mating by GSK1016790A, but not on Ca2+ fluorescence levels, indicating that there could be an unknown off-target effect of GSK101670A that interferes with mating in another way. The same trend appears within the joint treatment vs no treatment, implying that the 2nM GSK1016790A/0.05% OPC31260 joint treatment also may have a different offtarget effect on mating. While this experiment had been conducted under the belief that C. elegans mating was a Ca2+ driven process, and it has been shown that Ca2+ plays a major role in C. elegans fertilization, Ca2+ may not have an effect on mating ability and behaviors. However, further repetitions and larger sample sizes would need to be performed in order to support that hypothesis. 5 Conclusion and Future Work Through the use of C. elegans as a model organism, ADPKD was modeled. The 2nM GSK1016790A/0.05% OPC-31260 joint treatment and 2nM GSK1016790A single treatment was found to effectively restore and improve mating ability in untreated lov-1, pkd-2, and lov-1/pkd-2 mutant C. elegans, but did not demonstrate statistically significant improvements compared to either treatment individually, in all three mutant strains. Between joint treatments vs single treatments and joint treatment vs no treatment comparisons, Ca2+ fluorescence levels were found to be statistically significant in N2 and lov-1/pkd-2 mutant worms. The discrepancy in results from both restoration mating ability and Ca2+ fluorescence levels demonstrates that the joint treatment of 2nM GSK1016790A/0.05% OPC-31260 likely has an off-target effect on mating ability. Furthermore, it could be that an increase of intracellular Ca2+ levels does not directly correlate with an increase in the percentage of males or mating ability. With these results, future work should be conducted to investigate the off-target effect that GSK1016790A and OPC-31260 jointly have on mating ability. C. elegans should be used as a model organism to explore this, as effects on mating are easily visualized. In addition, the potential secondary function TRPV4 plays in C. elegans should be further investigated. 30 | 2021-2022 | Broad Street Scientific
In the future, the Wnt signaling pathway should be investigated to uncover the mechanisms by which ADPKD functions, and to possibly discover a new signaling pathway to approach ADPKD treatments. 6 Acknowledgements I would like to thank Dr. Kimberly Monahan for the dedication and patience she demonstrated when guiding me throughout the research process. In addition, I thank Dr. Heather Mallory for assisting me in my research throughout the school year and for serving as my mentor during the Summer Research and Innovation Program. Thank you to the Research in Biology Program class of 2022 for supporting me inside and outside of research. Finally, I would like to thank the North Carolina School of Science and Mathematics, GlaxoSmithKline and the NCSSM Foundation for providing me the opportunity to perform research. 7 Bibliography [1] Barr, M. M., and Sternberg, P. W. (n.d.). A polycystic kidney-disease gene homologue required for male mating behaviour in C. elegans. Nature News. Retrieved April 28, 2021, from https://www.nature.com/ articles/43913. [2] Bergmann, C., Guay-Woodford, L. M., Harris, P. C., Horie, S., Peters, D. J. M., and Torres, V. E. (2018, December 6). Polycystic kidney disease. Nature reviews. Disease primers. Retrieved November 9, 2021, from https://www.ncbi.nlm.nih.gov/pmc/articles/ PMC6592047/. [3] Fedeles, S. V., Gallagher, A.-R., and Somlo, S. (2014, May). Polycystin-1: A master regulator of intersecting Cystic Pathways. Trends in molecular medicine. Retrieved November 9, 2021, from https://www.ncbi. nlm.nih.gov/pmc/articles/PMC4008641/. [4] Irazabal, M. V., and Torres, V. E. (2013, February). Experimental therapies and ongoing clinical trials to slow down progression of ADPKD. Current hypertension reviews. Retrieved April 28, 2021, from https://www.ncbi.nlm.nih.gov/pmc/articles/ PMC4067974/. [5] Liedtke, W., Tobin, D. M., Bargmann, C. I., and Friedman, J. M. (2003, November 25). Mammalian TRPV4 (VR-OAC) directs behavioral responses to osmotic and mechanical stimuli in Caenorhabditis elegans. PNAS. Retrieved November 10, 2021, from https://www.pnas.org/content/100/suppl_2/14531. BIOLOGY
[6] Mangolini, A., de Stephanis, L., and Aguiari, G. (2016, January 6). Role of calcium in polycystic kidney disease: From signaling to pathology. World journal of nephrology. Retrieved May 5, 2021, from https://www. ncbi.nlm.nih.gov/pmc/articles/PMC4707171/. [7] Patel V; Li L; Cobo-Stark P; Shao X; Somlo S; Lin F; Igarashi P; (n.d.). Acute kidney injury and aberrant planar cell polarity induce cyst formation in mice lacking renal cilia. Human molecular genetics. Retrieved November 9, 2021, from https://pubmed. ncbi.nlm.nih.gov/18263895/. [8] Patel, V., Chowdhury, R., and Igarashi, P. (2009, March). Advances in the pathogenesis and treatment of polycystic kidney disease. Current opinion in nephrology and hypertension. Retrieved April 28, 2021, from https://www.ncbi.nlm.nih.gov/pmc/articles/ PMC2820272/.
[14] Gattone, Vincent H., Xiaofang Wang, Peter C. Harris, and Vicente E. Torres. “Inhibition of Renal Cystic Disease Development and Progression by a Vasopressin V2 Receptor Antagonist.” Nature Medicine 9, no. 10 (October 2003): 1323–26. https://doi. org/10.1038/nm935. [15] Ward, C. J., and Sharma, M. (2015, December 21). Polycystic kidney disease: Lessons learned from caenorhabditis elegans mating behavior. Current biology : CB. Retrieved September 27, 2021, from https://www. ncbi.nlm.nih.gov/pmc/articles/PMC6077978/. [16] Yamaguchi T; Hempson SJ; Reif GA; Hedge AM; Wallace DP; (n.d.). Calcium restores a normal proliferation phenotype in human polycystic kidney disease epithelial cells. Journal of the American Society of Nephrology : JASN. Retrieved May 7, 2021, from https://pubmed.ncbi.nlm.nih.gov/16319189/.
[9] Polycystic kidney disease. National Kidney Foundation. (2018, September 14). Retrieved May 6, 2021, from https://www.kidney.org/atoz/content/ polycystic. [10] Schindler, Adam J, and D. R. Sherwood. “Morphogenesis of the C. elegans Vulva.” Wiley Interdisciplinary Reviews. Developmental Biology 2, no. 1 (2013): 75–95. https://doi.org/10.1002/wdev.87. [11] Singaravelu, G., and Singson, A. (2013, January). Calcium signaling surrounding fertilization in the nematode caenorhabditis elegans. Cell calcium. Retrieved October 27, 2021, from https://www.ncbi. nlm.nih.gov/pmc/articles/PMC3566351/. [12] Tomilin, V., Reif, G. A., Zaika, O., Wallace, D. P., and Pochynyuk, O. (2018, August). Deficient transient receptor potential vanilloid type 4 function contributes to compromised [ca2+]i homeostasis in human autosomal-dominant polycystic kidney disease cells. FASEB journal : official publication of the Federation of American Societies for Experimental Biology. Retrieved October 27, 2021, from https://www.ncbi.nlm.nih.gov/ pmc/articles/PMC6044056/. [13] Torres, V. E., and Harris, P. C. (2011, December 13). Polycystic kidney disease in 2011: Connecting the dots toward a polycystic kidney disease therapy. Nature reviews. Nephrology. Retrieved May 7, 2021, from https://www.ncbi.nlm.nih.gov/pmc/articles/ PMC4096714/. BIOLOGY
Broad Street Scientific | 2021-2022 | 31
SELECTIVE AND COST-EFFECTIVE DEPOLYMERIZATION OF LINEAR POLYETHYLENE VIA TANDEM DEHYDROGENATION AND METATHESIS DUAL CATALYST SYSTEM Meghana Chamarty Abstract Efficient methods to aid in the safe degradation of polyolefins are necessary in preventing ecological harm and environmental stress. To address this need, this research presents a cost-effective metathesis-dehydrogenation tandem catalyst system composed of relatively safe and abundant first-row transition metals that can selectively depolymerize polyolefins into smaller molecular hydrocarbons. This process will not only break down the polymers efficiently, but also create chemical precursors that can be used in the plastics economy to produce plastic consumer products. In the present work, a synthesis pathway for novel cobalt and nickel pincer ligand catalysts was devised and proven. The catalysts were synthesized, purified, and subsequently tested for their efficacy in olefin dehydrogenation as a part of a tandem catalyst system with Re2O7 (metathesis catalyst). Each of these potential catalysts was evaluated in the extent of polyethylene (PE) degradation, PE recovery, and the nature of products produced. Cobalt dehydrogenation pincer catalyst was found to be effective for the catalytic degradation of PE and is a promising alternative to the cost-intensive and toxic iridium pincer catalysts. Comparatively, the nickel catalyst had low performance, rendering it unlikely to be useful in successful PE dehydrogenation and ultimately de-polymerization. 1. Introduction Inexpensive to produce and comprising many single-use commodities, polyolefin waste is abundant, making up about 57% of the 380 million tons of plastic waste produced annually. [1] By 2050, it is estimated that plastic production will reach over 1.1 billion tons per year. [2] While the strength and versatility of polyolefin plastics are the reason why polyolefin products are so useful, the sp3 hybridized nature of the carbon-carbon (C-C) bond between the monomers in the polyolefins causes them to be highly resistant to biodegradation. [3] As a result, plastics can take hundreds of years to naturally decompose. Efficient methods to aid in the safe degradation of polyolefins are necessary to prevent further ecological harm and environmental stress. [4][5] Currently, the plastics recycling economy is a linear system (Fig 1, top). Each step in this economy is energy-intensive, produces waste, and can only be utilized in the production of lower-quality materials, making this process highly unsustainable. [6] A better approach is to convert the polymeric material back into starting products so that it can be reused as a raw material. This approach is referred to as a circular recycling process (Fig 1, bottom). About 14% of produced plastics get recycled, 40% of which end up in landfills and contaminate the environment. [7] An efficient method of recycling that can transform the polymer structure of plastic polyolefins into small monomer alkanes will allow for the plastics 32 | 2021-2022 | Broad Street Scientific
economy to be transformed into a circular recycling economy. This will decrease loss of energy and environmental harm and increase overall sustainability.
Figure 1: Linear plastics economy (top) versus circular plastics economy (bottom). Current methods of recycling polyolefins can be categorized into three broad groups: mechanical recycling, thermochemical recycling, and catalytic recycling (Figure 2).
CHEMISTRY
mild reaction conditions, and fine control of degradation products, showing distinct advantages over traditional pyrolysis processes. However, expensive and toxic metals such as iridium are the core of the catalysis method. Therefore, if utilized at larger scales, the proposed catalyst system would pose environmental and health risks. Figure 2: Overview of current recycling processes (mechanical [8], thermochemical, and catalytic [24]). Mechanical recycling is the current method of mass recycling; but extremely inefficient, with only 16% of plastic actually getting recycled. [9] Additionally, recycled plastics have limited uses (concrete, [10] lumber, [11] clothing [12]). These methods simply displace plastics, so this is not sustainable. Thermochemical processes require high amounts of energy and have low product selectivity. These processes, which include pyrolysis and thermal cracking, enable the depolymerization of polyolefin by C−C bonds in PE to produce lightweight liquid alkanes as fuel sources or to integrate into chemical refineries. [13] [14] [15] [16] Catalytic recycling of polyolefins, unlike mechanical and thermochemical recycling, can selectively and efficiently cleave the C-C bonds. This method has been studied for short-chain alkanes [17] [18] [19] and lignin [20], but has not been extensively explored for the depolymerization of high-weight polyolefins. Dufaud and Basset studied catalytic degradation of PE and PP into mid-tolow weight alkanes by a zirconium hydride supported on silica-alumina, finding moderate activity under mild conditions (190 °C). [21] Celik et al. studied catalytic hydrogenolysis using Pt nanoparticles supported on SrTiO3 nanocuboids for the depolymerization of PE at 170 psi H2 and 300 °C, obtaining high yields of liquid hydrocarbons over a 96 hour period. [22] However, the mechanism of catalysis and investigation of other supported noble metals for hydrogenolysis at even milder conditions is required for further analysis. Rorrer, Beckham, and Román-Leshkov demonstrate the selective depolymerization of polyethylene to liquid hydrocarbons under mild conditions (200 °C and 30 bar H2) using Ru nanoparticles supported on carbon, although investigation of catalyst stability over time at milder pressures is required. [23] Recently, Jia et al. reported a mild and efficient degradation of PE into liquid fuels and waxes using light alkanes, iridium-based pincer ligand dehydrogenation catalysts, and Re2O7/γ-Al2O3 metathesis catalysts. [24] Alkane metathesis is the process by which alkanes are covalently rearranged to give a new distribution of alkane products. [25] [26] [27] This research demonstrated high efficiency, CHEMISTRY
Our research study shows the tandem metathesis and dehydrogenation catalytic depolymerization of linear, high-density polyethylene (HDPE) to processable lightweight liquid and mid-weight waxy hydrocarbons using relatively safe, cheap, and abundant cobalt in a metallopincer ligand. This method of PE degradation is based on a tandem catalytic cross alkane metathesis process developed by Goldman et al. [28] and Huang et al. [29] , which involves one catalyst for alkane dehydrogenation and another catalyst for olefin metathesis. The catalyst of focus for this research study was the alkane dehydrogenation catalyst. First, a one-step synthesis pathway was developed. Then, this synthesis pathway was tested and proved. Finally, PE recovery and dehydrogenation activity were investigated as a function of conversion. 2. Methods 2.1: Dehydrogenation Pincer Catalyst Screening Pincer complexes have been studied as catalysts for many applications. They promote pathways, such as Heck and Suzuki coupling, [30] [31] [32] [33] allylic additions, [34] [35] [36] and reduction of CO2. [37] Pincer complexes are characterized by coordinate bonding (chelation) that binds tightly to three adjacent coplanar sites. [38] Jia et al. explored the application of a dual catalyst system containing a supported pincer iridium complex and Re2O7/ γ-Al2O3. In this work, we optimized their iridium pincer complex (Figure 3) by removing the t-Bu2PO group on the benzene, and substituting the iridium with firstrow transition metals (nickel and cobalt).
Figure 3: Iridium dehydrogenation pincer complex.
Broad Street Scientific | 2021-2022 | 33
2.2: One-Step Synthesis Pathway Development Pincer complex synthesis generally requires a standard two-step synthetic protocol that requires equivalents of base to remove the three equivalents of acid generated during ligand synthesis, followed by the metalation step (shown at the top of Figure 4). Requiring tedious precision and filtration, this process is time consuming and difficult. Vabre et al. developed a one-pot, one-step synthesis method for a class of nickel pincer catalysts. Utilizing resorcinol and chlorodiisopropylphosphine ClP(i-Pr)2 , they detail a practical methodology for the preparation of the pincer complex. In this work, we utilized their one-step synthesis pathway, substituting ClP(i-Pr)2 with di-tert-butylchlorophosphine to achieve the desired ligand. This synthesis method was also adapted to the cobalt pincer catalyst. Figure 5: Nickel catalyst synthesis flowchart Flash silica gel column chromatography with a chloroform solvent was used to purify the desired product. The elution fractions from the column chromatography were collected in a round-bottom flask and the desired fractions evaporated in a rotary evaporator. 2.4: Cobalt Dehydrogenation Pincer Catalyst Synthesis and Purification
Figure 4: Traditional pincer complex synthesis pathway (adapted from Jia et al. [24]) vs novel one-pot, one- step synthesis pathway. 2.3: Nickel Dehydrogenation Pincer Catalyst Synthesis and Purification Resorcinol (Fisher Scientific), nickel powder, acetonitrile, and di-tert-butylchlorophosphine (Sigma-Aldrich) were used in synthesizing the nickel pincer dehydrogenation catalyst (Figure 5). Acetonitrile was dried with grade 56 three- angstrom molecular sieves for 6 hours. In a glovebox under inert conditions, reactants were combined into a three-neck Schlenk flask, transferred to a silicon heat bath at 70 °C, and placed on a stir plate at 400 rpm for 24 hours. The temperature inside the flask was monitored to ensure a constant temperature. Constant pressure and an inert atmosphere were maintained using a nylon balloon.
34 | 2021-2022 | Broad Street Scientific
Resorcinol (Fisher Scientific), cobalt powder, acetonitrile, and di-tert-butylchlorophosphine (Sigma- Aldrich) were used in synthesizing the cobalt pincer dehydrogenation catalyst (Figure 6). Acetonitrile was dried with grade 56 three-angstrom molecular sieves for 6 hours. In a glovebox under inert conditions, reactants were combined into a three-neck Schlenk flask, transferred to a silicon heat bath at 70 °C, and placed on a stir plate at 400 rpm for 24 hours. The temperature inside the flask was monitored to ensure a constant temperature. Constant pressure and an inert atmosphere were maintained using a nylon balloon. Products were purified through crystallization.
CHEMISTRY
Figure 7: Nickel pincer catalyst (blue) vs Resorcinol (red) FT-IR.
Figure 6: Cobalt catalyst synthesis flowchart 2.5: Nickel Catalyst Testing 0.24 g of 3 mm nominal granule size High-Density Polyethylene (Sigma-Aldrich), 0.00386 g Nickel Pincer Catalyst, 0.0552 g Re2O7 Catalyst (Fisher Scientific), and 8 mL octane or decane were reacted in a 25 mL Teflon chamber in an autoclave and set in an oven at 175 °C for 4 days. After 4 days the contents were centrifuged, and the resulting pellet was decanted, vacuum dried for 48 hours and massed. 2.6: Cobalt Catalyst Testing 0.24 g of 3 mm nominal granule size High-Density Polyethylene (Sigma-Aldrich), 0.00386 g Cobalt pincer catalyst, 0.0552 g Re2O7 Catalyst (Fisher Scientific), and 8 mL decane or octane were reacted in a 25 mL Teflon chamber in an autoclave and set in an oven at 175 °C for 4 days. After 4 days the contents were centrifuged, and the resulting pellet was decanted, vacuum dried for 48 hours, and massed. 3. Results and Discussion 3.2: Nickel and Cobalt Catalyst Synthesis and Purification The contents of the nickel reaction flask were initially grey/black. However, upon immediate addition to the flask the nickel turned blue. After 5-10 minutes on the stir plate/heat bath, the mixture turned emerald green. At the conclusion of the 24-hour reaction, the mixture appeared condensed. After 18 hours of normal evaporation, 0.866 g of catalyst was collected and analyzed through FT-IR (Figure 7).
CHEMISTRY
Due to the absence of the OH- peak (3300-3000 1/cm), and an evident peak in the P-O-C (aryl) functional zone (1240–1190 1/cm) in the Nickel Catalyst, the synthesis was concluded to have succeeded with a 38% yield. The contents of the cobalt reaction flask were initially grey/black. However, upon immediate addition to the flask the cobalt turned blue. Contents of the flask were purified through crystallization. Purity was evaluated with thin-layer chromatography using a 10% Dichloromethane/ 90% Methanol solution (Figure 8). Contents were dried and stored in inert conditions.
Figure 8: Cobalt pincer catalyst (green) vs Resorcinol (red) FT-IR Due to the absence of the OH peak (3300-2500 1/cm), and an evident peak in the P-O-C (aryl) functional zone (1240–1190 1/cm), the synthesis was concluded to have succeeded.
Broad Street Scientific | 2021-2022 | 35
3.3: Nickel and Cobalt Dehydrogenation Catalyst Testing Table 1 shows the cross metathesis of HDPE (Mw 4000) with varying combinations of 4.2 μmol metallopincer dehydrogenation catalyst and 57 μmol Re2O7 metathesis catalyst. Table 1: Cross metathesis of HDPE (Mw 4000) with varying combinations of 4.2 μmol metallopincer dehydrogenation catalyst and 57 μmol Re2O7 metathesis catalyst.
Sample
Control # 1 (no catalysts) 0.0977
Control #2 (only metathesis catalyst) 0.0958
Nickel dehydrogenation catalyst 0.0369
Cobalt dehydrogenation catalyst 0.0278
PE recovery (g) % Recovery Normalized PE Loss N (Number of trials)
100%
98.06%
52.93%
28.45%
0%
0.48%
8.04%
27.02%
3
3
3
3
PE recovery (g) refers to the mass of PE remaining in the sample at the conclusion of testing. % Recovery is determined in comparison to control #1 (no catalysts). Normalized PE loss quantifies the loss of PE in the recovered sample. Because of the low Mw of the PE (Mw = 4000), the smaller hydrocarbons converted by the catalysts are either dissolved in the decane or in liquid state. Qualitatively, there was a visible color change in the dehydrogenation + metathesis catalyst samples that was not as present in the controls. Rotor vaporizing the liquid component of the cobalt catalyst sample, the decane boiled off, although there was still liquid left in the flask. This indicates that some of the PE degraded into smaller liquid-at-room temperature hydrocarbons (10- 16 carbon alkanes). This did not occur in nickel catalyst trials. Based on the results of these trials, there is considerable evidence for the successful degradation of PE by the cobalt pincer dehydrogenation catalyst. Converting 27% of the PE sample into light-to-mid weight hydrocarbons 36 | 2021-2022 | Broad Street Scientific
in 96 hours with 28% recovery, this catalyst shows promise for mass scale depolymerization of PE at mild temperatures and pressures. 4. Conclusion In this work, a novel cobalt pincer dehydrogenation catalyst was developed to depolymerize linear polyethylene (HDPE) via a tandem metathesis dehydrogenation catalyst system. In contrast to previous work, this catalyst is cost-effective, non-toxic, easy to synthesize, and functions under mild conditions. Converting 27% of the PE sample into light-to-mid weight hydrocarbons in 96 hours, with 28% recovery, this catalyst shows promise for mass scale depolymerization of PE at mild temperatures and pressures. This is an important step in developing an efficient method of recycling that can transform the polymer structure of plastic polyolefins into small monomer alkanes, resulting in decreased loss of energy, decreased environmental and humanitarian harm, and increased overall sustainability. Future directions include computationally developing and screening more variations of these cobalt pincer complexes, computationally determining the mechanistic model of the tandem dehydrogenation metathesis catalyst system, developing non-toxic metathesis catalysts, testing the catalysts for recyclability/reusability, and characterizing products with mass spectrometry to precisely quantify the product distribution, analyze selectivity, and determine how selectivity behaves over various reaction time durations. 5. Acknowledgements I would like to thank the Burroughs Wellcome Fund and the NCSSM Foundation for providing me with this research opportunity, and my research instructor, Dr. Tim Anglin, for his dedication to my continued academic growth. 6. References 1 Geyer, R.; Jambeck, J. R.; Law, K. L. Production, use, and fate of all plastics ever made. Sci. Adv. 2017, 3 (7), No. e1700782 2 Hong, M.; Chen, E. Y. X. Future Directions for Sustainable Polymers. Trends Chem. 2019, 1 (2), 148−151. 3 Williams, P.T. Hydrogen and Carbon Nanotubes from Pyrolysis-Catalysis of Waste Plastics: A Review. Waste Biomass Valor 12, 1–28 (2021). 4 Ormonde, E.; DeGuzman, M.; Yoneyama, M.; Loechner, U.; Zhu, X. Chemical Economics Handbook (CEH): Plastics Recycling; IHS Markit: 2020. CHEMISTRY
5 Andrady, A. L. Microplastics in the marine environment. Mar. Pollut. Bull. 2011, 62, 1596−1605.
Pretreated Ru/TI02 Catalysts. J. Phys. Chem. 1986, 90, 4877−4881.
6 Guneet Kaur, Kristiadi Uisan, Khai Lun Ong, Carol Sze Ki Lin, Recent Trends in Green and Sustainable Chemistry & Waste Valorisation: Rethinking Plastics in a circular economy, Current Opinion in Green and Sustainable Chemistry, Volume 9, 2018, Pages 30-39.
18 Bond, G. C.; Yide, X. Effect of Reduction and Oxidation on the Activity of Ruthenium/Titania Catalysts for nButane Hydrogenolysis. J. Chem. Soc., Chem. Commun. 1983, 1248−1249.
7 Sheldon R. A; Norton M. Green chemistry and the plastic pollution challenge: towards a circular economy, Journal of Green Chemistry, Issue 19, 2020.
19 Egawa, C.; Iwasawa, Y. Ethane hydrogenolysis on a Ru(1,1,10) surface. Surf. Sci. 1988, 198 (1), L329−L334
8 Schyns Zoe, Shaver Micheal. Mechanical Recycling of Packaging Plastics: A Review; 2020. 9 Ormonde, E.; DeGuzman, M.; Yoneyama, M.; Loechner, U.; Zhu, X. Chemical Economics Handbook (CEH): Plastics Recycling; IHS Markit: 2020. 10 Siddique Rafat, Khatib Jamal, Kaur Inderpreet, Use of recycled plastic in concrete: A Review, Waste Management, Volume 28, Issue 10, 2008, Pages 1835-1852. 11 Hag-Elsafi O, Elwell DJ, Glath G, Hiris M. Noise Barriers Using Recycled-Plastic Lumber. Transportation Research Record. 1999;1670(1):49-58. 12 Young C, Jirousek C, Ashdown S. Undesigned: A Study in Sustainable Design of Apparel Using Post-Consumer Recycled Clothing. Clothing and Textiles Research Journal. 2004;22(1-2):61-68. 13 Serrano, D. P.; Aguado, J.; Escola, J. M. Developing Advanced Catalysts for the Conversion of Polyolefinic Waste Plastics into Fuels and Chemicals. ACS Catal. 2012, 2 (9), 1924−1941. 14 Anuar Sharuddin, S. D.; Abnisa, F.; Wan Daud, W. M. A.; Aroua, M. K. A review on pyrolysis of plastic wastes. Energy Convers. Manage. 2016, 115, 308−326. 15 Kunwar, B.; Cheng, H. N.; Chandrashekaran, S. R.; Sharma, B. K. Plastics to fuel: a review. Renewable Sustainable Energy Rev. 2016, 54, 421−428. 16 J. A. Onwudili, N. Insura, P. T. Williams, Composition of products from the pyrolysis of polyethylene and polystyrene in a closed batch reactor: Effects of temperature and residence time. J. Anal. Appl. Pyrolysis 86, 293– 303 (2009). 17 Bond, G. C.; Rajaram, R. R.; Burch, R. Hydrogenolysis of Propane, n-Butane, and Isobutane over Variously CHEMISTRY
20 Xiao, L.P.; Wang, S.; Li, H.; Li, Z.; Shi, Z.-J.; Xiao, L.; Sun, R. C.; Fang, Y.; Song, G. Catalytic Hydrogenolysis of Lignins into Phenolic Compounds over Carbon Nanotube Supported Molybdenum Oxide. ACS Catal. 2017, 7 (11), 7535−7542. 21 Dufaud, V.; Basset, J.-M. Catalytic Hydrogenolysis at Low Temperature and Pressure of Polyethylene and Polypropylene to Diesels or Lower Alkanes by a Zirconium Hydride Supported on Silica-Alumina: A Step Toward Polyolefin Degradation by the Microscopic Reverse of Ziegler± Natta Polymerization. Angew. Chem., Int. Ed. 1998, 37 (6), 806−810 22 Celik, G.; Kennedy, R. M.; Hackler, R. A.; Ferrandon, M.; Tennakoon, A.; Patnaik, S.; LaPointe, A. M.; Ammal, S. C.; Heyden, A.; Perras, F. A.; Pruski, M.; Scott, S. L.; Poeppelmeier, K. R.; Sadow, A. D.; Delferro, M. Upcycling Single-Use Polyethylene into High Quality Liquid Products. ACS Cent. Sci. 2019, 5 (11), 1795−1803. 23 Julie E. Rorrer, Gregg Beckham, and Yuriy Román-Leshkov; Conversion of Polyolefin Waste to Liquid Alkanes with Ru-Based Catalysts under Mild Conditions; JACS Au 2021 1 (1), 8-12 24 X. Jia, C. Qin, T. Friedberger, Z. Guan, Z. Huang, Efficient and selective degradation of polyethylenes into liquid fuels and waxes under mild conditions. Sci. Adv. 2, e1501591 (2016). 25 V. Vidal, A. Théolier, J. Thivolle-Cazat, J.-M. Basset, Metathesis of alkanes catalyzed by silica-supported transition metal hydrides. Science 276, 99–102 (1997). 26 J. M. Basset, C. Copéret, L. Lefort, B. M. Maunders, O. Maury, E. Le Roux, G. Saggio, S. Soignier, D. Soulivong, G. J. Sunley, M. Taoufik, J. Thivolle-Cazat, Primary products and mechanistic considerations in alkane metathesis. J. Am. Chem. Soc. 127, 8604–8605 (2005).
Broad Street Scientific | 2021-2022 | 37
27 M. C. Haibach, S. Kundu, M. Brookhart, A. S. Goldman, Alkane metathesis by tandem alkane-dehydrogenation– olefin-metathesis catalysis and related chemistry. Acc. Chem. Res. 45, 947–958 (2012). 28 A. S. Goldman, A. H. Roy, Z. Huang, R. Ahuja, W. Schinski, M. Brookhart, Catalytic alkane metathesis by tandem alkane dehydrogenation-olefin metathesis. Science 312, 257–261 (2006) 29 Z. Huang, E. Rolfe, E. C. Carson, M. Brookhart, A. S. Goldman, S. H. El-Khalafy, A. H. R. MacArthur, Efficient heterogeneous dual catalyst systems for alkane metathesis. Adv. Synth. Catal. 352, 125–135 (2010). 30 D. Morales-Morales, C. Grause, K. Kasaoka, R. Redón, R. E. Cramer and C. M. Jensen, Inorg. Chim. Acta, 2000, 300, 958; 31 R. B. Bedford, S. M. Draper, P. N. Scully and S. L. Welch, New J. Chem., 2000, 24, 745; 32 F. Miyazaki, K. Yamaguchi and M. Shibasaki, Tetrahedron Lett., 1999, 40, 7379; 33 A. Naghipour, S. J. Sabounchei, D. Morales- Morales, D. Canseco-González and C. M. Jensen, Polyhedron, 2007, 26, 1445. 34 O. A. Wallner, V. J. Olsson, L. Eriksson and K. J. Szabò, Inorg. Chim. Acta, 2006, 359, 1767; 35 Z. Wang, M. R. Eberhard, C. M. Jensen, S. Matsukawa and Y. Yamamoto, J. Organomet. Chem., 2003, 681, 189; 36 O. A. Wallner and K. J. Szabò, Org. Lett., 2004, 359, 1829. 37 S. Chakraborty, J. Zhang, J. A. Krause and H. Guan, J. Am. Chem. Soc., 2010, 132, 8872. 38 David Morales-Morales, Craig Jensen (2007). The Chemistry of Pincer Compounds. Elsevier.
38 | 2021-2022 | Broad Street Scientific
CHEMISTRY
APPLYING QUANTUM COMPUTATIONAL METHODS TO ANALYZE THE REACTION PROGRESSION OF AN SN2 FUNCTIONAL GROUP TRANSFORMATION REACTION Sophie Scherer Abstract Nucleophilic substitution 2nd order kinetics (Sn2) reactions afford many useful applications in organic chemistry synthesis, from functional group transformations to stereochemical inversions. Organic chemistry typically focuses on the general Sn2 reaction process without analyzing the specific energetic and physical mechanisms driving the reaction. To develop a more comprehensive understanding of the quantum mechanics behind the Sn2 reaction of 2-bromobutane to butan-2-ol, computational quantum chemistry methods were applied to analyze the intramolecular physical and energetic processes influencing this reaction. This paper presents the findings of molecular orbital analyses, transition state calculations, and NMR spectrum predictions applied to the Sn2 reaction of 2-bromobutane to butan-2-ol. Key observations included electron flow from the hydroxide (nucleophile) HOMO to the bromobutane LUMO; a combination of nucleophilic interactions, hydrogen bonding, and energetics driving the reaction; and a predicted NMR spectrum visualization for the product butan-2-ol with four major peaks. This computational approach to analyzing Sn2 reactions can be applied to better understand both the Sn2 reaction of 2-bromobutane to butan-2-ol and other Sn2 reactions used in pharmaceutical and research settings to reduce the time and monetary costs of wet lab experiments. Introduction This paper focuses on the application of quantum computing tools to better understand the physical and energetic progression of a specific Sn2 reaction. Visualization of the physical transition state, energetic pathways, molecular orbital interactions, and product NMR spectra enables scientists to approach this reaction more effectively in wet labs. The following research predicts the flow of electrons during the reaction’s progression, the transition state structure, the transition state energy profile, and the NMR spectrum of products for the Sn2 reaction of 2-bromobutane to butan-2-ol. Nucleophilic substitution 2nd order kinetics reactions (more commonly referred to as Sn2 reactions) are important reactions in organic chemistry synthesis due to their versatility in manipulating reactants to form the desired target molecule(s). The foundation of an Sn2 reaction involves a nucleophile attacking an electrophilic substrate, resulting in a configuration inversion (5). The reaction mechanism consists of a transition state in which the bond of the leaving group is broken simultaneously with the formation of the new bond of the product (Fig. 1) (5).
Figure 1: General progression of an Sn2 reaction (5) As seen in Fig. 1, both the nucleophile and substrate are CHEMISTRY
contained within the transition state, indicating that the Sn2 reaction is a bimolecular reaction that is first order with respect to the substrate and nucleophile and second order overall, hence the name of nucleophilic substitution 2nd order kinetics (5). Sn2 reactions afford many useful applications in organic chemistry. Most simply, this reaction allows for the substitution of one group for another within a molecule (5). Depending on the nucleophile used in the reaction, an Sn2 reaction can also enable functional group transformations, such as from a halide to an alcohol (5). Sn2 reactions can also be used to invert the stereochemical properties of a chiral molecule (5). A carbon is said to be chiral when there are four unique groups bonded to the central carbon, enabling the carbon to have two mirror-image orientations of the four bonding groups. Each of the four groups are assigned descending priority orders following the Cahn Ingold Prelog priority rules to describe which of the two possible configurations is present at a specific chiral center (5). The chiral center is said to have an R stereochemistry if the priority groups increase moving clockwise around the chiral carbon when the lowest priority group is held at the back. Conversely, a chiral center has an S stereochemistry if the priority groups increase moving counterclockwise around the chiral carbon when the lowest priority group is held at the back. Due to the inversion caused by the progression of an Sn2 reaction, this reaction can be applied to invert the stereochemistry of a chiral carbon from R to S or vice versa. This paper will specifically focus on the Sn2 reaction of 2-bromobutane to butan-2-ol. This reaction was Broad Street Scientific | 2021-2022 | 39
chosen because it illustrates the power of Sn2 reactions to change both the functional group and stereochemistry of a molecule. This Sn2 reaction progresses as a nucleophilic hydroxide attacks 2-bromobutane, ejecting bromine as the leaving group and resulting in butan-2-ol as the product (Fig. 2) (5).
Figure 2: Sn2 reaction of 2-bromobutane to butan2-ol (5) The presence of bromine in the reactant classifies the initial molecule as a halide, whereas the addition of a hydroxide via the Sn2 reaction transforms the molecule’s functional group classification from a halide to an alcohol. Additionally, following the Cahn Ingold Prelog rules of priority assignment, 2-bromobutane is determined to have an S stereochemistry which is inverted to an R stereochemistry in the product, butan-2-ol, as a result of the Sn2 reaction. To analyze this reaction’s progression, a variety of molecular energies will be calculated, interpreted, and visualized using computational quantum chemistry tools. The highest occupied molecular orbital (HOMO, characterized as the highest-energy molecular orbital containing electrons) and lowest unoccupied molecular orbital (LUMO, characterized as the lowest-energy molecular orbital not containing electrons) energies will help predict the flow of electrons within the reaction. Reactions with a smaller difference between the HOMO and LUMO energies (known as the HOMO-LUMO energy gap) are energetically favored. Thus, electrons will most likely flow between the HOMO and LUMO with the smallest energy gap. The transition state, as denoted by brackets in Figure 2, can also be computationally predicted using these calculated energy values. Understanding the transition state mechanism through which this reaction progresses enables scientists to better predict the intramolecular interactions and physical mechanisms driving the reaction. Nuclear magnetic resonance (NMR) imaging spectra are commonly used in wet labs to identify the structure of molecules obtained from a reaction. Specifically, NMR spectra are used to confirm that the resulting molecular geometry from a reaction matches the expected structure. An NMR spectrum is characterized by various peaks, each corresponding to the location of a carbon atom in the molecular structure. The location and height of each peak give insight to the specific structure of a compound and enable scientists to differentiate among structurally similar molecules. The peak shifts are commonly obtained 40 | 2021-2022 | Broad Street Scientific
by comparing the shielding constants determined for the reference molecule to a standard compound such as tetramethylsilane (TMS) (7). The ability to computationally predict the NMR spectrum and peak shifts for the desired products of this Sn2 reaction enables scientists to verify that their wet lab products match their computationally predicted products. This paper describes and interprets the computational results from the previously described analyses of the Sn2 reaction of 2-bromobutane to butan-2-ol. Methodology Gaussian 16 (1), MOPAC (2), and WebMO (3) were used to perform the calculations for this work. Mathematica (4) was used to analyze data. Three main computational approaches were applied to study the progression of the selected Sn2 reaction. First, a molecular orbitals calculation was performed on each of the two reactants to determine the highest occupied molecular orbital (HOMO) and lowest unoccupied molecular orbital (LUMO) energies. In combination with qualitative HOMO and LUMO visualization, these numerical data were used to predict the flow of electrons in the course of the reaction’s progression. Next, a transition state calculation was performed using Gaussian with model chemistry B3LYP 6-31G to predict the transition state and visualize the Sn2 reaction pathway. Lastly, a specialized single point energy (SPE) calculation was run on the product butan-2-ol to determine the approximate NMR peaks that would be obtained from a successful reaction in a wet lab. Results Molecular Orbitals Calculations The data obtained from the molecular orbital calculations indicate that the HOMO of 2- bromobutane is orbital 16 with an energy of -10.798 eV and the LUMO of 2-bromobutane is orbital 17 with an energy of -0.134 eV. The HOMO of the hydroxide was determined to be orbital 4 with an energy of -1.010 eV and the LUMO of hydroxide is orbital 5 with an energy of 15.385 eV. Using Mathematica, the HOMO-LUMO gap between the HOMO of 2-bromobutane and the LUMO of hydroxide was calculated to be -26.183 eV. The HOMO-LUMO gap between the HOMO of hydroxide and the LUMO of 2-bromobutane was calculated to be -0.876 eV. The quantitative molecular orbital calculation data are summarized in Tables 1 and 2 below. Table 1: HOMO and LUMO 2-bromobutane and hydroxide
energies
for
CHEMISTRY
Table 2: HOMO-LUMO gap energies between 2-bromobutane (halide) and hydroxide The calculated quantities were also used to perform visualizations of the HOMO and LUMO for both hydroxide and 2-bromobutane (Fig. 3 and 4). These visualizations show that the HOMO lobes of hydroxide and the LUMO lobes of 2-bromobutane would overlap well and be likely to interact due to their complementary shapes and polarities.
Figure 3: HOMO of hydroxide (left) compared with LUMO of 2-bromobutane (right) Figure 3 highlights that the physical configuration of the hydroxide HOMO and 2- bromobutane LUMO are compatible in their sizes, orientations, and polarities. This compatibility indicates that these orbitals will likely interact in this Sn2 reaction.
Figure 5: Energy profile of the transition state for 2-bromobutane to butan-2-ol The energies calculated from the coordinate scan were also compiled to make an energy profile graph of the Sn2 reaction in order to visualize the reaction’s energetic progression (Fig. 5). The transition state reaction energy appears to decrease exponentially as the reaction progresses. NMR Shielding Tensor Calculations Because there are four carbons on the product butan2-ol, there will be four peaks on its NMR spectrum. The isotropic shielding value for the TMS standard was calculated to be 199.9965 parts-per-million (ppm). The isotropic shielding values for each of the four carbons in butan-2-ol were calculated to be 138.2308 ppm, 170.4696 ppm, 187.7832 ppm, and 177.2140 ppm. The predicted relative carbon peak shifts for butan-2-ol were then calculated to be -61.7657 ppm, -29.5269 ppm, -12.2133 ppm, and -22.7825 ppm. These data are summarized in Table 3 below. Table 3: Calculated carbon isotropic shielding values for butan-2-ol and the predicted carbon peak shifts relative to standard TMS carbon isotropic shielding values
Figure 4: HOMO of 2-bromobutane compared with LUMO of hydroxide Transition State Calculations The calculated energies from the coordinate scan allowed for the creation of an animated visualization of the transition state of the Sn2 reaction of 2-bromobutane and butan-2-ol. A video of the animation can be viewed at the following link: https://youtu.be/K5GDXikhXB8. This animation shows the nucleophile (hydroxide) being attracted to the chiral carbon of 2-bromobutane. The bond of the leaving group, bromine, extends throughout the reaction’s progression. The bond between the hydrogen labeled atom 12 and the carbon labeled atom 7 is also extended towards the oxygen of the nucleophile and contracts throughout the reaction’s progression. CHEMISTRY
The NMR shielding tensor and peak shift calculations were then visualized to create an approximate NMR spectrum for the predicted butan-2-ol product of this Sn2 reaction (Fig. 6). The peaks visualized on this spectrum correspond with the calculated peak shifts of -61.7657 ppm, -29.5269 ppm, -12.2133 ppm, and -22.7825 ppm (Table 3).
Broad Street Scientific | 2021-2022 | 41
Figure 6: Predicted carbon NMR spectrum for the butan-2-ol product Discussion The results of each computational analysis performed on this Sn2 reaction allow for a more comprehensive understanding of the physical and energetic processes driving the reaction of 2-bromobutane to butan-2-ol. The molecular orbitals calculations indicate that the smaller HOMO-LUMO energy gap exists between the HOMO of hydroxide and the LUMO of 2-bromobutane (Fig. 2). Reactions with smaller magnitude HOMO LUMO gaps are more energetically favored. Thus, the HOMO - LUMO gap from the HOMO of the hydroxide to the LUMO of 2- bromobutane will be more energetically favored. Furthermore, Figure 3 visualizes that the HOMO of hydroxide and the LUMO of 2- bromobutane are compatible based on their similar size, shape, and polarity. From these quantitative and qualitative results, it can be predicted that the Sn2 reaction will occur between the HOMO of the hydroxide and the LUMO of 2-bromobutane, indicating that electrons will flow from the hydroxide to 2-bromobutane as the reaction progresses. The transition state animation highlights how the attraction between the nucleophilic hydroxide and chiral carbon of 2-bromobutane is a driving factor in this Sn2 reaction. As the hydroxide is drawn closer to the chiral carbon, the bond between bromine and this central carbon extends and weakens, priming this bond to be broken as the hydroxide is incorporated into the product (butan-2-ol) and bromine is ejected as a leaving group. The lengthened bond between hydrogen 12 and carbon 7 towards the nucleophilic hydroxide indicates that hydrogen bonding may be involved as a physical mechanism driving the attraction of hydroxide towards 2-bromobutane. The energy profile of the transition state progression for the reaction of 2-bromobutane to butan-2-ol indicates that the reaction is also energetically driven. This profile shows that the energy of the transition state decreases as the reaction progresses (Fig. 5). Molecular structures naturally seek lower-energy states, thus the reduction in energy as the reaction progresses indicates that the Sn2 reaction of 2-bromobutane to butan-2-ol is driven by 42 | 2021-2022 | Broad Street Scientific
energetics in addition to physical mechanisms. The predicted carbon NMR relative peak shifts indicate that a wet lab NMR of a successful Sn2 reaction of 2-bromobutane to butan-2-ol should contain four peaks at approximately -61.7657 ppm, -29.5269 ppm, -12.2133 ppm, and -22.7825 ppm (Table 3). These peaks are visualized on the predicted NMR spectrum (Fig. 6). A key difference between this approximated NMR spectrum and a spectrum obtained from a physical sample is that the experimental spectrum would contain smaller interference peaks in addition to the main carbon peaks that indicate nuanced structural orientations. Because the peak shifts are negative, it can be concluded that there is more shielding present in butan-2-ol than in the tetramethylsilane comparison standard. Conclusions The computational results indicate that the Sn2 reaction of 2-bromobutane to butan-2-ol is driven by both energetic and physical mechanisms. The molecular orbitals analyses demonstrate that electrons will most likely flow from the hydroxide HOMO to the 2-bromobutane LUMO as a result of a physical interaction between these compatible orbitals. The transition state visualization indicates that the hydroxide is drawn to the chiral carbon via a combination of nucleophilic and hydrogen-bonding interactions. The bromine is ejected as a leaving group due to increasing bond length throughout the reaction’s progression. The energy profile diagram of the transition state reaction progression also shows that this Sn2 reaction is energetically driven because the transition state decreases in energy throughout the reaction. The predicted carbon NMR spectrum of the expected product, butan-2-ol, indicates four peaks located at -61.7657 ppm, -29.5269 ppm, -12.2133 ppm, and -22.7825 ppm. This predicted NMR spectrum could be used in a wet lab setting to verify that the experimentallyobtained NMR spectrum identifies the expected product (butan-2-ol) for this Sn2 reaction. These computational approaches could be applied to other similar Sn2 reactions before experimentally conducting the reaction in a lab to reduce time and money wasted on errors in the lab that could have been prevented by better understanding the mechanisms driving the reactions. For example, if the calculations indicate that a reaction is not energetically favored, scientists would know to adjust the experimental conditions in order to obtain a successful initial trial. Predicting the NMR spectra would enable researchers to more quickly verify whether or not the experimental trial successfully produced the desired product. Because of the wide range of applications for Sn2 reactions, from pharmaceutical drug design to nanotechnology synthesis, developing a more comprehensive understanding of the mechanisms CHEMISTRY
driving an Sn2 reaction prior to experimentation in a wet lab allows for reduced time and costs associated with research in the variety of fields employing Sn2 reactions. Acknowledgements The author expresses appreciation and gratitude to Mr. Robert Gotwals, NCSSM Computational Science Educator, and the NCSSM Online Program for support of this opportunity. Additional thanks are given to Dr. William Sponholtz for his assistance in locating resources to aid in the research of specific Sn2 reactions. References 1. Gaussian 16, Revision C.01, M. J. Frisch, G. W. Trucks, et al. Gaussian, Inc., Wallingford CT, 2016. 2. MOPAC2016, J.J.P. Stewart, Stewart Computational Chemistry, Colorado Springs, CO, USA. 3. Polik, W.F., Schmidt, J.R. WebMO: Web-based computational chemistry calculations in education and research. WIREs Comput Mol Sci. 2021;e1554. https://doi. org/10.1002/wcms.1554. 4. Wolfram Research, Inc., Mathematica, Version 12.3.1, Champaign, IL (2021). 5. Sponholtz, William. “Part I What is Organic Chemistry: The Sn2 Reaction.” AP Chemistry, 2 December 2021, Forsyth Country Day School. Class handout. 6. “WebMO 04 - (Transition State of an SN2 Reaction).” YouTube, uploaded by Science Aura, 27 July 2020, https://www.youtube.com/watch?v=XhTnxlKSGsoab channel=ScienceAura. 7. Foresman, James B., and Æleen Frisch. Exploring Chemistry with Electronic Structure Methods. Second edition, Gaussian, 1996.
CHEMISTRY
Broad Street Scientific | 2021-2022 | 43
RATIONAL DESIGN AND SYNTHESIS OF A NOVEL CLASS OF BORONIC-ACID CONTAINING TUBULIN INHIBITORS AS TUMOR VASCULAR DISRUPTING AND ANTIPROLIFERATIVE AGENTS Winnie Wang Abstract Due to its vital role in tumor vasculature structure and cell proliferation, the tubulin colchicine complex is a key target in the development of antitumor drugs. Once tubulin polymerization is inhibited, the vascular system within tumoral masses can be made to collapse, leading to necrosis in the tumor’s core and mitotic arrest in proliferating cells. Though there are a number of existing vascular disrupting agents, such as Combretastatin A-4 (CA-4), the limitations of these drugs are their low Oral Non CNS Scores and associated pharmaceutical profiles, rendering them poor candidates for cancer therapies. However, research has suggested structure-activity relationships for CA-4 that point to the possible success of inhibitory analogs with specific structural modifications. Hence, this study proposes a novel class of boronic acid containing small molecule vascular disrupting and antiproliferative agents that have the potential to improve the inhibition and oral administration efficacy of CA-4. Inhibitors were designed using a 4-phase drug design framework in Schrodinger Maestro. The docking scores of the compounds were predicted in Schrodinger Maestro and the compounds were further assessed for pharmaceutical efficacy through StarDrop scoring profiles. Analysis of docking scores and scoring profiles identified a class of novel compounds as the most optimal candidates for tubulin depolymerization and vascular disruption. One of the most successful and synthetically viable compounds, D8, was selected for synthesis. This particular candidate demonstrates better predicted binding and oral administration capabilities over multiple clinically relevant tubulin inhibitors. This study successfully synthesized the top compound through a 3-step pathway featuring three reactions using commercially available reactants with a 60% yield. 1. Introduction 1.1 Cancer and Targeted Cell Therapies In a population that continues to see a rising lifespan, cancer is becoming a progressively high risk in many demographics (National Cancer Institute 2020). Approximately 40% of people will be diagnosed with cancer during their lives, and it has become the second leading cause of death worldwide. In 2020, there were 6.1 million cancer-related deaths and 18.1 million new cases (Sung et al. 2021). As global expenditures for cancer care continue to climb, the development of novel and more effective cancer treatments lies at the forefront of oncological research. Given the heterogeneity of cancer, current cancer therapy is largely limited to surgery, radiation, and chemotherapy, all of which aim to diminish cancerous cells and remove tumors. However, treatment failure is persistent due to the difficulty associated with gaining complete local control of tumors and reducing metastatic spread. An emerging form of treatment is Targeted Cancer Therapies (TCTs) (Group et al. 2015). TCTs are drugs that prevent the proliferation of cancer cells and combat existing vessels by binding to specific targets that are overexpressed or involved in cancerous pathways. In con44 | 2021-2022 | Broad Street Scientific
trast to existing forms of chemotherapy that largely fail to distinguish between proliferating and healthy cells, TCTs specifically target molecular elements associated with cancer. Furthermore, in contrast to the cytotoxicity of chemotherapy in relation to killing tumor cells, TCTs are cytostatic: they inhibit cell proliferation. Hence, TCTs are generally less toxic and have a higher specificity to tumor cells, offsetting the adverse effects of chemotherapy (National Cancer Institute 2020). 1.2 Characterization of the Tumor Vasculature and Endothelial Cells A continuously expanding vessel network is essential for tumor development, growth, survival, and spread, through the provision of oxygen and nutrients (Mita et al. 2013). Conventional anticancer therapies are impaired by the tumor endothelium, and often lead to cells with aggressive phenotypes and greater metastatic potential. Vascularization is required for tumor growth and metastasis (Hinnen and Eskens 2007). Tumor vasculature is composed of endothelial cells, which are planar cells that span the luminal surface of blood vessels, thus ensuring laminar blood flow (Jameson et al. 2003). These cells carry out their essential funcCHEMISTRY
tions (migration, proliferation, adhesion, and cohesion) with a strong dependency on the structural and functional integrity of their cytoskeletons. Microtubules are dynamic structures that constitute the cellular cytoskeleton, together with actin microfilaments and intermediate filaments (Dumontet and Jordan 2010). Beyond their well-known role in making up the mitotic spindle for cell division, they also maintain cell shape and morphology, cellular motility, and trafficking of vesicles and organelles (Choi et al. 2013). Significant differences exist between normal vessels and vessels found in the tumor microenvironment. The normal vasculature is a well-organized network that is strictly balanced between proangiogenic and antiangiogenic factors (Segaoula et al. 2016). In contrast, tumor blood vessels are immature and tortuous, given that they are composed primarily of rapidly proliferating endothelial cells and lack the structural support of connective tissue (Fig. 1). Overexpression of proangiogenic factors in the tumor environment also results in an unevenly distributed and highly chaotic vasculature. They are typically structurally abnormal as they lack pericytes and proper basement membrane, resulting in fragile and immature vessels. These abnormalities significantly impact vessel function — blood vessel permeability is unusually high, leading to increased interstitial pressure (Siemann 2011). Blood flow is also reduced. Thus, any alteration of blood vessel perfusion that would not affect normal vessels may be catastrophic for the intratumoral vascular system, leading to ischemia and tumor cell death (Fig. 1) (Siemann 2011). The inherent differences in physiological and structural characteristics between tumorous blood vessels and normal tissues provide unique targeting opportunities for novel therapeutic strategies (Schaaf et al. 2018). The high specificity of tumor vasculature is also accompanied by the accessibility and genetic stability of endothelial cells, making this approach very appealing. Highly specific targeting also minimizes systemic side effects, and could provide a universal therapy for solid tumors, regardless of histology (GM Daenen et al. 2010). 1.3 Background on Antiangiogenesis Agents, Spindle Poisons, and Vascular Disrupting Agents Two main approaches have been proposed to target tumor vascularization: (1) inhibition of angiogenesis by inhibiting the formation of new blood vessels, and (2) vascular disruption, in which the existing vasculature is targeted (Hinnen and Eskens 2007). Both approaches are directed towards endothelial cells due to their high selectivity and genetic stability. Although inhibiting angiogenesis has been successful in many conditions, clinical data indicate that treatment is often accompanied by lack CHEMISTRY
of efficacy, resistance development, and toxicity, hence an urgent need for a more optimal treatment in metastatic disease (GM Daenen et al. 2010).
Figure 1: Normal vasculature has a well-organized network of vessels that efficiently deliver nutrient and oxygen supplies. These vessels are well-matured, and the layer of endothelial cells is surrounded by a basement membrane and pericytes, in comparison to less organized, structurally lacking tumor vasculature. This diagram illustrates the pathway of certain physiological traits in tumorous vascular endothelial cells. Adapted from Schaaf et al. 2018. Created using BioRender. Research also exists on mitotic spindle poisons, which aim to halt cell cycle progression by preventing the polymerization of tubulin and disrupting the function of microtubules. Though many effective mitosis inhibitors have been developed, the majority have failed in clinical trials due to a poor therapeutic index (Lu et al. 2012). On the other hand, vascular disrupting agents (VDAs) target endothelial cells of the established tumor vasculature, exhibiting an immediate cytotoxic effect on the existing tumor cells. VDAs induce changes in endothelial cell shape via cytoskeleton disruption, leading to increased permeability to proteins and increased interstitial fluid pressure, minimizing vessel diameter (Chaplin et al. 2006). They do not seem to elicit many of the adverse effects associated with cytotoxic agents or angiogenesis inhibitors, and drug delivery is likely to be uncompromisable due to the accessibility and selectivity of the tumor vasculature. The combined inhibition of blood flow associated with decreased vessel diameter and the subsequent compromised supply of oxygen and nutrients will induce cell death of many tumor cells. Theoretically, this vascular collapse will cause a massive downstream tumor cell killing, leading to necrosis at the core of the tumor. Furthermore, complementary efficacy has been hypothesized, including efficacy in combination treatment with existing chemotherapies, Broad Street Scientific | 2021-2022 | 45
and synergy with both angiogenesis inhibitors and cytotoxic agents. Hence, the indirect killing of tumors through compromising their vascularization is an attractive anticancer treatment approach (Dumontet and Jordan 2010). 1.4 Tubulin Colchicine Complex — Inhibition Microtubules are formed through the polymerization of heterodimers of α and β tubulin, which undergo a “curved to straight” structural transformation to form the core. Composed of linear rows of alternating α- and β-tubulin, microtubules are dynamic structures and rapidly assemble/disassemble to meet the cell’s needs. Disruption of tubulin polymerization thus disrupts the formation of tumor vasculature, which depends on microtubules that form the cytoskeleton (Wang et al. 2016). Small molecule inhibitors have shown to bind at three major binding sites on tubulin: the vinca, taxane, and colchicine sites. Colchicine is a natural alkaloid, and was originally used for the identification of protein tubulin. This compound was discovered to be a potent cytotoxic agent that binds tubulin at the αβ-interface and causes microtubule depolymerization by inhibiting assembly. Colchicine binding to β-tubulin results in a curved tubulin dimer, preventing it from adopting a straight structure. Due to a steric clash between colchicine and α-tubulin, microtubule assembly is inhibited. However, colchicine’s low therapeutic index and high toxicity disqualifies it as an effective anticancer agent. Though the vinca and taxane sites have well-explored potentials in mitotic poison and anticancer treatment development, the therapeutic potential of the colchicine site remains largely undiscovered (Fig. 2) (Hinnen and Eskens 2007). In their work, Pettit et al. uncovered the antiproliferative and tubulin inhibitory effects of active compounds extracted from the south African tree Combretum caffrum (Pettit et al. 1989). The cis-stilbene Combretastatin A- 4 was found to be more potent than colchicine, while also having antimitotic activity in rapidly proliferating cells through interference with mitotic spindle formation and chromosome alignment in metaphase. The binding of CA-4 to tubulin primarily inhibits tubulin polymerization, activating the protein RhoA which is crucial for coordinating the interactions between F-actin and microtubules in the cytoskeleton (Nguyen et al. 2005). However, clinical trials were inhibited by its poor solubility. A water soluble phosphate ester, CA- 4P was developed for clinical use, but this thermodynamically favored isomer is much less active, and has adverse side effects related to cardiovascular and neurological damage (Fig. 3) (Chaplin et al. 2006). Inhibitors designed as analogues of combretastatin A-4 have the potential to reduce the adverse pharmacodynamic and pharmacokinetic interactions associated 46 | 2021-2022 | Broad Street Scientific
with combretastatin A-4 and many other VDAs, and fill an urgent need for more effective tubulin depolymerizing agents (Hinnen and Eskens 2007). Furthermore, employing a phenstatin-based scaffold to design this inhibitor has potential to reduce the adverse effects of highly cytotoxic stilbene analogues, as the phenstatin class of compounds are hypothesized to be less cytotoxic while still highly effective (Tron et al. 2006). From a pharmaceutical perspective, any form of TCT should be designed for oral administration. Orally administered drugs are ideal as they are safe, convenient, and reduce additional costs such as equipment and medical personnel time (Cyriac et al. 2014).
Figure 2: Structural overview of the colchicine-tubulin complex and detailed insight into the binding site of combretastatin. Tubulin’s α subunit is represented in green and pink, and the β subunit in cyan and yellow. Combretastatin is represented in red in ball-andstick form. The expanded view of the binding pocket displays the protein surface within 2.50 Å of the ligand, and has a color scheme corresponding to atom partial charge in the surrounding protein. The colchicine binding site is located in the intermediate interface between the α and β subunits of each tubulin dimer, and is largely buried in the domain of the αβ subunit. In this binding mode, the methoxy group of ring A establishes a hydrogen-bond interaction with Cys241, while ring B interacts with the N-terminal nucleotide binding domain within the β subunit by a H-bond. This model also shows the growth space outside of the ligand, displaying how combretastatin binding might affect tubulin dynamics. Adapted from Pērez-Pērez et al. 2016. Created using BioRender.
CHEMISTRY
3. Methodology
tions (mol MW < 500, QPlogPo/w < 5, donorHB ≤ 5, accptHB ≤ 10). After the ligands were designed and docked in Maestro, they were then imported into Stardrop, which is an optimization software featuring numerous pharmacokinetic scoring profiles. Profiles computed include Lipinski’s Rule of Five, Oral Non CNS Score, Human Intestinal Absorption (HIA), and Solubility. In StarDrop, an ideal drug will achieve a score of 1 (Tan and Kirchmair 2014). The Oral Non CNS is a relative predictive value of a drug’s ability to function as an orally administered, non-central nervous system targeting drug. A drug with a higher Oral Non CNS Score is preferable, and would have a higher predictive success as an orally administered compound. StarDrop glowing molecules were used to analyze the contribution of each modification to the aforementioned scoring profiles, and furthermore manipulated in the computation of structure-activity relationship determinations for the scaffold compound. Red areas of the glowing molecule indicate a greater contribution to the respective scoring profile whereas blue areas indicate minimal contribution (Obrezanova and Segall 2009). When a top compound had been selected, Schrodinger Reaction Based Enumeration was further utilized to generate possible synthetic paths of the different moieties in the final compound. To analyze the correlation between pharmaceutical scoring profiles and create idealized comparisons, charts and graphs were generated using the Visualization module on StarDrop and the Ligand Designer and Display tools on Maestro.
3.1 Assessment of Pharmaceutical Efficacy
3.2 Protein Preparation and Ligand Docking
200,000 novel compounds were designed in Schrodinger Maestro by the 2D Ligand Sketcher and Reaction Based Enumeration tools. The Schrodinger Suite’s Maestro software was utilized for its ligand building, visualization, and scoring capabilities (Schrödinger 2012). Drug design occurred via skeleton modification with the 2D Ligand Sketcher, in which modifications were made to the base scaffold to produce analogs. The Ligand Designer tool has a diverse range of capabilities such as 3D Ideation, Form Protein-Ligand Interactions, Cyclize Ligands, and Reaction Based Enumeration. Though 3D Ideation, Form Protein-Ligand Interactions, and Cyclize Ligands were used to analyze general ligand-protein interactions, Reaction Based Enumeration was used to quickly generate new analogs. In Reaction Based Enumeration, the drug profile was calculated using Qikprop, a tool that predicts numerous descriptors of pharmaceutically relevant properties that can identify and eliminate compounds falling outside the ranges for normal drugs. The scoring profile in focus was Lipinski’s Rule of Five, a set of criteria used to evaluate a ligand’s drug-likeness, relating to outcomes such as oral activity and cytotoxicity. In Maestro, each rule defines specific recommendations for drug interac-
The 3D crystal structure of the tubulin colchicine complex (PDB: 402B) was retrieved from the RCSB Protein Data Bank, and was selected based on its fulfillment of metrics according to the Data Bank’s validation assessment. This structure was imported into Schrodinger to complete computer-simulated docking analysis of compounds. Following the preparation of the protein and creation of the docking grid, combretastatin was prepared using Maestro’s LigPrep, following standardized criteria (desalt, generate tautomers, target pH 7.0+/-2.0). To validate the crystal structure, well-binding tubulin inhibitors and non-binding decoys were sourced from the DUD-E database and docked using the Enrichment Calculator. Prior to ligand docking, the proteins were prepared by removing waters, assigning bond orders, adding hydrogens, and filling in missing side chains and loops. Irrelevant ligands and other hET groups attached to the structure were also removed for purposes of efficiency. The binding site of the domain was determined through the presence of pre-docked ligands in conjunction with amino acid presence and cofactor proximity based on previously described binding sites on tubulin
Figure 3: (Left) Structure of CA-4 pharmacophore with key moieties identified in dark green. (Right) Structure of benzene with carbon position numbers. Created using BioRender. 2. Objectives This study aims to design a novel tubulin inhibitor with potential for oral administration. This is the first known study focused on applying rational targeted modifications to design a class of anticancer drugs having dual effects: tumor vasculature disruption and mitotic arrest. Combretastatin A-4 served as a base scaffold to design this compound, with phenstatin as a parallel reference. The rationale behind the design framework formulated and implemented in this work may be a useful reference for the development of dual function inhibitors targeting other proteins, as it accounts for the inhibition, pharmaceutical efficacies, and synthesis of potential drugs.
CHEMISTRY
Broad Street Scientific | 2021-2022 | 47
(Lu et al. 2012). Using Maestro’s Receptor Grid Generator, a receptor grid was calculated for docking using colchicine as the grid-defining ligand. Default van der Waals radius scaling parameters were used (scaling factor of 1, partial charge cutoff of 0.25). Ligands pre-docked in the crystallized structures were retained to serve as benzene analog references for ligand pose orientation. Combretastatin, sourced from PubChem, was imported using its SMILES code (COC1=C(C=C(C=C1)/C=C2=CC(=C(C(=C2) OC)OC)OC)O) and superimposed onto colchicine using Maestro’s Ligand Alignment. After combretastatin was incorporated, ligands were docked using Maestro’s Ligand Docking function. Standard precision was used, and flexible ligand sampling was preferred with an emphasis on nitrogen inversions and ring conformations, such that multiple generated poses were available for comparison. The most accurate pose of each ligand was determined based on the physical orientation alignment of the ligand to the analogs present in the structure, as well as its state penalty, which measures how optimized the pose is in regards to its free energy. In Maestro, a state penalty closer to 0 is better, with 0.00 being the optimal value. Quantitative and positional information about interactions with protein residues and ligands were obtained using Maestro’s Ligand Interaction module. Visualizations of hydrogen bond strengths/levels, waters, and solvent exposures were viewed using this docking workspace, defining an ideal binding cavity. Furthermore, ligand interactions with the binding cavity were explored using Maestro’s Ligand Designer, which provides a visual of growth space and ligand-receptor interactions.
Oral Administration: Due to its low Oral Non CNS Score, CA-4 is a poor candidate for oral administration primarily due to a negative HIA in conjunction with extremely low solubility reported in clinical trials — hence the necessity for a prodrug formulation in testing. However, based on the Oral Non CNS glowing molecule from StarDrop, phenols at every position on A and B rings contribute the most to the Oral Non CNS score (Tron et al. 2006). Solubility: CA-4 was not successful in clinical trials as its poor solubility indicated a low bioavailability, resulting in suboptimal drug delivery. Thus, an effective concentration of the drug may not reach the tumor site. The aliphatic linker is predicted to have a large contribution to the low solubility of the original CA-4 (Tron et al. 2006). hERG Score: CA-4 is very cytotoxic, achieving a suboptimal hERG score. Based on these findings, a 4-phase drug design framework aimed to create a ligand with improved docking score to tubulin, potential for oral administration, and improved solubility and hERG scores, while adhering to Lipinski’s Rule of 5 to confirm drug viability.
3.3 Preliminary Analysis of Combretastatin
Figure 4A: CA-4 Oral Non CNS Score glowing molecule exported from StarDrop. Blue = limited contribution. Green = moderate contribution Red = major contribution. Figure 4B: CA-4 bound to tubulin. Figure 4C: CA-4 orientation in relation to nearby residues and the T5 loop.
Preliminary analysis of CA-4 was conducted to guide the rational development of the 4-phase drug design framework employed in this research. This approach was designed based off of the assessment of docking score, oral administration capability, solubility, and hERG score of CA-4 (Fig. 4). Below are the key findings: Docking Score: CA-4 has a moderate docking score value to tubulin. Analysis of protein-ligand interactions revealed that substituents, especially at the carbon linker, display little meaningful interaction with surrounding residues in tubulin, despite having a large contribution to a negative hERG score (Karatoprak et al.). Alternatively, the 3,4,5-trimethoxy on the methoxylated benzene ring in CA-4 provides key H-bond interactions, and is hypothesized to serve as an anchor to place the B ring into the proper orientation within the binding site (Tron et al. 2006). Optimizing the carbon bond and B ring substituents while preserving the interactions between the methoxylated benzene ring and the protein appears to be a promising avenue for the improvement of binding. 48 | 2021-2022 | Broad Street Scientific
3.4 Drug Design Framework Phase I: To be an effective inhibitor, it is most essential that the ligand has a high docking score to tubulin. Thus, improving docking score was selected as the overarching goal of Phase I. The 3- and 5-positions of the B benzene were selected for Phase I modifications based on preliminary analysis in conjunction with prior research that reported significant contribution to docking score based on modifications made at these sites. Modifications included the expansion of the carbonyl bridge linking the two benzene rings, the substitution of the methoxy groups on the methoxylated A benzene, and the addition of -ortho, -meta, and -para substituted boronic acids, bromines, and fluorines. In this phase, chalconoids and fluorinated CHEMISTRY
cis-restricted moieties were explored (Alloatti et al.), but many analogs featuring the latter conformation were discarded due to poor state penalties. Boronic acid quickly became a focus of this phase, due to the highly directional and specific interactions with nearby residues that it formed, which explains the large improvement in docking score (Silva et al. 2020). Phase II: Phase II of this drug design framework aimed to maintain or improve the docking score, and improve the Oral Non CNS Score. Oral administration is not only more viable than intravenous administration, but it is the preferable route of administration for Targeted Cancer Therapies. Furthermore, orally administered vascular disrupting agents have been successful in clinical trials in comparison to their intravenous counterparts. Thus, Phase II integrated the goals of inhibition and oral administration, focusing on the creation of compounds that had improved Oral Non CNS Scores while conserving or improving the refined docking scores achieved in Phase I. Preliminary analysis of CA-4 indicates that the carbon linker contributes minimally to binding, while contributing greatly to the Oral Non CNS Score. Conversely, this moiety is also the greatest contributor to cytotoxicity and hERG score. Thus, this double bond was selected for modification to boost the Oral Non CNS Score and counteract the unfavorable properties associated with it. Modifications specifically targeted designing different varieties of carbon linkers for the two benzene rings. Fluorinated chalcones, inverse chalcones, and expanded carbonyl analogs were explored, as was the addition of an anisole at different positions on the B benzene. Various functional groups were chosen, many of which were polar and had hydrogen bonding capabilities, as substituents for the B ring as well (Tron et al. 2006). In respect to the Oral Non CNS Score, ionizable and polar groups were selected to increase solubility (ie. boronic acid), and transitively, the Oral Non CNS Score improved as well (Ghose et al. 2017). Phase III: For phase III of this drug design framework, the goal was to further optimize the Oral Non CNS Score. Since binding to tubulin had been optimized, as had drug likeability in terms of solubility and lipophilicity in Phases I and II, the objective of this phase was to create ligands with improved oral administration potential as well as reduced cytotoxicity. The human ether-a-gogo (hERG) gene encodes the voltage gated potassium channels in the heart, which are involved in cardiac repolarization. If the hERG current is overly inhibited, QT interval prolongation resulting in potentially fatal ventricular tachyarrhythmia could occur (Engelbrechtsen et al. 2018). Chalconated structures had carbon linkers that had large contributions to a nonoptimal hERG score, so modification to the carbon linker was once again the focus of this phase. Previous research has hypothesized that two carbon linkers yielded the most active combretastaCHEMISTRY
tin analogs, so a benzophenone structure was explored in the making of these compounds, and was found to be the most potent thus far. Compounds with the phenone moiety also had the most improved Oral Non CNS Scores, while having minimal contribution to cytotoxicity (Kalepu and Nekkanti 2015). Phase IV: Phase IV aimed to enumerate analogs of the most successful compound from Phase III. Due to the ease of synthesis of combretastatin analogs, testing is relatively uncomplicated, and a large amount of compounds can be synthesized and tested at once. Previous research has proposed that boronic acid derivatives should receive considerable attention in drug development, due to their moderate pH of 9-10 and general air stability (Silva et al. 2020). Ideally, all derivatives would achieve docking scores and Oral Non CNS Scores in a small range. The most successful ligand from Phase III was enumerated to create a class of similarly structured and potent compounds, featuring boronic acid bioisosteres and a phenone carbon linker.
Figure 5: 4-step drug design framework. 3.5 Synthesis Pathway Framework Following the creation of a successful, novel small molecule inhibitor of tubulin, reaction based enumerations of the compound were run to determine possible synthetic pathways. Finding that Aryl Miyaura-Suzuki Cross-Couplings and Borylations of arylboronic acids were the basis of a synthetic pathway for this structure, previous research was referenced in conjunction with Maestro’s reaction-based enumeration to establish reactions (Kwon et al. 2010) (Mfuh et al. 2016). Though D8 (Fig. 9) was only the 3rd most successful ligand in terms of both Oral Non CNS Score and docking score, it had reactants that were commercially abundant and a straightforward synthesis plan. A three-step synthesis pathway was identified for D8, using commercially available reagents and catalysts. All chemicals and reagents were sourced from Fisher Scientific and Sigma-Aldrich. Objectives: In synthesizing D8, emphasis was placed on ease of synthesis execution, with feasibility of a successful synthesis given the tools available. High yields are also preferable, as were well-formulated purification reactions to maximize yield. Rationale: Reaction based synthesis enumeration in Schrodinger originally yielded primarily synthesis pathways with fewer steps. However, many of those pathways utilized organo-tin compounds, which are highly Broad Street Scientific | 2021-2022 | 49
volatile and difficult to handle. A more efficient synthesis was sacrificed in the interest of safety. The photoinduced borylation, proposed by Mfuh et al., had an experimental yield of 92% (Mfuh et al. 2016). Adaptations were made to optimize the synthesis for a 3,5-dibromophenol, but the reaction remained additive-free and scalable, which was ideal given that it would be repeated in a later step with a differing quantity and structurally different bromophenol. The modified Suzuki cross coupling was originally proposed by Yoon et al., and was adapted for a trimethoxybenzoic acid from its original intended pathway for benzoic acids (Kwon et al. 2010). Vascular disrupting agents as derivatives of Combretastatin A4 are typically synthesized through Friedel-Crafts acylations and Wittig and Perkin condensations. However, Wittig reactions require tedious further purification procedures, often requiring a high-temperature, decarboxylative process to obtain high yields of a final product (Tron et al. 2006). Furthermore, the Friedel-Crafts acylations have significant drawbacks, such as limited regioselectivity, harsh reaction conditions, and difficulty in separation and purification. The Suzuki reaction has an average yield of 80%, and boasts a straightforward purification. Therefore, it is proposed that this alternative pathway is of greater eas2e and synthetic efficacy than existing pathways for vascular disrupting agents and CA-4 analogs, and may be modified for further research in boronic acid containing analogs or benzophenones. (Fig. 6)
3.6 Synthesis Scheme Step 1: Scalable, Metal and Additive Free, Photoinduced Borylation Reactants: 3,5-Dibromophenol, Tetrahydroxydiboron, anhydrous methanol. The synthesis started with 10 ml of 3,5-Dibromophenol and 1.793 ml tetrahydroxydiboron dissolved in 100 ml anhydrous methanol under 254 mm ultraviolet irradiation for 24 hours. The anhydrous methanol was prepared with methanol in 3 Å molecular sieves. Reaction conditions recommended a temperature of 15° C, so the reaction occurred in an ice bath that was heated up to 15° C throughout the experiment duration. The resulting reaction mixture was purified in additional steps. First, the reaction mixture (scale 1mmol) was transferred to a round bottom flask with 10.08 grams of sodium bicarbonate (3 mmol). The remaining methanol was evaporated from the flask using a rotary evaporator, and the remaining solid was dissolved in 1M aqueous solution of fructose (15 mL), and 1M aqueous solution of sodium carbonate (15 mL). Ethyl acetate (50 mL) was added along with the dissolved solution into a separatory funnel, and the organic portion was separated and discarded. The remaining aqueous phase was acidified to pH 3 using 6M aqueous hydrochloric acid, and then extracted with ethyl acetate again (2 x 30 mL). The combined organic portions resulting from the subsequent separations were dried over anhydrous sodium sulfate, and solid sodium bicarbonate (252 mg) was added. The mixture was filtered and concentrated under reduced pressure in the rotary evaporator to yield 5g of the desired boronic acid, a dark brown solid (55% yield) (Mfuh et al. 2016). Step II: Palladium-Catalyzed Synthesis of Aryl Ketones from Carboxylic Acid Reactants: 3,4,5-trimethoxybenzoic acid, N-Ethoxycarbonyl-2-ethoxy-1,2-dihydroquinoline (EEDQ), Tetrakis (triphenyl phosphine) palladium, Water, and Dimethylformamide.
Figure 6: 3-step synthesis scheme for D8, including reagents and reaction conditions. I = Step I; II = Step II; III = Step III. Created using ChemDraw and BioRender. 50 | 2021-2022 | Broad Street Scientific
To a solution of 3,4,5-trimethoxybenzoic acid (1.773 g, 0.66 mmol), EEDQ (3.09938 g, 0.99 mmol), and Pd(PPh3)4 (0.2926 g, 0.02 mmol) in DMF (12.66 mL) and water (0.38 mL, 1.66 mmol) in a 50 mL reaction vessel was added aryl boronic acid (5.03 g, 0.79 mmol). The resulting solution was purged with argon via a nylon balloon. The reaction mixture was covered with aluminum foil to shield the light-sensitive palladium catalyst, and was stirred for 15 hours at 60° C. The reaction was monitored by TLC (ethyl acetate: n-hexane = 1:10). After the reaction stopped, the CHEMISTRY
mixture was quenched with 10 mL water, and the aqueous solution was extracted with ethyl acetate via separatory funnel (5 mL x 3). The combined ethyl acetate organic phases were concentrated, and further purification of the product was achieved by column chromatography on silica gel (ethyl acetate : n-hexane = 1:10) to yield a boronic acid containing 3,4,5-methoxybenzophenone (86%) as a yellow solid. The remaining DMF was removed by extracting the aqueous phase in a separatory funnel with water and ether (Kwon et al. 2010). Step III: Scalable, Metal and Additive Free, Photoinduced Borylation Part 2 Reactants: Tetrahydroxydiboron, anhydrous methanol. The result of the previous step was added to a test tube containing 90 mL of anhydrous methanol (see Step I above) and 1.45 g of tetrahydroxydiboron. The first step of the synthesis was repeated with this new product. The same purification steps were repeated to yield a white solid (60%) (Mfuh et al. 2016).
The B-ring methoxy was hypothesized to have contributed largely to the cytotoxicity of the compound, so the removal of that methoxy group reduced the hERG score significantly. The inverted chalcone structure did not form H-bonds well within the binding site, so this structure was neglected in future phases. Analysis of this phase supports the hypothesis that the carbon linker between the two aromatic rings contributes significantly to both cytotoxicity and docking score, and future phases will focus on double bond modifications to improve both states. Phase I was successful in designing ligands with improved binding. Though the focus was on docking score, significant improvements were made in terms of hERG score and solubility as well. Due to their high docking score values, compounds D1, D2, D3, and D4 were selected as the starting points for Phase II (Fig. 7). Based on analysis of the Lipophilicity glowing molecules in StarDrop, it is proposed that the presence of H-bond accepting groups (methoxy) improve lipophilicity and transitively, intestinal permeability represented by the HIA (Fig. 8) (Submanian and Kitchen 2006).
4. Results & Discussion 4.1 Phase I Overview and Analysis Goal: Improve binding to tubulin. Most compounds docked in Phase I have improved docking scores to tubulin. Phase I explored modifying the double bond by replacing it with a carbonyl group, attempting to replicate the chemically privileged chalcone structure. Hundreds of chalcone analogs were designed or enumerated with a carbonyl expanded linker, as well as an inverse carbonyl expansion. To improve solubility, a boronic acid was added at ortho, meta, and para positions on the B ring, though many of these chalconated structures do not achieve a Lipinski’s Rule of Five score of 1 due to the heavy molecular weight of borons and chalconoids in general (Schobert et al.). To improve binding, the trimethoxy benzene A ring was modified, placing methoxy groups at 1, 3, 5 positions as well as 2, 3, 4 positions, which would differentiate the mode by which the compound orients itself in the binding site. The trimethoxy A ring is hypothesized to anchor the ligand into the binding site via hydrogen bonding with the nearby proton acceptor (Tron et al. 2006). In comparison to CA-4, compounds D1, D3, and D4 demonstrated improved binding to topo II. All compounds maintain original H-bond interactions, and the strength of the Pro175 H-bond is improved. Furthermore, the boronic acid at position 3 in D1 also introduces two new H-bond interactions with Asn350 and Asn349, not seen between combretastatin A-4 and the protein. CHEMISTRY
Figure 7: Binding affinity values, hERG scores, logS values, Lipinski’s Rule of 5 values, and Oral Non CNS Scores of the best compounds from each phase. Green = improved value/score compared to the parent ligand.
Broad Street Scientific | 2021-2022 | 51
Figure 9: Molecular structures of the most successful ligands from each phase exported from Schrodinger 2D Designer. Figure 8: StarDrop glowing molecules (from left to right): HIA (Human Intestinal Absorption) = ideal score is +. Lipophilicity = ideal score is +. Blue = limited contribution. Green = moderate contribution Red = major contribution. 4.2 Phase II Overview and Analysis Goals: Improve the Oral Non CNS Score 2, maintain or improve docking score, reduce cytotoxicity on the double bond. Most ligands created in this phase displayed improved Oral Non CNS Scores in comparison to their parent ligands D2, D3, and D4 (Fig. 7). Therefore, all compounds fulfilled the primary goal of improving the Oral Non CNS Score. In the process of improving the Oral Non CNS Score, modifications were made that surpassed the second goal of maintaining docking score — all docking scores are significantly better. Though D6 did not achieve a Lipinski’s Rule of 5 Score of 1, it did provide important insight into the structure-activity relationship of combretastatin analogs. Its StarDrop glowing molecule revealed that the four carbon linkers have a significant contribution to cytotoxicity and hERG score, and provided a reference for modification to the double bond in future phases. While many other compounds display improved docking score to tubulin, the compounds ranked lower than others in respect to the Oral Non CNS Score. Given that improving the Oral Non CNS Score was the primary goal of this phase, the relatively low scores of those compounds render them poor candidates for future development. Out of the remaining compounds in Phase II, D5 and D7 were the most promising candidates designed. D5 is the same as D2 but contains a fluorine substitution at the second carbon linker between the two aromatic rings (Fig. 9). D7 was one of the first compounds to utilize a phenone moiety — two carbon linkers instead of four, a modification that created noticeable decrease in cytotoxic activity (Fig. 11). 52 | 2021-2022 | Broad Street Scientific
4.3 Phase III Overview and Analysis Goals: Improve the Oral Non CNS Score, Optimize solubility. Of all the compounds created in Phase III, D8 ranked 2nd in terms of docking score and 3rd in terms of Oral Non CNS Score. D9 and D10 were other notable compounds, having slightly higher Oral Non CNS Scores. However, the docking scores of both compounds were significantly compromised, making them unideal candidates for synthesis (Fig. 10). D8 adopts the phenstatin structure of D7. A notable modification made is an added phenol at the 5th position on the B ring. Based on analysis of the Oral Non CNS and glowing molecules in StarDrop, it is proposed that the addition of the phenol increases the soluble molecular surface area contributing to the improved Oral Non CNS Score (Fig. 11). Furthermore, the 2,4,6- trimethoxy was modified to a 3,4,5-trimethoxy, which seemed to have larger surfaces of high contribution to soluble area. Therefore, based on the results of this phase, D8 was chosen as the most promising candidate for oral inhibition of tubulin (Fig. 12).
Figure 10: Binding affinity values, hERG scores, logS values, Lipinski’s Rule of 5 values, and Oral Non CNS Scores of the best compounds from all derivatives created of D8. The original ligand, D8, is highlighted in green. CHEMISTRY
Figure 11: StarDrop glowing molecules from Phase II and Phase III of solubility and Oral Non CNS Score. Blue = limited contribution. Green = moderate contribution Red = major contribution.
Figure 12: CA-4 vs. D8 docking scores and multi-parameter optimization scores. Green = improved score. 4.4 Phase IV Overview and Analysis Goals: Enumerate phase III’s most successful compound to generate a class of pharmaceutically relevant compounds similar in structure. Of all the compounds created in Phase III, D8 was the most improved when taking into account both Oral Non CNS Score and docking score. It should be noted that D8 has four additional hydrogen bonds in comparison to CA4, and is largely oriented so that the carbonyl group is pointed towards the α-tubulin T5 loop. Conformational changes of the T5 loop are known to interact with the assembly of tubulin heterodimers, indicating that the carbonyl group of the phenone moiety may be critical for its biological activity (Choi et al. 2013). The 3,5-bromophenol demonstrated good cytotoxicity, but when these substituents were moved to different positions, such as 2,6-, 2,4-, or 2,3-, these compounds became less potent, indicating that a steric factor is critical for the cytotoxicity of these derivatives. Therefore, a series of analogs containing the various substituted benzophenone and boronic acid addition were designed. The top nine compounds, taking into consideration docking score, Oral Non CNS Score, and Lipinski’s Rule of Five scoring, were selected to represent this novel class of compounds (Fig. 13). CHEMISTRY
Figure 13: Molecular structures of the most successful ligands created as derivatives of D8, exported from Schrodinger 2D Designer. Created using BioRender. 4.5 Comparison of D8 and CA-4 D8 displays significantly improved docking score values and H-bonding to tubulin compared with CA-4. In terms of solubility, Oral Non CNS Score, and pharmacokinetic properties, D8 is significantly more viable as a tubulin inhibitor. 4.5.1 Binding Comparison, Pharmaceutical Efficacy Comparison, Clinically Relevant Inhibitors D8 has a significantly better docking score to tubulin over CA-4. Maintaining an original H-bond with Asn337, new H-bond interactions with Ser340, Gln336, and Asn349 are seen, with a particularly potent bonding between the boronic acid and Gln336. Increased steric interactions also contribute to the significant improvement in docking score of D8 to tubulin over CA-4. It is proposed that D8’s significantly high solubility also indicates that this compound may be active in vivo without the substituent of a prodrug. This may be attributed to the presence of the boronic acid moiety, as well as the double carbon linker. Poor solubility was a setback in the clinical development of CA-4; thus, D8 shows significant increased promise in comparison. A comparison between chalnoids, fluorinated cis-restricted combretastatin analogs, and these benzophenone compounds suggests that the double carbon linker may create the most successful binding mechanism and have the most potency. D8 also has a higher Oral Non CNS Score over CA-4. Analysis of the glowing molecules from StarDrop suggests that the targeted modifications made to CA-4 and its subsequent derivatives are key to improving various criteria used to calculate the Oral Non CNS Score. It is seen that the methoxylated benzene ring, the phenone bridge, the Broad Street Scientific | 2021-2022 | 53
boronic acid, and the phenol on Ring B contribute greatly to the improved Oral Non CNS Score. In comparison to clinically relevant tubulin inhibitors, D8 displays improved docking scores to tubulin (Fig. 14). Vinblastine, Ombrabulin, and 2-Methoxyestradiol are tubulin inhibitors undergoing clinical trials as microtubule polymerization inhibitors: the improved binding of D8 to tubulin in comparison indicates that D8 has the potential to be a more effective microtubule depolymerization agent (Lu et al. 2012). D8 also has a higher Oral Non CNS Score than all existing compounds tested. Notably, D8 outscores ABT751, a drug achieving high oral bioavailability in clinical testing. Thus, D8 is a promising candidate for oral administration as it achieves a very competitive Oral Non CNS Score.
Figure 15: Fourier-transform infrared spectroscopy (FTIR) spectra of D8 (in blue) overlaid with the beginning compound, 3,5-dibromophenol (in orange). 5. Conclusion
Figure 14: Tubulin docking score and Oral Non CNS Score of D8 and clinically relevant tubulin inhibitors. 4.6 Synthesis Results FTIR spectra confirmed the successful synthesis of D8 based on the presence of peaks indicating formation of the substituents of the compound. The benzophenone (Ph-CO-Ph) structure was confirmed by the presence of the peaks at 1650 and 1280 cm-1. According to the peak at 3280 cm-1, the phenol on the phenylboronic acid is present. The boronic acid group on the B benzene ring is indicated by the peaks in the 1350-1370 cm-1 boron-oxygen stretch, and it is proposed that there was successful formation of boron-carbon bonds based on a peak at 1089 cm-1. There is a clear change in spectra from the initial 3,5-dibromophenol compound. Overall, there was a final yield of 60% (Fig. 15).
54 | 2021-2022 | Broad Street Scientific
This is the first known study that applies rational targeting modifications to design a novel class of tubulin inhibitors with potential for oral administration. A 4-phase design framework was employed in designing this inhibitor. Based on the results of Phase I and II, it was determined that the ortho-substituted boronic acid was most successful in improving docking score to both proteins by providing essential H-bonds or steric interactions, as well as increasing solubility. In Phase II, it was found that a 3,4,5-trimethoxy on Ring A and a benzophenone structure were most successful in improving the Oral Non CNS Score. In phase III, it was determined that the extension of the B benzene into a phenol improves the Oral Non CNS Score and contributes to an additional H-bond, improving the docking score value of D8 to tubulin. At the end of this targeted design process, D8 was determined as the most promising candidate for oral tubulin-colchicine complex inhibition, given its synthetic viability. In contrast to CA-4, D8 has higher docking score to tubulin, Oral Non CNS Score, HIA, and solubility, traits that are caveats in the clinical development of CA-4 and many analogs. In fact, D8 has better docking score values and Oral Non CNS Scores than multiple clinically relevant tubulin inhibitors. Not only that, but D8 can be easily synthesized using commercially available reactants in a few steps. D8 is part of a class of novel boronic acid containing compounds that have a phenone linker, all of which were successful in binding and Oral Non CNS Scores. D8 and the other compounds in its class are meaningful to the development of vascular disrupting and antiproliferative agents in the growing field of TCTs, as they have the potential to inhibit both tumor growth and cell proliferation simultaneously, working in combination with existing treatments or alone to treat cancer. Furthermore, it is proposed that combination of D8 as a vascular disrupting agent with existing antiangiogenesis agents may lead to complete tumor shutdown, leaving no risk of a peripheral CHEMISTRY
rim of viable cancer cells. The intended catalytic mechanism of D8 may reduce toxic cardiovascular damage associated with current vascular disrupting agents and chemotherapies, thus limiting secondary malignancies associated with or caused by cancer and its treatment methods. Having better solubility, a prodrug development in clinical trials is unnecessary, indicating that D8 will be more active and stable than most existing VDAs. In conclusion, D8’s potential for oral administration may lead to continued significant developments in the field of vascular disruption, as well as improve the practicality and convenience of cancer treatment for both medical professionals and patients. While D8 demonstrates significant improvement in Oral Non CNS Score and docking score, future work must be done to optimize and complete more analysis of hERG score and solubility, to confirm that a prodrug will not be necessary for clinical trials. Doing so is imperative to increasing the potential success of in vitro and in vivo experimentation. Future work will be focused on optimizing the synthesis scheme of D8, and working out viable synthetic schemes for the other 9 compounds in this class of boronic acid containing benzophenones. The efficacy of the dual inhibition effect should be tested using incubated human leukemia cells (HL-60) to evaluate the inhibitory effect on tumor cell growth in vitro. The vascular disrupting effect can be measured using human umbilical vein endothelial cells (HUVEC), and the compounds should be tested using tubulin polymerization assays to assess inhibition of polymerization. Finally, a microsomal stability test should be performed to confirm the compound is viable as an orally administered agent. If D8 demonstrates potential in these in vitro studies, further in vivo studies using mice is essential to evaluate the efficacy of D8 from a pharmacokinetic perspective.
small-molecule vascular disrupting agents.” Current opinion in investigational drugs (London, England: 2000), vol. 7, no. 6, 2006, pp. 522–528.
5. Acknowledgements
Hinnen, Petra and FALM Eskens. “Vascular disrupting agents in clinical development.” British journal of cancer, vol. 96, no. 8, 2007, pp. 1159–1165.
I would like to thank the NCSSM Foundation for their continued support, Dr. Tim Anglin for his constant guidance throughout this project, Dr. Darrell Spells and Dr. Michael Bruno for their insightful advice, Mr. Antonio Lopez for helping me locate chemicals, Mr. Bob Gotwals for assistance with software and computational development, and my wonderful Research in Chemistry peers for being there with me every step of the way. 6. References Alloatti, Domenico, et al. “Synthesis and biological activity of fluorinated combretastatin analogues.” Journal of medicinal chemistry, vol. 51, no. 9, 2008, pp. 2708–2721. Chaplin, David J, et al. “Current development status of CHEMISTRY
Choi, Min Jeong, et al. “Synthesis and biological evaluation of aryloxazole derivatives as antimitotic and vascular-disrupting agents for cancer therapy.” Journal of medicinal chemistry, vol. 56, no. 22, 2013, pp. 9008–9018. Dumontet, Charles and Mary Ann Jordan. “Microtubule-binding agents: a dynamic field of cancer therapeutics.” Nature reviews Drug discovery, vol. 9, no. 10, 2010, pp. 790–803. Engelbrechtsen, Line, et al. “Common variants in the hERG (KCNH2) voltage-gated potassium channel are associated with altered fasting and glucose-stimulated plasma incretin and glucagon responses.” BMC genetics, vol. 19, no. 1, 2018, pp. 1–9. Ghose, Arup K, et al. “Technically extended multiparameter optimization (TEMPO): an advanced robust scoring scheme to calculate central nervous system druggability and monitor lead optimization.” ACS chemical neuroscience, vol. 8, no. 1, 2017, pp. 147–154. GM Daenen, Laura, et al. “Vascular disrupting agents (VDAs) in anticancer therapy.” Current clinical pharmacology, vol. 5, no. 3, 2010, pp. 178–185. Group, US Cancer Statistics Working, et al. “United States cancer statistics: 1999–2012 incidence and mortality web-based report.” Atlanta (GA): Department of Health and Human Services, Centers for Disease Control and Prevention, and National Cancer Institute, 2015.
Jameson, MB, et al. “Clinical aspects of a phase I trial of 5, 6-dimethylxanthenone-4-acetic acid (DMXAA), a novel antivascular agent.” British journal of cancer, vol. 88, no. 12, 2003, pp. 1844–1850. Kalepu, Sandeep and Vijaykumar Nekkanti. “Insoluble drug delivery strategies: review of recent advances and business prospects.” Acta Pharmaceutica Sinica B, vol. 5, no. 5, 2015, pp. 442–453. Karatoprak, Gökçe Şeker, et al. “Combretastatins: an overview of structure, probable mechanisms of action and potential applications.” Molecules, vol. 25, no. 11, 2020, p. 2560. Broad Street Scientific | 2021-2022 | 55
Kwon, Young-Bum, et al. “Palladium-Catalyzed Synthesis of Aryl Ketones from Carboxylic Acids and Aryl- boronic Acids Using EEDQ.” Bulletin of the Korean Chemical Society, vol. 31, no. 9, 2010, pp. 2672– 2674.
Siemann, Dietmar W. “The unique characteristics of tumor vasculature and preclinical evidence for its selective disruption by tumor-vascular disrupting agents.” Cancer treatment reviews, vol. 37, no. 1, 2011, pp. 63–74.
Lu, Yan, et al. “An overview of tubulin inhibitors that interact with the colchicine binding site.” Pharmaceutical research, vol. 29, no. 11, 2012, pp. 2943–2971.
Silva, Mariana Pereira, et al. “Boronic acids and their derivatives in medicinal chemistry: synthesis and biological applications.” Molecules, vol. 25, no. 18, 2020, p. 4323.
Mfuh, Adelphe M, et al. “Scalable, metal-and additive-free, photoinduced borylation of haloarenes and quaternary arylammonium salts.” Journal of the American Chemical Society, vol. 138, no. 9, 2016, pp. 2985–2988.
Subramanian, Govindan and Douglas B Kitchen. “Computational approaches for modeling human intestinal absorption and permeability.” Journal of molecular modeling, vol. 12, no. 5, 2006, pp. 577–589.
Mita, Monica M, et al. “Vascular-disrupting agents in oncology.” Expert opinion on investigational drugs, vol. 22, no. 3, 2013, pp. 317–328.
Sung, Hyuna, et al. “Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality world- wide for 36 cancers in 185 countries.” CA: a cancer journal for clinicians, vol. 71, no. 3, 2021, pp. 209–249.
Nguyen, Tam Luong, et al. “A common pharmacophore for a diverse set of colchicine site inhibitors using a structure-based approach.” Journal of medicinal chemistry, vol. 48, no. 19, 2005, pp. 6107–6116. Obrezanova, Olga and Matthew D Segall. “Automated QSAR modeling to guide drug design.” 237th National Meeting of the American Chemical Society. Salt Lake City, UT, USA, 2009, pp. 22–26. Pettit, GR, et al. “Isolation and structure of the strong cell growth and tubulin inhibitor combretastatin A-4.” Experientia, vol. 45, no. 2, 1989, pp. 209–211.
Tan, Lu and Johannes Kirchmair. “Software for metabolism prediction.” Drug Metabolism Prediction, 2014, pp. 27–52. Tron, Gian Cesare, et al. “Medicinal chemistry of combretastatin A4: present and future directions.” Journal of medicinal chemistry, vol. 49, no. 11, 2006, pp. 3033–3044. Wang, Yuxi, et al. “Structures of a diverse set of colchicine binding site inhibitors in complex with tubulin provide a rationale for drug discovery.” The FEBS journal, vol. 283, no. 1, 2016, pp. 102–111.
Schaaf, Marco B, et al. “Defining the role of the tumor vasculature in antitumor immunity and immunother- apy.” Cell death & disease, vol. 9, no. 2, 2018, pp. 1–14. Schobert, Rainer, et al. “Pt (II) complexes of a combretastatin A-4 analogous chalcone: effects of conjugation on cytotoxicity, tumor specificity, and long-term tumor growth suppression.” Journal of medicinal chemistry, vol. 52, no. 2, 2009, pp. 241–246. Schrödinger, Maestro. 9.3, User Manual. 2012. Segaoula, Zacharie, et al. “Synthesis and biological evaluation of N-[2-(4-Hydroxyphenylamino)-pyridin-3yl]-4-methoxy-benzenesulfonamide (ABT-751) tricyclic analogues as antimitotic and Antivascular agents with potent in vivo antitumor activity.” Journal of medicinal chemistry, vol. 59, no. 18, 2016, pp. 8422–8440.
56 | 2021-2022 | Broad Street Scientific
CHEMISTRY
DEVELOPMENT OF A BIOACTIVE, BIODEGRADABLE, AND VARIABLE-DENSITY 3D PRINTER FILAMENT FOR PATIENT-SPECIFIC BONE RECONSTRUCTIVE IMPLANTS Jacob Rose Abstract Current orthopedic implants are restricted in the level of patient specificity and biological interaction achievable. The demand for more effective and less intrusive reconstructive orthopedic implants has been increasing with an aging population. This project aims to develop a bioactive, biodegradable and variable density 3D printer filament for use in patient-specific bone implants. Using Polylactic-acid (PLA) as a biodegradable structural polymer, the bioactive agent hydroxyapatite (HA) to increase osteoblast integration and a chemical-foaming-agent (CFA) as a temperature-sensitive foaming agent, a 3D printer filament was composed of their properties. This filament was assayed to determine how physical properties of 3D printed parts change based on manufacturing temperature. It successfully demonstrated controllable density, modulus of elasticity and ultimate tensile stress when printing at temperatures chosen based on the decomposition temperature of the CFA. The ability to tune these properties on a patient-specific scale can greatly improve the efficacy of orthopedic implants as surrounding bone tissue will more readily accept material that matches its properties. 1. Introduction Thermoplastics such as polyether ether ketone (PEEK) and polylactic acid (PLA) have shown substantial benefits when used in medical applications. PEEK has regular use in orthopedic implants for bone replacement [5], and PLA has shown promise in applications such as cardiovascular implants, orthopedic interventions and medical equipment [2]. A particular interest for PLA is in use for regenerative medicine as a biodegradable and fully bioabsorbable material. PLA has been shown to last for time periods of approximately 1 year in the body before complete biosorption [2]. Due to this property, it has been investigated for use in bone regenerative scaffolding and found to have promising physical properties for orthopedic implants [2]. Currently, the majority of implants are non-biodegradable permanent fixtures. This leads to a necessity for either a secondary surgery to remove the scaffold or to design parts that are the structural component of bone for the rest of the patient’s life. By developing a biodegradable scaffolding, implants will be able to be resorbed by the patient’s body after bone growth and replacement [6]. PEEK has been used extensively in bone implants, as plastics such as PEEK closely resemble mechanical properties of bone [7]. Problems with conventional metal implants include stress shielding, bone resorption, and poor osseointegration caused by a mismatch in the modulus of elasticity between the implant and surrounding bone tissue [1]. Designing polymers with mechanical propENGINEERING
erties that adhere as closely to the natural properties of bone as possible reduces negative impacts on surrounding bone tissue [1]. This can be a challenge however, as these mechanical properties, such as the modulus of elasticity, are highly patient-specific. Not only does modulus vary based on type of bone and location in the body, but also age and osteoporotic bone structures [3][4]. Other research has shown that the modulus of elasticity also depends on bone density [8]. Therefore, the ability to adjust both density and modulus of elasticity of an implant would greatly improve the implant’s efficacy and reduce the negative impact on surrounding tissue. While these plastics can resemble the mechanical properties of bone, they do not chemically interact with the body. Most metal and plastic implants are biologically inert, which leads to poor binding of bone. One advancement in implant technology is the addition of a hydroxyapatite (HA) coating. HA is a major component of normal bone that when used as a coating enhances osteoblastic cell adhesion, growth and differentiation [12]. Integrating HA into the polymer could enhance bone growth by providing HA binding throughout the implant. HA not only improves biological interactions with surrounding tissue, but also has an effect on the modulus of elasticity and degradation behavior [9][10]. Therefore, by adjusting the proportion of HA in a composite material, the material properties may be tuned to the application. An additional issue with current implants is the lack of patient specificity. Most current implants are manufactured with injection molding. Due to requiring a Broad Street Scientific | 2021-2022 | 57
pre-manufactured mold, these parts are limited in the specificity with which they can replicate patient bone loss. 3D printing is of recent interest because of its ability to produce non-standardized components from computer-designed models. Currently, most 3D printed bone implants are manufactured using a titanium sintering process. This process is expensive and time-intensive and requires parts to be sourced from an outside manufacturer. Fused-Deposition-Modeling (FDM) 3D printing greatly improves upon these issues. Requiring relatively cheap and easy to run machinery, FDM 3D printing could bring production of implants to hospitals, reducing both the time and cost of manufacturing implants. Additionally, FDM 3D printing allows much more specificity in replicating a patient’s bone over conventional methods of implant manufacturing such as injection molding. Injection molding is not able to vary properties such as density, surface structure, and modulus of elasticity on a region-by-region basis. By integrating a temperature-sensitive agent, the 3D printer is able to adjust these properties by adjusting printing temperature. This is a simple and quick adjustment that can be completed automatically during printing. Using this method of varying density, implants could be manufactured with different material properties in specific regions, so implants created using this method could closely replicate patient-specific bone. More accurate replicates of bone as regenerative scaffolding would improve the interaction and regeneration of a patient’s bone with the implant. This project aims to incorporate chemical foaming agents into printable plastics to develop a bioactive, biodegradable, and variable density 3D printer filament for use in patient-specific bone implants that reflect patients’ individual bone physical properties. Materials and Methods Preparation of polylactic acid polymer pellets containing hydroxyapatite powder and chemical foaming agent. Hydroxyapatite powder (HA) was obtained from Sigma Aldrich (Product Number 04238) for use as the bioactive agent. Citric Acid (CA) was obtained for use as a chemical foaming agent (CFA). In preparation for the dissolution of polylactic acid polymer, ~16 g total of HA and CFA in varying weight-to-weight ratios were continuously stirred with a magnetic stir bar in 200 mL of dichloromethane (DCM) for 10 minutes to produce a fine slurry. 60 grams of commercial grade polylactic acid (PLA) polymer produced by 3DXTech Additive Manufacturing (PPLA15000) were added to the slurry and stirred with a magnetic stir bar for 10 minutes at 270 rpm. After stirring, the PLA was left to dissolve overnight. DCM (50 mL) was added to the beakers and the solutions were stirred for 50 minutes. Samples were then poured into flat-sheet 58 | 2021-2022 | Broad Street Scientific
molds and air-dried for 24 hours. After air drying, plastic sheets were shredded to pellets with dimensions of 7mm or less for subsequent processing. Polymer pellets were then dried in a vacuum oven at 13.51 psi vacuum and 75° C for 4 hours to fully remove solvent. Extrusion of polymer pellets into filament strand Polymer pellets were extruded using a single-screw and single-temperature filament extruder from Filastruder (V2.0) at 150° C and 6 rpm. This temperature was chosen to be above the melting temperature of pure PLA and below the decomposition temperature of the CFA. The rotational speed was chosen to produce a constant diameter of filament. The extruded filament was wound onto a spool via an automatic spool winder to maintain constant tension on the filament. Tensile strength testing Samples were printed for tensile testing at temperatures from 182° C to 220° C at increments of 10° C to test the effect of printing temperature. Feed rate was adjusted based on the foaming characteristics to achieve a constant extrusion width. Samples were printed in a dogbone shape to measure breakage in a known cross-sectional area. Stress and strain data were collected using a Vernier Structure and Materials Tester through tensile testing samples until failure. Stress-strain curves were determined from stress and strain data determined by the equations σ = F/A₀ and ϵ = ΔL/L₀. Modulus of elasticity was determined from the slope of the linear portion of stress-strain curves. Ultimate strength measurement Ultimate strength was determined from the stressstrain curves measured in tensile strength testing. This is the maximum stress held by the sample. Density measurement of extruded polymer Density measurements were performed on the printed dog bone tensile test samples to gather density information. Volume was determined from computer-aided-design (CAD) predicted volume and validated with measurements expected from CAD model. Mass was measured on an analytical balance. Results Using the method described above, filament samples were prepared at printing temperatures from 182° C to 220° C. Serial numbers were assigned based on filament type and printing conditions as follows: Filament type CA (Citric Acid - type of foaming agent) 795 (percentage by weight of PLA - 79.5%) -200 (percentage by weight HA ENGINEERING
- 20.0%) -05 (percentage by weight CFA - .5%) :5 (generation of filament sample). Printing conditions: 182 (degrees C - printing temperature) - 130 (percentage flow rate) :1 (trial number for specific filament at these conditions). 3D printed dog bone tensile testing shapes were assayed using the methods described above to gather information on modulus of elasticity, ultimate tensile strength, and density. Samples are pictured in figures 1 and 2.
Figure 3: Raw data panel Figure 1: Picture of fractured tensile test samples of CA 795-200-05:5
Figure 2: Picture of fractured tensile test samples of CA 785-200-15:5 The relationship between printing temperature and modulus of elasticity, ultimate tensile stress, and density was graphed, as shown in figure 3. Comparing the .5% CFA filament with the 1.5% CFA filament demonstrates a reduction in printing temperature required to change physical properties due to CFA foaming.
ENGINEERING
In modulus of elasticity, graph A (.5% by mass CFA) has a peak modulus at 200° C, compared to a peak modulus in graph B (1.5% by mass CFA) at 190°C. The modulus of elasticity is then reduced by 25% in graph A and 33% in graph B. This relationship between concentrations of CFA is shown in density as well. In graphs E and F, a similar peak was observed at 200° C for the .5% CFA sample and at 190° C for the 1.5% CFA sample. The density then reduces by 9.7% and 14.8%, respectively. A similar relationship exists between printing temperature and ultimate tensile stress, as shown in graphs C and D in figure 3. While the peak ultimate tensile stress exists for both graphs at 190° C, graph D (1.5% by mass CFA) shows a greater decrease in stress as compared to graph C (.5% by mass CFA). In graph D, there is a 26.2% reduction in ultimate tensile stress observed. This is much greater than in graph C, where there is a 7.0% decrease in ultimate tensile stress. Relationships between density and both modulus of elasticity and ultimate tensile stress can be drawn as well. As shown in figure 4, positive linear relationships between density and both modulus of elasticity and ultimate tensile stress were demonstrated.
Broad Street Scientific | 2021-2022 | 59
C. Comparing these two images, a clear view of the foaming action can be observed. This foaming action also has the effect of changing surface structure. This provides a surface structure similar to trabecular bone as shown in figure 7.
Figure 4: Modulus of Elasticity and Ultimate Tensile Stress vs Density Conclusions and Future Directions This project demonstrates control of modulus of elasticity and ultimate tensile stress with change in printing temperature. Using filaments produced in this study, implants may be created that are tuned to the patient’s specific bone properties. This is especially impactful in patients with osteoporosis. Due to the low modulus of elasticity and high variability of modulus of elasticity of osteoporotic bone[4], current implants cannot match their modulus of elasticity to the modulus of osteoporotic bone. In a study determining the impact of osteoporotic bone structures of the pelvic-hip complex on stress distribution, modulus of elasticity was shown to vary from about 10,707 MPa in healthy bone to 4,701 MPa in osteoporotic bone[4]. In this study, PLA-HA-CFA composite filament was demonstrated to have a modulus of elasticity that varied between 7,046 MPa and 4,810 MPa. With this filament, implants may be specifically tuned during printing to match the modulus of elasticity of the patient’s bone. In addition to tuning modulus of elasticity and ultimate tensile stress, density was shown to have a similar relationship to printing temperature. Shown in graphs C and D in figure 3, as the CFA reaches decomposition temperature and begins to foam, there’s a reduction in the density of the samples. This foaming behavior is also shown in figures 5 and 6. Figure 6 depicts a single layer of the 1.5% CFA by mass filament printed at 182° C. Figure 5 depicts a single layer of the same filament printed at 220° 60 | 2021-2022 | Broad Street Scientific
Figure 5: Picture of a .2 mm layer of CA 785-200-15:5 printed at 220° C
Figure 6: Picture of a .2 mm layer of CA 795-200-15:5 printed at 182° C
ENGINEERING
terial properties may be produced in very specific regions in a print. A diagram of a 3D printing nozzle is shown in figure 9, also depicting how this temperature-sensitive foaming filament reacts to the high temperatures in production of an implant. This behavior is further shown in figures 10 and 11, as it depicts a section of filament removed from the nozzle during printing and allowed to cool. These figures show how, as the filament is melted and eventually reaches the decomposition temperature of the CFA, the filament foams.
Figure 7: SEM Image of trabecular bone [11] As increasing printing temperature past CFA decomposition temperature is shown to improve surface structure while decreasing ultimate tensile stress, implants can be created with a balance of each property. Varying printing temperature in each region of an implant allows an implant to either directly reproduce the placement of cortical and trabecular bone as in the patient’s original bone or create an entirely new structure designed for high-stress or low-stress environments as shown in figure 8.
Figure 9: Diagram of foaming process during usage
Figure 10: Picture of partially-foamed 3d-printing filament composite
Figure 8: Diagram of final part manufactured with varying temperature
Figure 11: Picture of fully foamed portion of composite
Using this filament in 3D printing applications opens further avenues for patient-specificity. As depicted in figure 8, different regions of a 3D print may have varied physical properties. As FDM 3D printers can automatically and quickly change nozzle temperature, different ma-
Controlling each of these properties in a region-specific manner allows final implants to be directly tuned to the needs of a patient. Additionally, as this technique uses FDM 3D printing, manufacturing of implants can happen on-site, within hours and much cheaper than conven-
ENGINEERING
Broad Street Scientific | 2021-2022 | 61
tional methods. Combining each of these benefits into one reconstructive implant changes the way that bone replacements (especially for traumatic injuries when pre-manufacturing implants is impossible) will be done in the future. To implement this composite filament, testing of compressive strength will need to be performed. Additionally, testing on the decomposition rate of this composite and how this rate influences cell integration will need to be performed to ensure stability of the implant throughout its life cycle.
6. Dana da Silva, Maya Kaduri, Maria Poley, Omer Adir, Nitzan Krinsky, Janna Shainsky-Roitman, Avi Schroeder. Biocompatibility, biodegradation and excretion of polylactic acid (PLA) in medical implants and theranostic systems. Chemical Engineering Journal, Volume 340, 2018, Pages 9-14, ISSN 1385-8947, https://doi.org/10.1016/j. cej.2018.01.010.
Acknowledgments
7. Heary, R. F., Parvathreddy, N., Sampath, S., & Agarwal, N. (2017). Elastic modulus in the selection of interbody implants. Journal of spine surgery (Hong Kong), 3(2), 163–167. https://doi.org/10.21037/jss.2017.05.01.
I would like to thank Dr. Timothy Anglin, Dr. Michael Bruno, the NCSSM Foundation, the NCSSM Science Department, the NCSSM Summer Research and Innovation Program, Dr. Sarah Shoemaker, and my Research in Chemistry peers for making this research possible.
8. Morgan, E. F., Unnikrisnan, G. U., & Hussein, A. I. (2018). Bone Mechanical Properties in Healthy and Diseased States. Annual review of biomedical engineering, 20, 119–143. https://doi.org/10.1146/annurev-bioeng-062117-121139.
Works Cited
9. Po-Liang Lin, Hsu-Wei Fang, Tiffany Tseng, Wun-Hsing Lee. Effects of hydroxyapatite dosage on mechanical and biological behaviors of polylactic acid composite materials. Materials Letters, Volume 61, Issues 14–15, 2007, Pages 3009-3013, ISSN 0167-577X, https://doi.org/10.1016/j. matlet.2006.10.064.
1. Shi, L., Shi, L., Wang, L., Duan, Y., Lei, W., Wang, Z., Li, J., Fan, X., Li, X., Li, S., & Guo, Z. (2013). The improved biological performance of a novel low elastic modulus implant. PloS one, 8(2), e55015. https://doi.org/10.1371/journal.pone.0055015 2. Vincent DeStefano, Salaar Khan, Alonzo Tabada. Applications of PLA in modern medicine. Engineered Regeneration, Volume 1, 2020, Pages 76-87, ISSN 2666-1381, https://doi.org/10.1016/j.engreg.2020.08.002. 3. Valdes, Rogelio & Solis, A.L. & Godínez, Francisco & Martínez, E. & Villegas, C.H. & Navarrete, Margarita. (2010). Evaluation of Modulus of Elasticity, Mineral Composition and Bone Mineral Density of Trabecular Bone L3- Vertebrae Samples Extracted From Mexican Men. Materials Research Society Symposium Proceedings. 1242. 133-138. 10.1557/PROC-1242-S4-P81. 4. Arkusz, Katarzyna & Klekiel, Tomasz & Niezgoda, Tadeusz & Bedzinski, Romuald. (2018). The influence of osteoporotic bone structures of the pelvic-hip complex on stress distribution under impact load. Acta of bioengineering and biomechanics / Wroclaw University of Technology. 20. 10.5277/ABB-00882-2017-02.
10. Russias, J., Saiz, E., Nalla, R. K., Gryn, K., Ritchie, R. O., & Tomsia, A. P. (2006). Fabrication and mechanical properties of PLA/HA composites: A study of in vitro degradation. Materials science & engineering. C, Biomimetic and supramolecular systems, 26(8), 1289–1295. https://doi. org/10.1016/j.msec.2005.08.004. 11. Whitehouse, W. J., Dyson, E. D., & Jackson, C. K. (1971). The scanning electron microscope in studies of trabecular bone from a human vertebral body. Journal of anatomy, 108(Pt 3), 481–496. 12. Hasegawa, S., Tamura, J., Neo, M., Goto, K., Shikinami, Y., Saito, M., Kita, M., & Nakamura, T. (2005). In vivo evaluation of a porous hydroxyapatite/poly-DL-lactide composite for use as a bone substitute. Journal of biomedical materials research. Part A, 75(3), 567–579. https://doi. org/10.1002/jbm.a.30460.
5. Honigmann P, Sharma N, Okolo B, Popp U, Msallem B, Thieringer FM. Patient-Specific Surgical Implants Made of 3D Printed PEEK: Material, Technology, and Scope of Surgical Application. Biomed Res Int. 2018 Mar 19;2018:4520636. doi: 10.1155/2018/4520636. PMID: 29713642; PMCID: PMC5884234. 62 | 2021-2022 | Broad Street Scientific
ENGINEERING
ZIEGLER-NICHOLS TUNING IMPLEMENTATION ON ARDUINO-BASED PID CONTROLLER FOR DC MOTOR ROTATION Pracheeti Shikarkhane Abstract Control systems are the core regulatory mechanisms in many engineered devices. One widely implemented industrial control system is the Proportional-Integral-Derivative (PID) controller, known for its efficient tuning outputs, situational versatility, and relative ease in logic. Nevertheless, despite its prevalence, the PID framework remains mainly inaccessible to secondary school students and early undergraduates. This research study, aimed at expanding the educational reach of control systems, presents a control experiment implementing the PID algorithm on a DC motor for fine-tuned speed regulation. The system is designed using the Arduino microcontroller to lower cost barriers for replication. Upon specifying the target speed, a photoelectric sensor is used to enable closed-loop feedback control, in which the system receives and adjusts the rate of motor rotation based on calculated steady-state error. The algorithm is developed to reduce sinusoidal oscillation at the setpoint and minimize overshoot. The Ziegler-Nichols (ZN) tuning experiment is implemented on the algorithm to predict gain parameters, after which it is observed that the motor self-corrects with reduced settling time and thus increased efficiency. The derived gain parameters lead to minimized overshoot and desirable logarithmic settling trajectories as the setpoint is reached. The success of this two-part implementation demonstrates feasibility for classroom applications. 1. Introduction The rapid advances in technology propagating over the last several decades have subsequently propelled developments in autonomous robotics and intelligent systems. Robotics, often seen as the intersection of electrical and mechanical engineering, has found permeating applications in an appreciable array of fields, ranging from space and underwater exploration to health care. NASA’s Mars rover Perseverance [7] is a prime paradigm of industrial robotics, as is the Da Vinci robot intended to aid surgeons in performing complex surgeries of the neck, heart, and head [8]. But regardless of the field in which robotics is applied, every autonomous robot is driven by an intelligent control system. One of the most extensively applied such control system is the Proportional-Integral-Derivative (PID) controller, which has found increasing applications in industry-grade robotics as well as upper undergraduate education. The PID controller is based on closed-loop feedback control theory, an iterative computational method in which the output of a process is fed back as the input. When implemented, this corrective method allows for intelligent control and increasingly accurate results compared with an open-loop control system. While a primitive PID controller may serve the needs of a minimally taxing robotic system, more resource-intensive systems require greater control autonomy. The efficiency of the PID controller, as defined by the time it takes to produce the desired output and the proximity of the output to its intended value, varies with the scope of the algorithm ENGINEERING
used to implement and optimize it. In the past, classical techniques such as the ZN-tuning method and Cohen Coon methods have been used for PID controller optimization, as well as more computationally intensive optimization techniques such as artificial neural networks (ANN), support vector machines (SVN), particle swarm optimization (PSO), fuzzy logic control (FLC), etc. [4] While these methods have been used in industrial robotics for applications such as trajectory tracking [3] and stability control, [6] the control theory framework remains inaccessible to a younger audience, specifically high school robotics teams and early undergraduate students. As such, this research study aims to provide accessible cyber-physical experimentation of control systems and their applications to autonomous robotics. To this end, the study investigates how PID control may be implemented on a low-cost Arduino microcontroller to regulate DC motor rotation. A three-part experiment was performed to achieve this: a physical system was first constructed, after which the PID control was programmed, followed by an implementation of the Ziegler-Nichols tuning method. 2. Control Theory Control theory addresses the behavior of dynamic systems in which the designed controller manipulates one or more variables to achieve the desired reference output. Digital processing platforms and analog IO devices often perform the intensive computations necessary for controller design. Therefore, the framework is commonBroad Street Scientific | 2021-2022 | 63
ly applied in mechanical, electrical, and aerospace engineering settings. The importance of control systems can be depicted through a simple example considering a fan and knob system; when an analog knob is set so that a fan is expected to turn at 450 revolutions per minute (RPM), the observer is unable to determine the precision of speed. In other words, no quantitative method can be applied to determine whether the fan is turning at the desired 450 RPM or at an arbitrary 500 RPM. While the precision in fan rotation may not be of exceeding importance, precise control becomes of utmost importance when used for systems such as guided military missiles, satellites, or surgery robotics. Consequently, using an intelligent control method is necessary to guarantee the accuracy of electrical systems, ensuring that they perform as desired. An additional benefit of employing control systems is reducing dependency on a power source. While systems without control algorithms are often limited by the efficiency and stability of their power source, systems employing intelligent control methods demonstrate autonomous self-correction regardless of a depleting or fluctuating battery voltage [17]. It can also be noted that systems with intelligent control methods can correct the setpoint even after experiencing an external disturbance. These factors have led control theory to become extensively practiced in the industrial setting.
eters to the function, also referred to as the gain values. The output of the proportional controller, Kp, is directly based on the error of the system, where the error is defined as the difference between the reference point and current output. The output of the integral controller, Ki, is based on the accumulated error in the system. The derivative controller, Kd, measures the rate of change of error, i.e., the slope of the error. The function can also be expressed where Ki and Kd are replaced by Kp/Ti and KpTd respectively. Ti and Td represent integration time and differentiation time respectively [12].
2.1 PID Control Since its conception in the early 20th century, the PID control method has advanced significantly in terms of complexity and efficiency. The control algorithm was first conceived for steam engines and ship steering, quickly adopting widespread application when wideband highgain amplifiers were developed. The control method has maintained a strong industrial presence with applications across medicine and aerospace for trajectory tracking purposes [1] and subsystem automation. The PID control method is also frequently applied in non-industrial projects. The method has been used to teach control systems in the classroom at the graduate level. It has also been used broadly for hobby projects as well as in some cases of competitive high school robotics [5]. Evidently, the control method’s versatile applications render it a powerful tool for controlling dynamic systems. The PID controller is a three-term controller as the name suggests. It includes proportional, integral, and derivative components. The control function u(t) can be written as follows [2].
Figure 1. Feedback control loop [27]
2.2 Closed-Loop Control PID control is a closed-loop feedback system, meaning that after each iteration, its output y is fed back as an input to the controller with the use of a sensor. The reference input r is used to calculate the error e, where e = r −y. C(s) is the controller, in this case the PID controller, and G(s) is the system, commonly referred to as the plant (Fig. 1) [16].
From this, the transfer function between the input r and output y can be calculated when it is assumed that the system is linear [16].
Knowledge of the transfer function is useful because it enables the implementation of more mathematically intensive control methods such as the Root Locus method [22]. However, in most cases the plant G(s) is unknown. In such cases, PID control presents itself as a suitable control method since it does not require a transfer function for the plant. Instead, the parameters of the plant are identified through closed-loop experimentation that can be done either explicitly or implicitly. Implicit identification of the system’s parameters can be done using ZN-tuning, which proposes PID controller parameter values optimized for efficient correction to a setpoint. The PID controller can be graphically expressed in closed-loop control as seen in Fig. 2.
Kp, Ki, and Kd denote the non-negative constant param64 | 2021-2022 | Broad Street Scientific
ENGINEERING
Figure 2. PID controller feedback control loop. The Process block can be considered the plant G(s) [25] The proportional gain is multiplied to the error term to determine the proportional response. In common control practice, it has been observed that increasing the proportional gain will lead to a faster control system response. However, if the proportional gain Kp is increased beyond a system-specific threshold, the system will display oscillating behavior, never reaching the setpoint. It is possible that this oscillating behavior will extend out of control into chaotic instability. The point at which the oscillations are controlled and consistent is referred to as neutral instability. During autonomous corrective control, the integral gain often fluctuates, until the reference point is reached. Fluctuations are determined by whether the iterative error is positive or negative, considering that the parameter represents the cumulative error. The integral gain works to drive the system’s steady-state error to zero. Steadystate error is defined as the final difference between the input and the output as time approaches infinity, ef = r − yf. The derivative gain works to bring the rate of change of error to zero [20]. It dampens the system response by reducing applied corrective force and thus reducing overshoot. Ultimately it aims to mitigate oscillations around the setpoint [24]. A controller need not necessarily contain all three gain parameters, proportional, integral, and derivative. For example, controllers consisting of P, PI, and PD can be constructed and compared, which is done in this study. It is important to note that while singular proportional and integral controllers can be constructed, it is not possible to construct a singular derivative controller, since it does not directly consider error [24]. It is predicted that a combination of all three components will produce the most efficient system optimization response. What presents a challenge is determining the gain values that will optimize the system. While most rudimentary implementations of PID controllers tend to choose near arbitrary values for these gains, an intelligent system response can be achieved with implicit identification of these parameters [18]. 2.3 ZN-Tuning The Ziegler-Nichols tuning method is an implicit opENGINEERING
timization technique for determining the gain values in a PID controller. This method is considered to be among the most popular as a result of its relative mathematical simplicity. The Ziegler-Nichols tuning method has two variations. The first variation begins with the application of a unit step input to the system. The response is graphed, upon which a tangent line is drawn at the point of inflection. The delay value L is measured as the distance from the x-intercept to the line’s stability point K (Fig. 3). The derived L and T values are used as determinants for the resulting gain [10].
Figure 3. ZN-tuning method variation I [26] The second variation, which was implemented in this study, includes an initial oscillation experiment and subsequent instability analysis. To achieve oscillation, the integral and derivative gains are set to 0, while the proportional gain is increased until consistent oscillations are visible i.e., neutral instability has been reached. The resulting proportional gain value (Ku) and period of oscillation (Tu) are measured and used to determine the gain values in accordance with the broadly accepted presets (Fig. 4). It is important to note that while ZN-tuning predicts the optimal gain parameters for an efficient system response, further tuning and testing are often conducted for more resource-intensive systems.
Figure 4. ZN-tuning variation II presets [9] 3. Methods A physical system was first designed to model applications for industrial robotics. The system was intentionally designed to focus solely on regulating motor rotation to mimic the effects of a control algorithm on one specific robotic subsystem. This implementation was further Broad Street Scientific | 2021-2022 | 65
designed to demonstrate a fundamental framework for complex systems. After designing and constructing the system, the PID control algorithm was programmatically implemented to determine if the motor would self-correct and reach the desired reference point based on closed-loop feedback. The performance of the PID system was used as a base comparison to quantify the system’s performance when a tuning algorithm was applied. The ZN-tuning algorithm was employed to predict the optimal gain parameters for the system and compare various controller designs. Five controller designs were implemented and tested for autonomous correction efficiency from the predicted gain parameters. 3.1 Construction of the System Increasingly, the Arduino microcontroller has proved an invaluable resource for hobbyists, students, and industry-grade electronics. The Arduino company offers a wide selection of boards, kits, and resources for its customers who have found endless applications for the palm-sized hardware unit [11]. The Arduino was deemed a suitable microcontroller for this study due to its accessibility and relatively low cost. A circuit was constructed using the Arduino Uno, which provided a 5V battery supply. An apertured wheel was attached to a DC motor, and a photoelectric sensor was placed beneath it to record motor speed. While the wheel was spinning, the photoelectric sensor measured the change in light between the solid and hollow surfaces in order to determine motor frequency on a scale of 120630. The sensor returned the measured frequency to the system, serving as the feedback mechanism. As shown in Figure 5, an H-Bridge chip was used to regulate the motor voltage and to catalyze a change in rotational direction if required. Two switches were used to control direction of rotation and on/off state. A 9V battery was attached to power the DC motor. Four 10K potentiometers were added to allow for initially controlling gain parameters manually. A mount and platform were created for the custom circuit (Fig. 6). The code was written in C++ using the Arduino IDE and then run on the Arduino Uno.
Figure 5. Custom circuit schematic 66 | 2021-2022 | Broad Street Scientific
Figure 6. System overview. The circuit-sensor mechanism can be seen in the bottom left, the Arduino can be seen in the top right, the potentiometers are visible on the rightmost breadboard, and the H-bridge mechanism is visible in the leftmost breadbaord. 3.2. PID Implementation A control algorithm was implemented on the system by creating a PID function within the code. The error value was computed after every iteration based on the output from the photoelectric sensor. The product of the error value and Kp constant was used to determine the proportional output. The integral output was determined from an integration of error multiplied by the Ki constant. Similarly, the derivative output computed the slope of the error, and the corrective value returned a summation of the three terms. This value was added to the RPM sent to the motor in each iteration [18][19]. 3.3. ZN Tuning Implementation A calibration experiment was first conducted in order to determine the relation between the input and output of the DC motor. For certain RPMs 55-255 (motor would not run below an RPM of 55), the resulting frequency reading from the photoelectric sensor was recorded on a scale of 120-630. The data were plotted and the equation for the line of best fit was used to convert the photoelectric sensor’s readings to a range of 0-255. An oscillation experiment was then conducted; out of the three constants, Ki and Kd were set to 0. Kp was manually increased until the system displayed a point of neutral instability, wherein oscillations had a consistent amplitude and period. The Kp value that produced neutral instability was regarded as Kmax. The resulting period, Tmax, was determined from the oscillations graph. Kmax and Tmax were used to find the optimal values for Kp, Ki, and Kd from Figure 4. The behavior of the system for each of the five controller designs, P, PI, PD, PID, and No Overshoot, was graphed, and the settling time was determined. 3.4. Autonomous Robotics Application The sensor-circuit mechanism [6] was designed to model a subsystem in a robotic device. To test applications for a multi-subsystem device, the control algorithm ENGINEERING
was implemented on an autonomous robot. The TI RSLK Max device was chosen for implementation, and the Pocket Beagle microcontroller, similar to the Arduino, was used. The PID control system was programmed and subsequently run on the microcontroller connected to the autonomous car to control motor speed and implement path following. A straight-line test was conducted to determine if the control algorithm could regulate and sync the motor speeds of both wheels in order to enable straight line travel (Fig. 7). Furthermore, the PID algorithm and an infrared sensor were implemented to allow the device to follow a black line (Fig. 8). This application demonstrates how the PID algorithm is feasible and successful for not only individual subsystems, but also for complex devices that involve the regulation of multiple motors and sensors.
Figure 7. PID controller implemented on both motor wheels to test straight line travel in reference to a ruler
Figure 8. Control algorithm and infrared sensor to enable line following 4. Results and Discussion The initial calibration experiment revealed that the system was linear (Fig. 9). The linear nature of the system enabled the derivation of a formula to convert photoelectric sensor frequencies into motor RPM. Conversion into motor RPM units was necessary for an accurate comparison of whether the reference RPM value had been reached as time approached infinity. The conversion formula was derived as follows:
Figure 9. Linear system visible through the calibration experiment The implementation of a programmatic PID algorithm revealed that the system was able to self-correct to the desired setpoint. For example, when given a target reference of 180 RPM, the motor was observed to initially overshoot the target RPM, but slowly reduce speed until the RPM was reached. However, there remained several deficiencies in the self-correcting process. Graphing the trajectory of measured RPM revealed that the motor significantly fluctuated speeds throughout the corrective process. This was likely because the input gain parameters were not tailored for the individual system, instead arbitrary values were initially chosen. Furthermore, the steady-state error did not reach a value close to zero, suggesting that the non-tailored gain values could not create an efficient system. This proved that while the PID algorithm allows for self-correction and control of electronic devices, the method will likely benefit from a tuning algorithm to determine optimal gain parameters. To predict these gain values, the Ziegler-Nichols (ZN) tuning method was implemented. The oscillation experiment conducted as a part of the tuning method (Fig. 10) revealed that the Kmax value for which the system reached neutral instability was 2.5, and the period was 0.6 seconds. This configuration was determined after testing Kp values ranging from 1-3.
Figure 10. Neutral instability reached when Kp = 2.5 Five controllers were tested to determine the most optimal design for this system. It is important to note that the gain parameters and controller design that produce
ENGINEERING
Broad Street Scientific | 2021-2022 | 67
the most efficient response vary from system to system. Using the ZN method presets (Fig. 4), the following Kp, Ki, and Kd values were determined (Table 1). Table 1. Gain values as determined by Kmax and Tmax
When set to a reference RPM of 109, the system displayed varied responses dependent on the controller type (Fig. 11). Three of the controllers presented a large initial overshoot with increasingly dampened oscillations, while one configuration demonstrated a near logarithmic increase towards the setpoint. Specifically, the PID and No Overshoot controllers had the shortest settling time, thus proving the most efficient. The optimal gain parameters for the PID and No Overshoot controller designs were as follows: Kp = 1.5, Ki = 5, Kd = 0.1125 which resulted in a small initial overshoot and a settling period of approximately 2.6 seconds, and Kp = 0.5, Ki = 1.67, Kd = 0.2778 in which there was no initial overshoot and a settling period of approximately 2.3 seconds. These gain parameters were predicted values from the ZN-tuning method, thus proving the method’s efficacy.
Figure 11. Results of ZN-tuning method. Optimal tuning configurations are bolded. The horizontal line represents the reference RPM Evidently, the system successfully self-corrected to the reference point while minimizing the steady-state error. The system’s persisting yet minimal steady-state error after 2.25 seconds is likely the result of the limitations in the system’s intrinsic precision capabilities. It can be noted that the system was able to autonomously reach the reference input both when it was undisturbed and after experiencing external interference, demonstrating the feasibility of this control method for robotics applica68 | 2021-2022 | Broad Street Scientific
tions. The ZN-tuning ultimately proved the practical benefits of systematic PID optimization rather than arbitrary guessing values for each gain parameter. The ZN-tuning implementation produced an efficient response from the system in which the settling period was between two and three seconds with minimal steady-state error. This result is significantly better compared to the PID system without ZN-tuning, in which settling time was approximated at seven seconds with persisting oscillations. While the oscillation experiment in this study was conducted on the physical system, in the future, it can instead be simulated using software such as MATLAB and the extension it offers for the Arduino [21]. Simulating the oscillation experiment is desirable if the ZN-tuning were to be performed on an expensive or resource-intensive system. Although the ZN method is significantly less mathematically intensive than methods like the Root Locus [22], it offers a practical enhancement to PID control for many common optimization problems. Because of the nature of ZN-tuning, the mathematical representation of the system need not be known to tune it, in contrast to the Root Locus method. As a result, utilizing ZN-tuning for PID systems offers an intelligent control method with demonstrated success. The findings of this implementation present avenues for replication at the high school and undergraduate levels as an introduction to control systems engineering. This experiment modeled motor rotation control as it would be used in an industrial robotics setting. The demonstration of the PID control algorithm and tuning method on the constructed sensor-circuit mechanism demonstrates applications for subsystem control while presenting a fundamental understanding of electronic device operation. Furthermore, the successful implementation of the PID algorithm on the TI RSLK Max for chassis control and line following demonstrates the algorithm’s success for dynamic robotic systems. It also successfully reveals how control systems enable precision control and autonomy in complex electronic devices. However, it is essential to note that PID control and thus ZN-tuning are only applicable to systems with a stagnant setpoint and defined steady-state. Computational methods of greater precision and intensity are needed when implementing control on autonomous systems attempting to reach a dynamic setpoint. The vision system of an autonomous car, for example, would require an adaptive control method such as a neural network to account for changing driving conditions [23]. 5. Conclusion In this study, the Ziegler-Nichols tuning method was applied to the Proportional-Integral-Derivative control algorithm used to achieve efficiency in DC motor rotation. An Arduino-based circuit was constructed as the test system upon an initial literature review in control ENGINEERING
theory. A photoelectric sensor measured the RPM from a wheel attached to the DC motor. The system was characterized as linear, after which PID control was implemented on the system through code. The ZN tuning method was applied to derive the optimal gain parameters for the PID algorithm. It was found that the system demonstrated a desirable response, as determined by the No Overshoot controller configuration and minimal settling time (2.3 seconds), converging to the setpoint when the gain values were as follows: Kp = 0.5, Kp = 1.67, Kp = 0.2778. Implementation of intelligent PID control was also successfully extended to an autonomous car. Line following and multiple motor regulation for straight line travel were demonstrated through closed-loop feedback from sensors. The success of this study demonstrates an array of further applications, such as employing a different control algorithm on the system and automating the tuning process with MATLAB. This study can be repeated in the classroom setting at a high school and undergraduate level to introduce circuits and control systems. It also serves as an understanding of control methods implemented on complex robotic devices and demonstrates the accuracy of their subsystems. 6. Acknowledgements The author thanks Mr. Robert Gotwals, Dr. Mariah King, and North Carolina School of Science and Mathematics for the opportunity to pursue this work. Gratitude is also extended to the author’s immediate family. References 1. Mohammed, Reham. Trajectory Tracking Control for Robot Manipulator Using FractionalOrder PID Control. 2015. 2. Paz, Robert. The Design of the PID Controller. Jan. 2001. 3. Swadi, Salah, et al. Design and Simulation of Robotic Arm PD Controller Based on PSO. 2017. 4. Bansal, Hari, et al. PID Controller Tuning Techniques: A Review. Nov. 2012, pp. 168–76. 5. WPILib, FRC. “PID Control in WPILib.” FIRST Robotics Competition Documentation,https://docs.wpilib.org/ en/latest/docs/software/advanced-controls/controllers/ pidcontroller.html. Accessed 19 Feb. 2021. 6. Kaul, Upender K. “First Principles Based PID Control of Mixing Layer: Role of Inflow Perturbation Spectrum.” 7th AIAA Flow Control Conference, American Institute of Aeronautics and Astronautics, 2014. DOI.org (Crossref), doi:10.2514/6.2014-2222.
ENGINEERING
7. mars.nasa.gov. Mars Exploration Rovers. https://mars. nasa.gov/mer/. Accessed 21 Feb. 2021. 8. “Da Vinci Surgery — Robotic Assisted Surgery for Patients.” DaVinci Surgery, https://www. davincisurgery.com/. Accessed 21 Feb. 2021. 9. Christopher Lum. Designing a PID Controller Using the Ziegler-Nichols Method. 2019. YouTube,https://www. youtube.com/watch?v=n829SwSUZ_c. 10. The Complete Guide to Everything. PID Tuning: The Ziegler Nichols Method Explained. 2015.YouTube,https:// www.youtube.com/watch?v=nvAQHSe-Ax4. 11. Arduino Official Store — Boards Shields Kits Accessories. https://store.arduino.cc/usa/. Accessed 20 May 2021. 12. “PID Controller.” Wikipedia, 19 May 2021. Wikipedia, https://en.wikipedia.org/w/index.php?title=PID_controller&oldid=1023906733. 13. Tricks for Controlling DC Motors - Arduino Project Hub. https://create.arduino.cc/ projecthub/tolgadurudogan/tricks-for-controlling-dc-motors-3a05a5?ref=search& ref_id=PID&offset=12. Accessed 20 May 2021. 14. Robot Room: Bipolar Transistor HBridge Motor Driver. http://stan.cropsci.illinois.edu/people/LSY_teaching/ Fall2008/bioengin2008/Top/Manuals/Motors/ DC_MotorDrivers/BipolarHBridge.html. Accessed 20 May 2021. 15. How to Build an H Bridge Circuit with Transistors. http://www.learningaboutelectronics. com/Articles/H-bridge-circuit-with-transistors.php. Accessed 20 May 2021. 16. Soares Augusto, Jose, et al. Arduino Implementation of Automatic Tuning in PID Control ofRotation in DC Motors. 2014. ResearchGate, doi:10.1007/978-3-319-103808_21. 17. Yamamoto, Toru, et al. “Design and Implementation of a Self-Tuning PID Controller.” IFACProceedings Volumes, vol. 31, no. 22, Aug. 1998, pp. 59–64. ScienceDirect, doi:10.1016/S14746670(17)35921-9. 18. Improving the Beginner’s PID – Introduction Project Blog. http://brettbeauregard.com/ blog/2011/04/improving-the-beginners-pid-introduction/. Accessed 20 May 2021. 19.PID - Arduino Reference. https://www.arduino.cc/reference/en/libraries/pid/. Accessed 20 May 2021. Broad Street Scientific | 2021-2022 | 69
20. Arduino PID Control Tutorial — Make Your Project Smarter. https://www.teachmemicro. com/arduino-pid-control-tutorial/. Accessed 20 May 2021. 21. Christopher Lum. Getting Started with the Matlab Support Package for Arduino Hardware.2018. YouTube, https://www.youtube.com/watch?v=8NQ1h0gGgX8. 22. Understanding and Sketching the Root Locus. 2019. YouTube, https://www.youtube.com/ watch?v=gA-KOk3SAb0. 23. Goyal, Manish, and Parasara Sridhar Duggirala. NeuralExplorer: State Space Explorationof Closed Loop Control Systems Using Neural Networks — SpringerLink. https://link. springer.com/chapter/10.1007/978-3-030-59152-6_4. Accessed 20 May 2021. 24. PID Theory Explained. https://www.ni.com/en-us/innovations/white-papers/06/ pid-theory-explained.html. Accessed 20 May 2021. 25. Figure 2. Generic closed loop control system with a PID controller. (n.d.). ResearchGate. Retrieved March 18, 2022, from https://www.researchgate.net/figure/Generic-closed-loop-control-system-with-a-PID-controller_ fig2_301866985. 26. Fig. 2. Response Curve of Ziegler-Nichols Method. (n.d.). ResearchGate. Retrieved March 18, 2022, from https:// www.researchgate.net/figure/Response-Curve-ofZiegler-Nichols-Method_fig2_337514918. 27. Figure 1. Control system with negative unity feedback. (n.d.). ResearchGate. Retrieved March 18, 2022, from https://www.researchgate.net/figure/Control-system-with-negative-unity-feedback_fig1_338191064.
70 | 2021-2022 | Broad Street Scientific
ENGINEERING
MINIMUM NUMBER OF TRIANGLES WITH INTEGER SIDE LENGTHS IN A SATURATED ARRANGEMENT: SOLUTION TO THE AMERICAN MATHEMATICAL MONTHLY PROBLEM 12257 Takumi Fujita, Sean Kim Abstract A saturated arrangement is defined as an arrangement of equilateral triangles such that the intersection of any two is either empty or a common vertex, and every vertex is shared by exactly two triangles. By generalizing the arrangement to sets of four triangles (denoted as a quadruple set) while utilizing the Law of Cosines and the solution to a specific Diophantine equation, we derive a function that outputs an integer-side length quadruple sets given two integer values x and y. By showing that n and disproving each of n = 4, 6, and 8, we prove that n = 10 is the minimum number of equilateral triangles required to form a saturated arrangement in a plane. 1. Introduction The American Mathematical Monthly (AMM) problem 12257 states as follows:
This law is applicable when expressing one side of a triangle in terms of two sides and the angle between them.
12257. Proposed by Erich Friedman, Stetson University, DeLand, FL, and James Tilley, Bedford Corners, NY. An arrangement of equilateral triangles in the plane is called saturated if the intersection of any two is either empty or is a common vertex and every vertex is shared by exactly two triangles. What is the smallest positive integer n such that there exists a saturated arrangement of n equilateral triangles with integer length sides?
2.2 Niven’s theorem: If (in radians) and sin(x) are both rational, then the sine takes values of 0, ±1/2, and ±1. [1]
Simply put, a triangle in a saturated arrangement must have each of its vertices intersecting with the vertex of exactly one other triangle, such that no triangles intersect at any point other than a vertex. With this, we can deduce that the perimeter of a saturated arrangement must form a polygon, based on the arrangement’s finite nature. In other words, the triangles must point inwards, with the outer edge of each triangle forming the polygon.
Quadruple set: n = 12
This is a useful theorem stating that the only rational outputs to a sine function are 0, ±1/2, and ±1. This theorem can be expanded for cosine functions as well, and will be used to restrict certain angles to obtain rational side lengths.
Let’s consider the following shape (Fig. 1).
2. Methodology We used the graphing function of Geogebra to construct different sets of saturated arrangements of triangles. With the calculator capabilities of Desmos and Wolfram Alpha, we were also able to model and solve various Diophantine equations. 2.1 Law of Cosines:
MATHEMATICS AND COMPUTER SCIENCE
Figure 1. A basic quadruple set If we can obtain an arrangement of this shape with integer sides, we can simply connect multiple units to form a saturated arrangement. Since a and b can be any given
Broad Street Scientific | 2021-2022 | 71
integers, let us consider when d and c will be integers as well. Finding d and c to be rational is sufficient since we can obtain integer values for d and c by multiplying the rational values by some natural number. Using the law of cosines, d and c can be written as
In a similar way, since we let
,
From Equations (3) and (4), (5)
To obtain d and c such that d, c , cos( ) and cos( ) must satisfy cos( ), cos( ) , as well. Since the internal angle of every triangle is , + = and < , < . By Niven’s theorem, it is known that the only angles that would give rational cosines are 0, , , , ... The only pairs of and that would satisfy this are = , = . Now we know that
Following a similar process, we obtain from Equations (1) and (3),
Now, we have a, b, c, and d in terms of x and y only. However, we must consider some restrictions on x and y. For a and b to be positive, also has to be positive.
Let b = a. Then
Since
has to be a perfect square, let
From Equation (5), if c > 0, then xy > 0. Therefore, since xy > 0 and x + y > 0, x and y should satisfy (1)
Therefore, Let
has to be a perfect square again. . Then (2)
Since tion
has to be a perfect square again, let . The solution to the Diophantine equais known to be [2]
For our purposes, we will use x, y instead of y1, y2. Therefore, for sides c, d to be rational, (3) Plug in tain
to our Equation (2), and we ob-
As long as they follow the conditions stated in Equation (6), any set of integer values of x and y will result in a quadruple set of triangles with rational sides as denoted below. (7) (8) To make b, c and d rational, simply change the value of a so that the remaining sides become integers. For simplicity, we will let a = 4xy. Then, (9) (10) (11)
(4)
72 | 2021-2022 | Broad Street Scientific
(6)
Finally, we can arrange 3 quadruple sets of triangles to MATHEMATICS AND COMPUTER SCIENCE
form a saturated arrangement (which can be scaled up to fulfill the integer side length condition).
can also be calculated by adding the two triangles FBD and BCD. First, let’s find
Figure 2: A saturated diagram of 12 triangles at x = 1, y = 5. (Scaled down by a factor of 1/4) Quadruple set: n = 10 Next, let’s consider an arrangement of 10 triangles. An example arrangement is shown below in Figure 3.
Figure 4: Half of the quadruple set Since
(13) Equations (12) and (13) are equal to each other Figure 3: The basic form of a 10 triangle arrangement. To investigate if this is possible or not, we need to find out the lengths e and f. In order to find theses lengths, length k is needed first. Now let’s consider the area of the quadrilateral BCDF to find the length k. Since is a sum of two triangles with the base (see Figure 4),
(12)
MATHEMATICS AND COMPUTER SCIENCE
(14) Because a, b, c
,k
.
Since the combination of sides e, f and k is identical to the relationship of sides a, b, c from the previous section, we can use Equations (9), (10), and (11). Instead of x and y, variables u and v will be used such that u, v . k corresponds to c, from Equation (7), Broad Street Scientific | 2021-2022 | 73
(15) From Equations (14) and (15) (16) Following the same step, Figure 5: Example of variables x, y, u, v that output side lengths for a n = 10 arrangement (17) Therefore, from Equations (9), (10), (11), (16), and (17), all 6 sidelengths for the n = 10 arrangement can be expressed in terms of x, y, u, v. However, these expressions only guarantee e, f , as they are expressed in fractions. To guarantee e, f , scale a again.
Figure 6 shows an example of a n = 10 saturated arrangement created with variables x = 2, y = 7, u = 1, v = 4. As discussed above, infinitely more possible saturated arrangements can be made by varying these four variables.
3. Results The above processes yield the following equations. (18) (19) (20) (21) (22) Equations (18) to (22) will produce all 6 integer side lengths for the n = 10 arrangement in terms of integers x, y, u, v with the domain expressed above. Figure 5 below shows a few of the infinite combinations of natural numbers x, y, u, v that output a set of integer side lengths. Note that they must follow y > 3x and v > 3u.
Figure 6: Saturated arrangement of 10 triangles at x = 2, y = 7, u = 1, v = 4 4. Discussion Now we have obtained an example of a saturated arrangement of 10 triangles. However, we have not yet proven that n = 10 is the minimum. In this section we will prove that it is impossible to form a saturated arrange-
74 | 2021-2022 | Broad Street Scientific
MATHEMATICS AND COMPUTER SCIENCE
ment with n < 10 sides. We will use:
5. Conclusion
s = number of sides of outer polygon, and n = number of equilateral triangles in the arrangement.
We were able to find the minimum number (n = 10) of triangles required to form a saturated arrangement within a plane. The solution was found by deriving a formula to produce arrangements of 10 triangles, then proving that it is impossible to form a saturated arrangement with less than 10 triangles.
Theorem 1. n Proof. The number of vertices in the arrangement can be given by: as every vertex is used exactly twice. For the number of vertices to be an integer, n must be even. Theorem 2. n > 6 Proof. For the equilateral triangles to fit inside the outer polygon, the sum of the interior angles must be greater than the sum of the two angles of the equilateral triangles. That means
6. Acknowledgements We would like to thank Professors T. Lee and N. Koberstein for their time and support and for providing us with this opportunity. Their assistance greatly contributed to the quality and extensiveness of our research. We would also like to thank the North Carolina School of Science and Mathematics (NCSSM) for making this research possible through the Summer Research and Innovation Program. 7. References
Since there must be 1 equilateral triangle for each side on the outer polygon, n > 6 as well. Theorem 3. n 8 Proof. Assume n = 8. From the proof of theorem 2, s > 6. Also, n = 8 s < 8. Thus, s must equal 7. This means that there will be 7 open vertices in the middle, with 1 equilateral triangle left.
Niven, Ivan (1956). Irrational Numbers. The Carus Mathematical Monographs. The Mathematical Association of America. p.41. MR 0080123 Abdelalim, S., & Dyani, H. (2014). The solution of the Diophantine equation . International Journal of Algebra.8, 729–732. https://doi.org/10.12988/ ija.2014.4779
Figure 7: Arrangement of s = 7 Since it is impossible to fill all 7 of these vertices with 1 triangle, n ̸= 8. Since n must be even and we have proven that n ̸= 2, 4, 6, 8, the minimum number of triangles that would form a saturated arrangement with integer sides is 10.
MATHEMATICS AND COMPUTER SCIENCE
Broad Street Scientific | 2021-2022 | 75
MAPPING THE LANDSCAPE OF RNA-SEQUENCING TOOLS Sherry Liu Abstract Advances in bioinformatics have been instrumental in better understanding large quantities of biological data. Software tools are at the heart of these discoveries, so trends in software usage provide valuable insights for researchers. One way to uncover these trends is to analyze published literature to extract the tools that are used in papers over time. However, current keyword-based approaches fail to identify tool names that are homonyms (e.g. “Star”). This project incorporates citation data to more accurately automate the identification of software tools in bioinformatics literature. The citation of anchor papers, or papers describing a tool’s creation, is used as an indicator for tool usage. This method was applied to tools for the bioinformatics task RNA Sequencing. The PubMed Central database was queried for the term “RNA-Seq” to obtain a corpus of papers and their citations from 2011-2020. Using a random forest decision tree, the anchor papers for each tool are identified with 95.6% accuracy. Any given paper is identified as using a tool if it cites the anchor paper for that tool. The resulting trends are then visualized for tools of interest. This novel method significantly reduces false positives found with keyword-based approaches. These results highlight the value of citation information in analyzing scientific literature to draw insights from textual data. 1. Introduction Software tools are a central component to the fields of bioinformatics and computational biology. Specifically, the task “RNA Sequencing” contains many tools throughout the analysis pipeline, from sequence aligners like “STAR” to differential expression like “DESeq2” (Fig. 1). Researchers select the best tool for a specific task by drawing on their knowledge of available tools, usage frequency, interaction with other tools, and other characteristics. These trends are also a major contributor to an understanding of the field as a whole and its evolution [1].
Figure 1. An overview of tools used in the RNA sequencing pipeline
76 | 2021-2022 | Broad Street Scientific
Although the UK is creating a comprehensive list of bioinformatics software with an accompanying universal ID system through the registry bio.tools, the US has no such registry [2]. Manually extracting the tools from a research paper is a straightforward task, but is tedious to carry out at a large scale. Thus, automated methods are explored to determine what tools are used in bioinformatics literature, a task known as Named Entity Recognition (NER). Previous research has attempted to automate the identification process using keyword search with a dictionary of known softwares. For instance, a series of rules may be applied based on features that positively or negatively impact the likelihood that the keyword is a software tool [3]. While these attempts were promising, the model frequently failed when the software tool was a homonym, such as “Network”, “Analysis”, or “Star”. This leads to false positives, such as a paper about astronomy being labeled as using the bioinformatics tool “Star” despite referring to a different definition of the word. Further studies of ambiguity and variability of database and software names has led to the conclusion that keyword search alone is insufficient to accurately carry out software recognition [4]. The unique structure of scientific literature presents an opportunity for improving NER with bioinformatics tools. In the past, citations have been used to identify and characterize scientific discoveries in the biomedical sciences. However, many identified discoveries were actually software tools, resulting in false positives [5]. This MATHEMATICS AND COMPUTER SCIENCE
is because when researchers use a tool in their project, they frequently cite the paper published when the tool was released, which we define as the anchor paper. Citing this paper is often recommended by creators of the tool. Thus, it is reasonable to assume that if a software’s anchor paper is cited and keywords are identified, then the software is used in the paper. This observation offers the basis of our novel citation-based approach. Although there is substantial research on the use of keyword search and text mining techniques to automate software identification, a citation-based approach has been minimally evaluated [6]. This research could potentially introduce a new means of detection and disambiguation of software tool names. The purpose of the study is to develop and assess a process to automatically identify what softwares are used in bioinformatics research papers and apply it to analyze trends in the usage of bioinformatics tools.
2.3 Identifying Anchor Papers To find the anchor paper for each tool, citation data from the corpus were analyzed through statistical tests and supervised machine learning. Papers likely to be anchor papers had a high proportion of their citations captured by a subset of the corpus that contains the tool name. This proportion is defined as the X-Statistic.
2. Methods 2.1 Tool Dataset Determining a suitable dataset of tools to analyze is crucial to this project. Specifically, precise keywords are important to get an accurate sense of how keyword search functions. The tools considered for this study consist of a list of 526 tools web-scraped from the webpage “List of RNA-Seq bioinformatics tools” by parsing through tool names and their corresponding types [7]. The tools are split into categories, so the name and the narrowest category are gathered for each tool. Next, each tool name is designated as either a homonym or non-homonym in order to assess the behavior of the novel method for homonym tools. A tool is considered a homonym tool if its name appears in the English dictionary, using the Natural Language Toolkit “words” corpus. 2.2 PMC Corpus The desired corpus, or text dataset, is a list of papers that involve RNA-Sequencing, along with what papers are cited by these papers. Since the term “rna-seq” is not a homonym, this corpus can be easily queried using keyword search. To obtain this corpus, keyword search was performed in PubMed Central for the term “rna-seq”. This corpus was then filtered for open access papers published in the 10 year timespan from 2011-2020. The resulting corpus was a list of 81,946 papers containing the term “rnaseq” and their corresponding citations. These papers are stored using their PubMed IDs (PMID), a unique identifier for papers in the PubMed database.
MATHEMATICS AND COMPUTER SCIENCE
Figure 2. Anchor Paper Classification Decision Tree Structure - A given paper moves from the top of the tree down to a specific class depending on its characteristics. A random forest decision tree was created to classify each paper as an anchor paper or not using many decision trees (Fig. 2). Random forest classifiers were chosen over regular decision trees to reduce the risk of overfitting. The classifier was trained on 1026 papers, 171 of which are anchor papers, with an 80/20 split between training and testing data. Some features are dependent on others as a preprocessing of features (Table 1). Table 1. Decision Tree Features Feature
Description
# Cited Total
Number of times the paper is cited in the corpus Number of times the paper is cited among papers in a keyword search of the corpus Proportion of citations captured by the tool subset (# Cited Total/ #Cited Subset) Number of papers in the subset The paper’s position compared to other potential anchor papers after being sorted by X-Statistic
# Cited Subset
X-Statistic
Subset Size Rank
This classifier was then applied to obtain anchor papers. A keyword search was conducted with the tool name on the corpus to create a tool subset. All papers cited at Broad Street Scientific | 2021-2022 | 77
least once in this subset are considered potential anchor papers, as they are cited by papers that contain the tool name. A one-tailed hypothesis test for proportions was conducted (p=0.05) to eliminate papers whose subsets do not show an enrichment of citations. Then, all statistics used in the decision tree were computed for each potential anchor paper. Papers that were classified as anchor papers were identified as such, allowing for multiple anchor papers. Duplicate papers were then removed from the list.
eliminated, and then all tools that were in the top 3 most cited tools in any year were included. These trends are compared to trends found using keyword search. Overall, the tool dataset and corpus are used to automatically generate anchor papers using the decision tree classifier, and then all three datasets are taken together to generate usage trends (Fig. 3).
2.4 Assessing Anchor Paper Accuracy To assess how accurately homonym and non-homonym tools are extracted using keyword search and anchor paper, manual labeling is required to determine whether a tool is actually used in a paper. A random sample of 100 papers was taken for 6 tools with a variety of popularity, functions, and homonym characteristics (Table 2). Each paper is labeled with whether the paper actually used the tool, whether it appears in a keyword search of the tool, and whether it cites the anchor paper for that tool. Table 2. Characteristics of Representative Tools for Labeling Tool Name BowTie2 HTSeq DESeq
Homonym Function No No No
QuoRUM Yes EdgeR Yes Salmon
Yes
Number of citations Short aligner 5927 Quality control 3490 Differential 3391 expression Error correcter 12 Differential 5848 expression Quantification 699 analysis
2.5 Extracting Trends In this study, we defined a paper as using a tool if it cites the anchor paper for that tool. This method is referred to as the “anchor paper method” for NER. Using the list of tools, their corresponding anchor papers, and the RNA-Seq corpus, temporal popularity trends were extracted for tasks and tools of interest. The data were normalized based on overall growth of the corpus. Trend analysis occurs in two steps. First, the total number of citations is compared between fields. Then, individual tools are analyzed within specific fields. This allows for comparison of overall RNA-Seq trends and the popularity of specific tools. To create a visualization, tools with few overall citations relative to other tools were 78 | 2021-2022 | Broad Street Scientific
Figure 3. Project Pipeline 3. Results 3.1 Tool Data Out of 526 tools, 79 (15%) were homonym tools. It is interesting to note that 284 (54%) of tool names have non-traditional capitalization, such as being in all caps or having capital letters in the middle of the word. 3.2 Paper Corpus In the corpus of papers containing the term “rna-seq”, there were over 1,500,000 papers cited at least once in total. The average paper had 55 citations and was cited 2.8 times. Around 40% of papers in the corpus, not including citations, were cited by other papers in the corpus. The size of the corpus has increased exponentially over time from 703 papers in 2011 to 20,497 papers in 2020 (Fig. 4). This reflects the expansion of the research community, increased digitization of research papers, and growth of the bioinformatics field.
MATHEMATICS AND COMPUTER SCIENCE
a)
Figure 4. Corpus Size over Time - The number of papers in the corpus for each year exponentially increases from 2011 to 2020.
b)
The top 5 most frequently cited papers in our corpus were the anchor papers for DESeq 2, edgeR, Bowtie2, STAR, and SAMtools. This indicates that tools and their anchor papers play an important role in the dataset. 3.3 Keyword Search for Homonym and Non-Homonym Tools The number of keyword hits over time for two sample tools in the corpus differs for homonym and non-homonym tools (Fig. 5). While non-homonym tools like HISAT show an increase in use after tool creation as expected, homonym tools like “STAR” have hits even before tool creation, displaying many false positives.
Figure 5. Keyword Hits Trends - Keyword search behaves as expected for non-homonym tools (a) but creates false positives in homonym tools (b). These false positives for the tool “Salmon” were investigated, and the following results were obtained: • 32 papers with the tool “salmon” as desired • 55 papers with the fish “salmon” • 9 papers with the color “salmon” often in figures • 3 papers with authors named Salmon • 1 paper with a corn gene called “salmon” This demonstrates the great variability and prevalence of alternate senses of homonym tools, and that false positives are indeed caused by these variations of the word. 3.4 Anchor Paper Classifier The anchor paper classifier achieved an accuracy of 95.6%. The features X-Statistic and # Cited Subset had the highest importance in the classifier. The number of citations for the obtained anchor papers is heavily skewed while the publication dates are approximately normal and centered at 2012 (Fig. 6).
MATHEMATICS AND COMPUTER SCIENCE
Broad Street Scientific | 2021-2022 | 79
a)
Table 3. Anchor Paper Accuracy - These tables display the results from manual labeling. The anchor paper method is found to have near perfect precision and lower recall. a)
b)
b)
Figure 6. Anchor Paper Distributions - These figures display the number of times our final anchor papers are cited (a) and their publication date (b). 3.5 Anchor Paper Accuracy for Homonym and Non-Homonym Tools Results from labeling papers found in keyword search for our 6 representative tools are shown in Table 3. There were many false positives in keyword search (68%), but these were eliminated with citation of anchor paper. For non-homonym tools, keyword search resulted in perfect classification. To assess the effectiveness of the model, standard metrics characterizing the levels of True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN) were used. The random sample experiments allow us to assess the precision - TP/(TP + FP) - and recall TP/(TP + FN) of the anchor paper method. Anchor papers have a very high precision, with few false positives, or papers that cite the anchor paper but do not use the tool (Table 3). The recall is lower.
For non-homonym tools, anchor papers were cited roughly 60% of the time. This can be taken as an inherent error rate for anchor papers due to researchers citing other papers or URLs. 3.6 Overall Trends in RNA-Sequencing The number of total citations for several significant fields of RNA-Seq display temporal trends (Fig. 7). Alignment seems to be on a decreasing trend as the task becomes more well-defined in recent years. Gene expression analysis is on the rise, indicating the increased use of RNA-Seq for tasks like biomarker discovery.
Figure 7. Overall Tool Trends - Different RNA-Seq tasks become more and less frequently cited over time
80 | 2021-2022 | Broad Street Scientific
MATHEMATICS AND COMPUTER SCIENCE
3.7 Task-Specific Trends In both task-specific visualization graphs, various tools can be seen competing for usage within their respective fields (Fig. 8). One clear trend is the relationship between Bowtie and Bowtie2, as BowTie2 overtakes a declining Bowtie around 2015. While this surge for Bowtie2 pushes tools like Maq to virtually no usage, Bowtie still retains some usage due to its niche application to shorter sequences. a) Figure 9. Keyword Search vs. Anchor Paper - The anchor paper method (light blue) eliminates the false positives found in the keyword search method (dark blue). 4. Discussion
b)
Figure 8. Short Unspliced Aligners and De Novo Splice Aligners – Specific tools in these categories vary in usage over time. Finally, trends generated from the anchor paper method were compared to those from the keyword only approach and were found to have significantly fewer false positives (Fig. 9).
MATHEMATICS AND COMPUTER SCIENCE
From visualizing keyword search trends, it is clear that keyword search alone is insufficient to determine usage trends for homonym tools. This is confirmed by the many incorrectly classified papers found in the labeling of papers with homonym tool names. This agrees with previous work done by the University of Manchester which found that NER in scientific literature is a difficult task due to ambiguous terms [7]. From labeling papers found in keyword search, the anchor paper method is found to have a very high precision and lower recall. Papers that cite the anchor paper almost always use the tool, while some papers that use the tool do not cite the anchor paper. These false negatives occur relatively frequently and are an inherent error to our process, as researchers sometimes cite URLs, other papers, or forget to cite. Even so, the near perfect precision of anchor papers is still promising, as the inherent researcher error is consistent across all tools and therefore does not affect the relative proportion trends. The random samples also reveal that the homonym characteristic is a trend, rather than a clear-cut line. For example, some homonym tools like EdgeR have no false positives in keyword search because while “Edger” is an English word, it is not commonly used in the corpus. Each tool can be on the spectrum of homonym character, depending on the popularity of the tool in the corpus and of the word in the English language. The anchor paper classifier achieved high accuracy and can be considered reliable. The misclassified papers were often anchor papers of other tools. The final usage trends demonstrate the rise and fall of various tools and fields over time, which are captured through anchor paper citations. It is found that there are many tools with few or no citations, a result that agrees Broad Street Scientific | 2021-2022 | 81
with previous surveys of bioinformatics software. These graphs tell the stories within bioinformatics as tools compete for usage, complement each other, and at times fall into irrelevancy. 5. Conclusion Overall, this experiment successfully demonstrated the robustness of using anchor papers as a proxy for tool usage in scientific literature. It is found that anchor papers are a highly precise means of assessing tool usage which can be reliably generated automatically through machine learning methods. The anchor paper method is especially useful when applied to homonym tools, eliminating many false positives. Our results demonstrate citation data as an efficient and accurate way to draw insights from scientific literature. This project has many applications to the bioinformatics field. While current databases such as bio.tools must manually label anchor papers, the anchor paper classifier created here can help automate this process. The same can be done for what tools a given paper uses. The final trend graphs are valuable to researchers to understand the field of RNA-Seq and its tools. Further, this method allows for tool extraction without the full text of a paper, a result that is unseen in previous attempts to extract trends through text mining and valuable in increasing the set of available papers for analysis. There are limitations to this work. The list does not include more recent tools that indicate emerging innovations in the field. The inherent error rate results in some mislabeling of tools used. Additionally, tool popularity is not the only characteristic relevant to understanding the overall value of a tool. Other factors such as performance tests and reproducibility should also be considered. This project contributes to a larger effort to automate the entire pipeline to catalogue, characterize, and analyze bioinformatics tools. In the future, the method described here can be easily extended to other fields, not only within bioinformatics but also in other contexts like diseases, discoveries, and inventions.
tools database. Plos Computational Biology. https://doi. org/10.1371/journal.pcbi.1006245 [2] Ison, J. et al. (2015). Tools and data services registry: a community effort to document bioinformatics resources. Nucleic Acids Research. https://doi.org/10.1093/nar/ gkv1116 [3] Duck, G., Nenadic, G., Brass, A., Robertson, D. L., & Stevens, R. (2013). Bionerds: Exploring bioinformatics’ database and software use through literature mining. BMC Bioinformatics, 14(1), 194. https://doi.org/10.1186/14712105-14-194 [4] Duck, G., Kovacevic, A., Robertson, D. L., Stevens, R., & Nenadic, G. (2015). Ambiguity and variability of database and software names in bioinformatics. Journal of Biomedical Semantics, 6(1), 29.https://doi.org/10.1186/ s13326-015-0026-0 [5] Small H, Tseng H, Patek M (2017). Discovering discoveries: Identifying biomedical discoveries using citation contexts. Journal of Informetrics, 11(1), 46–62. https://doi. org/10.1016/j.joi.2016.11.001 [6] Duck G, Nenadic G, Filannino M, Brass A, Robertson DL, Stevens R (2016) A Survey of Bioinformatics Database and Software Usage through Mining the Literature. PLoS ONE 11(6): e0157989. https://doi.org/10.1371/journal. pone.0157989 [7] List of RNA-Seq bioinformatics tools (2015). Transcriptome Sequencing Research and Industry News, https://www.rna-seqblog.com/list-of-rna-seq-bioinformatics-tools-2/
6. Acknowledgements Thank you to Dr. Steffen Heber at NC State for serving as my mentor throughout this project. Thank you to Dr. Bey, Dr. Fuchs, the NCSSM Mentorship team, the Burroughs Welcome Fund (BWF), and the NCSSM Foundation for their generous support of my work. 7. References [1] Zappia L, Phipson B, Oshlack A. (2018). Exploring the single-cell RNA-seq analysis landscape with the scRNA82 | 2021-2022 | Broad Street Scientific
MATHEMATICS AND COMPUTER SCIENCE
SIMULATING QUANTUM KEY DISTRIBUTION IN THREE POLARIZATION BASES Katherine Panebianco Abstract With the rise in quantum computing’s potential to break current encryption schemes, new methods of data encryption are needed to keep messages secure. Quantum Key Distribution (QKD) presents a solution that uses superposition of quantum states to ensure perfectly secure encryption. Our research simulated an established method of QKD (BB84 Protocol) and a proposed new method (3 Basis Protocol). Data collection focused on the key rate (number of particles sent to obtain a key of a certain length) in a non-eavesdropper setting and error rate (percentage of the key that contains errors) in an eavesdropper setting. We hypothesized that the 3 Basis Protocol would have a lower key rate, but that the BB84 Protocol would have a lower error rate. Five hundred simulations were run of each protocol, confirming the 3 Basis Protocol’s hypothesized lower key rate, but revealing both protocols to have the same error rate. These results suggest that there is no advantage in adding a third basis to QKD procedures. Further experimentation with the number of bases in QKD protocols has proved that the error rate is constant for any number of bases, supporting the conclusion that the two basis BB84 Protocol remains the most efficient.
1. Introduction 1.1 - Motivation In today’s increasingly technology-oriented world, being able to keep information secure online has become more important than ever before. Most current encryption methods rely on algorithms built with products of extremely large prime numbers [1] and the idea that these prime factorizations are too large for classical computers to find. Data are securely encrypted simply because it would take computers too long to decrypt the information. However, quantum computers offer a computational advantage. With a large enough working quantum computer, the current encryption algorithms would be rendered useless, since they would be breakable in a reasonable amount of time [1]. As technology continues to grow closer to constructing quantum computers large enough to do so, additional attention needs to be focused on encryption methods resistant to an attack by a quantum computer. Quantum cryptography is one such encryption method that can fill this gap. Since messages are kept safe by the physical principles of quantum mechanics, data would remain secure against an attack by a quantum computer. 1.2 - The BB84 Protocol The field of quantum cryptography is mostly focused on the idea of Quantum Key Distribution (QKD), distributing a key via quantum particles, like photons. Certain traits of these quantum particles, such as their polarizations, can be used to symbolize different bit values (either a zero or a one). By measuring the particles in multiple PHYSICS
bases, an element of randomness is added that catches eavesdroppers before the message has even been sent. One of the most commonly known methods of QKD is the BB84 Protocol, named after scientists Bennett and Brassard in 1984 [2]. BB84 uses two polarization bases rotated 45 degrees from each other to send a string of photons between two parties, Alice and Bob. In this procedure, Bob measures each photon with a randomly selected basis and records the bit value he obtains. Alice and Bob then compare the bases that she polarized in and he measured in, keeping only the bits for which their bases lined up. The result is a matching string of bits that can be used as a key (Fig. 1) [3].
Figure 1: Breakdown of BB84 into key steps [3]. Here, Alice generates a string of bits and bases to encode her particles, which she sends to Bob. Once he receives them, Bob measures the particles and he and Alice compare bases to finish up the transmission of the key. Broad Street Scientific | 2021-2022 | 83
The BB84 Protocol works based on the idea that measuring a quantum particle can fundamentally alter the state it is in. Alice and Bob only keep the particles for which their bases matched, because if Bob used the wrong basis, the particle may have collapsed into a bit value that does not match what Alice originally sent. This is why the BB84 Protocol is able to catch eavesdroppers. At the end of the protocol, Alice and Bob compare a subset of their key to make sure that they match. As long as there were no eavesdroppers on the line, there should be no errors. However, if there was an eavesdropper, Alice and Bob would encounter an issue with some of their bit values not matching. This is because when an eavesdropper intercepts and measures the particle, they too may pick the wrong basis, causing the particle to collapse into a state that does not match Alice’s. As a result, Bob ends up receiving some faulty particles from the eavesdropper, increasing the chance that his key differs from Alice’s, despite their bases matching. When an eavesdropper intercepts Alice’s photons, some of the particles’ states are scrambled in the process, causing Bob to obtain different key values (Fig. 2). [3]
Figure 2. Impact of an eavesdropper on the BB84 Protocol. When an eavesdropper intercepts the particles that Alice sends before Bob can measure them, the states of some of the particles that Bob receives may not match what was initially sent. As a result, some of the bits in Bob’s final key could differ from those in Alice’s. Here, Alice’s key is a string of bits, and the particles that Alice and Eve send are denoted as H, V, D, and A for horizontal, vertical, diagonal, and antidiagonal, respectively. Bob obtains a key of 0s and 1s based on the particles that he receives. 1.3 Research Goal This project explores the idea of a QKD protocol similar to the BB84 Protocol, with one key difference. Rather than measuring in two bases as in BB84, we propose a three basis protocol that utilizes three bases angled 30 degrees off from each other (Fig. 3). This Three Basis Protocol is otherwise identical to the BB84 Protocol, with Alice sending Bob a string of quantum particles to be measured, Bob measuring them, and Alice and Bob disposing of the bits for which their bases did not match.
84 | 2021-2022 | Broad Street Scientific
Figure 3. The Three Polarization Bases. For the proposed Three Basis Protocol, Alice can polarize the particles and Bob can measure them in one of three polarization bases: 0, 30, and 60. This project questioned how the addition of a third polarization basis affects the key rate (the number of particles that need to be sent to obtain a key of length n) between Alice and Bob. We also investigated how, in a situation with an eavesdropper, the third basis affects the error rate (the proportion of bits that do not match when keys are compared) between the two parties. We hypothesized that the key rate would be lower in the three basis setup than it would be for the two basis setup, as the addition of a third basis would decrease the probability that Alice and Bob would pick the same polarization and measurement bases. We also predicted that the additional element of randomness added by the third polarization would cause the error rate between Alice’s and Bob’s keys to increase in an eavesdropper situation, making it more likely that the eavesdropper would be detected. 2. Procedure In the BB84 Protocol, the particles can be polarized in one of four directions: horizontal, vertical, diagonal, and antidiagonal (H, V, D, and A). Together, the horizontal and vertical directions make up one basis, and the diagonal and antidiagonal directions make up the other. Using these bases, Alice is able to encode her string of bits to send to Bob. However, as mentioned earlier, when Bob chooses to measure in the wrong basis, he has a chance of obtaining either bit value (Fig. 4). For the BB84 Protocol, the two bases are rotated 45 degrees from each other. Notably, this rotation angle means that any particle polarized in one basis but measured in the other has a 0.5 probability of returning one measurement value (0 or 1) and a 0.5 probability of returning the other measurement value (0 or 1).
PHYSICS
Figure 4: Encoding Schemes for the BB84 Protocol [3]. In the BB84 Protocol, the two polarization bases (here numbered 0 and 1) combine to form four possible polarization directions: horizontal, vertical, diagonal, and antidiagonal. When Alice sends Bob one of these particles, he can only definitively obtain a certain measurement if he uses the same basis that Alice polarized the particle in (0 basis for H and V, 1 basis for D and A). For the Three Basis Protocol, each of the bases is rotated 30 degrees from the other two. Similar to the BB84 Protocol, each basis is composed of two potential polarization directions, meaning the Three Basis Protocol has a total of six possible polarization states. Once again drawing parallels to the BB84 Protocol, when Bob measures in the wrong basis, he could obtain either a 0 or a 1 for his final measurement (Fig. 5). The big difference between the two protocols lies in the effects of the angles between the bases. Since the Three Basis Protocol has every basis rotated 30 degrees from the others, any particle measured in a different basis than it was polarized in will have a 0.75 probability of returning one measurement value (0 or 1) and a 0.25 probability of returning the other measurement value (0 or 1).
To compare the key rates and error rates of the two protocols, we simulated BB84 and Three Basis in two different settings, one without an eavesdropper and one with an eavesdropper. These simulations were done in Python on a classical computer. Due to the probabilistic nature of quantum mechanics and the known various probabilities for the two protocols explained above, the random. randint function, which generates random numbers, mimicked the behaviors of a quantum particle for this experiment. We ran 500 trials for each setting, providing confidence that the programs were correctly simulating these quantum mechanical processes. Additionally, for each of the simulations, the program was manually checked before collecting data. The initial number of particles that Alice sent to Bob was set to 10, and each step of the QKD procedure was printed in order to compare the computer’s calculation to the protocol worked out by hand. All of the simulations were found to be working as they should. 3. Results and Analysis Once the accuracy of the programs was verified, the number of particles that Alice sent to Bob was raised to 100 in order to have a larger initial pool of data for each simulation. In the situation with no eavesdropper, we modeled the BB84 and Three Basis protocols 500 times each, and collected data for the key rate (what proportion of the 100 particles Alice sends make it into the final key). These key rates were then plotted on histograms to display the results (Fig. 6). The BB84 Protocol had a mean key rate of 0.501 and the Three Basis Protocol had a mean key rate of 0.336. Note that the variation between each trial within simulation sets is due to the fact that there is a strong element of randomness in QKD. As such, any outliers are simply due to random chance.
Figure 5: Encoding schemes for the Three Basis Protocol. Similar to Figure 4, the three polarization bases here combine to result in six possible polarization directions for the particles that Alice sends. The 0, 30, and 60 bases in Figure 3 correspond to the 0, 1, and 2 bases, respectively. Likewise, if Bob does not measure a particle in the same basis that Alice polarized it in, he cannot be certain that he will obtain a 0 or a 1 bit. PHYSICS
Broad Street Scientific | 2021-2022 | 85
a)
b)
a)
b)
Figure 6: Key rates for (a) the BB84 Protocol and (b) the Three Basis Protocol. The (a) BB84 Protocol had an average key rate of 0.501 from 500 simulations; the (b) Three Basis Protocol had an average key rate of 0.336 from its 500 trials. Both of these values do not differ significantly from what was expected (0.5 and 0.3333, respectively), confirming the original hypothesis.
Figure 7: Error rates for (a) the BB84 Protocol and (b) the Three Basis Protocol. The (a) BB84 Protocol had an average error rate of 0.248; the (b) Three Basis Protocol had an average error rate of 0.254. Neither of these results differ significantly from the expected 0.25 error rate for the BB84 Protocol, providing evidence that the true error rate in an eavesdropper setting is 0.25 for both protocols.
We then simulated the BB84 Protocol and Three Basis protocol with an eavesdropper intercepting and measuring Alice’s particles before sending them to Bob. For these simulations, data collection focused on the error rate (what proportion of the particles in Alice’s and Bob’s final keys do not match as a result of the presence of an eavesdropper). Similar to the no eavesdropper situation, we set the programs such that Alice initially sent 100 particles to Bob. Five hundred simulations ran for both protocols, and the error rates were recorded. We then plotted these error rates on histograms to display the results (Fig. 7). The BB84 Protocol had a mean error rate of 0.248 and the Three Basis Protocol had a mean error rate of 0.254.
For the 500 trials of the BB84 Protocol without an eavesdropper, the mean key rate was determined to be 0.501 with a standard deviation of 0.051. Since the BB84 Protocol is known to have a key rate of 0.5, this simulation was verified with a z-test. This resulted in a p-value of 0.314, affirming that the data collected are an accurate model of the BB84 Protocol’s known key rate. Similarly, the Three Basis Protocol was hypothesized to have a key rate of about 0.3333 (1⁄3), as about 1⁄3 of the time, Alice and Bob would happen to choose the same polarization and measurement bases, resulting in the bit being included in the key. The 500 non-eavesdropper trials for this protocol recorded a mean of 0.336 with a standard deviation of 0.046. Comparing this mean to the predicted 0.3333 via a z-test resulted in a p-value of 0.143, supporting the accuracy of the hypothesized key rate. In an eavesdropper setting, the BB84 Protocol is known to have an error rate of 0.25; the 500 trials had a
86 | 2021-2022 | Broad Street Scientific
PHYSICS
mean of 0.248 with a standard deviation of 0.061. A z-test with this data yielded a p-value of 0.219, confirming that the simulation was working as expected. For the Three Basis Protocol, we predicted that the error rate in an eavesdropper setting would be higher than the error rate for the BB84 Protocol (0.25). However, a z-test using the Three Basis Protocol’s trial mean of 0.254 and standard deviation of 0.077 revealed no difference between the error rates of the two protocols. 4. Discussion The original hypothesis suggested that the BB84 Protocol would have a greater key rate (of 0.5) compared to the Three Basis Protocol (predicted 0.3333). Running the simulations of both protocols affirmed this, presenting a statistic in favor of the BB84 Protocol. Since one aspect of cryptography is the efficiency of the steps required to encode a message, having a higher key rate in QKD is extremely important, as it minimizes the number of particles that Alice needs to send to Bob, decreasing the chance that someone else could intercept some of the particles along the way. Only having to send 2n particles for a key of length n for the BB84 Protocol (versus 3n particles for the Three Basis Protocol), slightly speeds up the time it takes for a key to be obtained. However, the data did not support the hypothesis that the Three Basis Protocol would have a greater error rate in eavesdropper situations, making it easier to catch an eavesdropper. Instead, the two protocols were both determined to have an error rate of about 0.25. Further investigation of this interesting result included drawing out a probability tree of the possible scenarios when an eavesdropper intercepts a particle as Alice sends it to Bob. This visually illustrated the math behind each step of QKD, and helped to more intuitively understand the simulation results. The probability tree breaks the process down into several sections: what basis the eavesdropper (“Eve”) chooses to measure in, the polarization of the particle Bob receives from Eve, and the bit value that Bob obtains, assuming he chooses to measure in the same basis as Alice (because otherwise the bit would be discarded) (Fig. 8).
PHYSICS
Figure 8: Probability tree of Three Basis Protocol with an eavesdropper. If Alice decides to encode a 0 bit using the 0 polarization basis (horizontal and vertical directions), she would obtain a particle polarized in a direction denoted here as ‘0x’ (equivalent to horizontal). From there, Eve has a 1⁄3 probability [4] of selecting each basis to measure in, and variable chances of obtaining each potential polarization of the particle (denoted ‘0x,’ ‘30x,’ ‘30y,’ ‘60x,’ and ‘60y’ next to their corresponding probabilities). When Bob measures, he is no longer certain to obtain the same 0 bit that Alice started with, even though he uses the same 0 basis that Alice did. 5. Extension to Four and Five Bases Based on the results of the two and three basis simulations, we then hypothesized that the key rate for any number n bases (where n is an integer) is 1/n, meaning that for each additional basis added, the number of particles that need to be sent in order to receive a key of some set length increases. Furthermore, we hypothesized that the error rate for any number n bases in an eavesdropper setting is a constant 0.25; the number of bases does not impact the probability that parts of Alice’s and Bob’s keys differ due to eavesdropper interception. In order to test these new hypotheses, we extended the simulation method applied for the earlier protocols to structurally similar methods of QKD that employ four and five basis setups. Parallel to what was done for two and three bases, 500 simulations ran for each protocol, collecting data on the key rate and error rate for each. From this sample, the Four Basis Protocol had an average key rate of 0.252 and average error rate of 0.246, neither of which were statistically significant differences from their expected values (0.25 and 0.25). Likewise, the Five Basis Protocol’s simulated average key rate of 0.201 and average error rate of 0.255 did not differ statistically significantly from their expected values of 0.2 and 0.25, respectively (Fig. 9).
Broad Street Scientific | 2021-2022 | 87
a)
b)
c)
4. Discussion Our results all follow the hypotheses described above for the generalization of key rate and error rate values to any number of bases. In fact, it can be explained in a straightforward manner why the general form for the key rate in QKD protocols similar to the ones simulated here is 1/n. As part of the protocol, Alice and Bob throw out all values for which the basis Bob measured the particles with does not match the basis that Alice polarized them in. This means that key rate can essentially be broken down into the probability that Bob picks the same basis that Alice did, out of the n options available. Assuming Bob picks completely randomly, as is expected of the protocol, there is a 1/n chance that he selects the same basis, meaning that the value he records will be included in the final key. However, the explanation as to why the error rate is 0.25 for a protocol using any number of bases is not as clear. Similar to what we did for the Three Basis Protocol, we drew sample probability trees for the Four and Five Basis protocols, providing secondary support that the results from the simulations are reasonable. Upon looking at the structure of the probability trees for the Three, Four, and Five Basis protocols, a pattern in how we calculated the error rates became evident. Specifically, the error rate calculations for all three protocols could be generalized into the following equation.
d)
Figure 9. Key rates and error rates for the Four and Five Basis Protocols. From 500 simulations, the (a) Four Basis Protocol had an average key rate of 0.252 and the (b) Five Basis Protocol had an average key rate of 0.201. Additionally, 500 simulations of the (c) Four Basis Protocol with an eavesdropper had an average error rate of 0.246. The (d) Five Basis Protocol’s 500 eavesdropper simulations had an average error rate of 0.255. These trends are all in line with what was expected from the hypotheses.
An explanation as to why this equation makes sense is as follows. When Eve, the eavesdropper, intercepts a particle, she has a 1/n chance of picking each possible basis. The angle between each polarization basis in a protocol with n bases is 90/n. Inclusion of k in the 90k/n term captures the iterative process that occurs when the probabilities of obtaining an error given Eve selects each basis, rotated 90(1)/n, 90(2)/n, ... , 90/(n-1)n degrees from the basis Alice picked, are summed. The sine and cosine terms are both squared because they are probability amplitudes, and the squares of probability amplitudes provide the probabilities that a particle collapses into various states [4]. Proof that this equation is equal to 0.25 for any n can be achieved by using Euler’s Formula for Complex Numbers, providing a generalized argument that the error rate is 0.25 for any number of bases. 6. Conclusion Quantum Key Distribution provides a groundbreaking new method of encrypting data that ensures complete secrecy via the physics of quantum mechanics. This
88 | 2021-2022 | Broad Street Scientific
PHYSICS
complete secrecy is increasingly important with the rise of quantum technologies that have the potential to break current encryption methods. As such, investigating various methods of QKD is of great value. While the BB84 Protocol is one of the most common and best known examples of QKD, this project investigated a new method, referred to as the Three Basis Protocol, to see if it has any advantages over the two bases of the BB84 Protocol. Further extensions from this project explored the addition of four and five bases, and summary results can be found in Table 1.
Quantum Mechanics. Jones and Bartlett Publishers International, London. UK. [3] Utama, A.N., J. Lee, and M.A. Seidler. 2020. A hands-on quantum cryptography workshop for pre-university students. American Journal of Physics. 88: 1094-1102. [4] Feynman, Richard P., R.B. Leighton, and M. Sands. 1965. Probability Amplitudes. The Feynman Lectures on Physics, Volume III. Addison-Wesley Publishing Company, Reading. MA.
Table 1. Summary of Key and Error Rates for All Simulations Protocol BB84 (2 Basis) 3 Basis 4 Basis
Key Rate
Error Rate
0.501 ± 0.002 0.336 ± 0.002 0.252 ± 0.002
0.248 ± 0.003 0.254 ± 0.003 0.246 ± 0.004
5 Basis
0.201 ± 0.002
0.255 ± 0.005
Since the BB84 Protocol has the highest key rate and there is no additional benefit in eavesdropper situations using the Three, Four, or Five Basis protocols, it appears that the BB84 Protocol remains the best of the protocols to use. Additionally, our proof that the error rate is 0.25 for any number of bases supports this conclusion more generally. The BB84 Protocol is simply the most efficient without sacrificing anything relative to the other setups. It could be argued that there is still value in using additional bases at times, as it adds another layer of complexity that could deter eavesdroppers from listening in. However, while this may be the case, the additional bases also increase the complexity of the setup for the two communicating parties. This, combined with the fact that QKD requires physical systems to be implemented, supports the conclusion that simplicity and efficiency are best for Quantum Key Distribution. 7. Acknowledgements Special thanks to Dr. Jonathan Bennett (NCSSM), Dr. Duane Deardorff (UNC Chapel Hill), the Burroughs Wellcome Fund, and the NCSSM Foundation. 8. References [1] Asfaw, A. et al. 2020. Quantum Key Distribution. Learn Quantum Computation Using Qiskit. http://community. qiskit.org/textbook. [2] Greenstein, George, and A.G. Zajonc. 2006. Quantum Information and Computation. Pages 245-277. The Quantum Challenge: Modern Research on the Foundations of PHYSICS
Broad Street Scientific | 2021-2022 | 89
EFFECTS OF UNDERLAYMENT ROUGHNESS AND ANGLE ON GRANULAR CHUTE FLOWS Noah Siekierski Abstract Granular materials such as sand are ubiquitous, appearing everywhere from beaches to industrial processing. Dense granular flows like landslides can devastate both people and infrastructure. We utilized a chute apparatus to study how the stopping thickness, hstop(θ), varied with inclination angle for 40 grit and 400 grit sandpaper underlayments. For the 40 grit underlayment, we found that flow occurs and sand is retained for angles between 29.7° and 39.5°, as opposed to between 25.1° and 29.7° for the 400 grit. We found that hstop(θ) was a monotonically decreasing nonlinear function for both underlayments. We fit our data using an empirical model derived for spherical particles, and found that it works exceptionally well for the 40 grit underlayment. Furthermore, we observed that flows on the 400 grit underlayment are significantly more susceptible to erosive effects. The results suggest that flows on smoother underlayments may be more sensitive to underlayment irregularities than flows on rougher underlayments, and that hstop(θ) for polydisperse flows is sometimes well approximated by a monodisperse hstop(θ) curve. A possible application of our work is to artificially control the roughness of surfaces along which hazardous flows can occur, which may significantly reduce their impact on surrounding infrastructure and populations. 1. Introduction Granular materials are systems of large conglomerates of particles (Jaeger et al., 1996). An archetypal example of a granular material is sand. Substances like these possess unique properties that do not align completely with any particular state of matter (Jaeger and Nagel, 1992). This lack of adherence to a solid, liquid, or gaseous nature makes granular physics a challenging field of study. It has been pioneered primarily over the last four decades, and as a result, there is still much to be uncovered. Granular flows occur when a stress is applied to the surface of a granular material. This stress will cause the surface layer to move, which will shear the layer below, and so on. Granular chute flows belong to a dense flow regime where interactions between grains are governed by interparticle collisions and friction (Andreotti et al., 2013). A simple diagram of a granular chute is shown below (Fig. 1).
In order for steady flow to occur, the thickness of the flowing material must exceed a minimum value that depends on the angle of inclination of the chute (Andreotti et al., 2013). Once the thickness of the granular material decreases to this critical value, known as hstop or the deposit function, the flow will cease. The function hstop(θ) is known to be a result of complicated boundary effects (Pouliquen 1999). These effects are highly related to the friction between the surface and the granular medium (Andreotti et al., 2013) (GdR MiDi, 2004). The precise way that the boundary conditions for the basal underlayment depend on the roughness of the surface remains unclear, as does the effect of those boundary of those boundary conditions on hstop(θ). Since the mass of granular material retained on the chute surface is directly proportional to the stopping thickness hstop(θ), measuring the mass can be used as a proxy for measuring hstop(θ). Most practical examples of granular flow, such as rock avalanches, landslides, and mudslides, are dense granular flows (Silbert et al., 2001). A better understanding of dense granular flows and how they are affected by different surface roughness conditions is a step towards developing better mitigation strategies for geophysical hazards. 2. Materials and Methods
Figure 1. A diagram of a granular chute flow inclined at an angle θ.
90 | 2021-2022 | Broad Street Scientific
To investigate how the end state of a granular chute flow depends on the angle of inclination and the roughness of the surface, we constructed a chute apparatus (Fig. 2). With the exception of the transparent acrylic walls that border the chute, the device is built of 1.9 cm PHYSICS
plywood. The dimensions of the chute are 110 cm long by 50 cm wide; the purpose of such a wide board is to reduce the effects of the walls. These effects are known to be significant (Pouliquen 1999).
Figure 2. A schematic of the granular chute apparatus used. We used Ottawa sand from GlobalGilson and sieved it between #30 and #45 ASTM sieves of mesh size 600 μm and 355 μm, respectively. We conducted three trials for each combination of underlayment and chute inclination angle. For each trial, we loaded the sand into the reservoir and then opened the gate to allow it to flow down the chute. We used two different underlayments, 400 grit and 40 grit sandpapers, which were attached to the chute via staples on the edges of the board. These grits correspond to underlayment grain diameters of 15.3 to 23.0 μm and 336 to 425 μm, respectively. We conducted trials on both sandpapers across the range of angles where steady, uniform flow could be observed. This angle, θ, could be varied and fixed easily with the use of a car jack. We calculated the angle using trigonometry, finding the uncertainty in the angle to be ±0.1°. We made markings on the vertical board at half-inch intervals, and we used the alignment arm to match the chute with these markings. We loaded the sand into the reservoir using a bin, and then released it by lifting the gate. After the flow concluded in a trial, we massed the sand that remained on the chute using a scale with a precision of ±0.1 kg. All trials were conducted in the same location, with the temperature varying by less than 2°C and the humidity varying by less than 5% over the course of all trials. Humidity, and to a lesser extent temperature, are known to have significant effects on the cohesive properties of granular materials (Andreotti et al., 2013). 3. Results and Discussion We plotted the results for each sandpaper underlayment on a graph (Fig. 3). Using the mass retained, we were able to compute the average thickness of the layer of sand left on the chute, hstop(θ), using the formula hstop(θ) = M/ρA. Here, M is the mass retained, ρ is the density of the sand (measured to be 1800 kg/m3), and A is the surface area of the board (0.55 m2). For the 40 grit sandpaper, we PHYSICS
observed that steady flow began at an inclination of 29.7°. On the 400 grit sandpaper, we observed steady flow to begin at 25.1°. As the angle of the chute was increased for the 40 grit sandpaper, the same fundamental behavior was observed over the entire range of steady flow angles. The flow would begin, some granular material would fall off the chute, and then the flow would quickly come to a clear halt. For the 400 grit sandpaper, however, we observed that for inclinations of 29.7° or higher, the flow would not end at a clear moment. At isolated locations on the remaining sand layer, typically towards the bottom of the chute, sand would avalanche and fall off. This avalanche would create an open space to which the surrounding sand would be pulled by gravity, which could lead to subsequent avalanches. This behavior is the reason for the much larger error bars that appear for some data points in Figure 3: the exact pattern of avalanching was inconsistent and would cause differing amounts of sand to fall off the chute. One explanation of how such inconsistency arises is that the variation in how the sand is loaded, coupled with the irregularities in the sandpaper, led sand grains to travel down the chute differently in each trial. This affected the 40 grit sandpaper less than the 400 grit sandpaper because the grains embedded in the 40 grit underlayment are large, whereas the grains embedded in the 400 grit sandpaper are small.
Figure 3. A graph of the stopping thickness versus the angle of inclination for both sandpaper underlayments. Both horizontal and vertical error bars are displayed, but many are too small to be seen. The blue data points were collected on the 40 grit sandpaper; the orange data points were collected on the 400 grit sandpaper. Each of the black curves represents a fit to one data series. The two curves have very different shapes (Fig. 3). While both show a trend of a decrease in mass retained (and consequently, stopping thickness) as the angle of inclination increases, the fit for the 40 grit sandpaper has a concave up shape, whereas the fit for the 400 grit sandpaper is concave down. It is clear that both of these Broad Street Scientific | 2021-2022 | 91
data sets are nonlinear, which is consistent with previous work that has been done in examining hstop(θ) as a function of inclination angle (Pouliquen 1999). It is unclear exactly what fit should be applied to these data, as no theoretical model exists. A known empirical fit, derived by Pouliquen, is shown below:
According to Pouliquen, θ1 is the lowest angle for steady flow to occur and θ2 is the angle at which hstop becomes zero. D is a characteristic distance that describes how hstop(θ) varies. The best fit for the data for the 40 grit underlayment is achieved with tan(θ1) = 0.572, tan(θ2) = 0.827, and D = 2.1 x 10-3 m. The values for tan(θ1) and tan(θ2) make physical sense, corresponding to θ1 = 29.8° and θ2 = 39.6°. These are close to the observed angles where steady flow starts and where hstop becomes zero, respectively. Assuming an average particle diameter of 500 μm, we find that D corresponds to around 4.2 particle diameters. Applying the same model to the 400 grit data, we achieve the best fit with tan(θ1) = 0.566, tan(θ2) = 0.564, and D = -3.1 x 10-3 m. These values for our fit parameters are physically nonsensical. The corresponding values for θ1 and θ2 are 29.5° and 29.4°, respectively, which do not match with the observed values of 25.1° and 29.7°. Furthermore, θ1 > θ2, which seems to imply that hstop(θ) equals zero at an angle where all of the sand is retained on the chute. This, coupled with a D value of -6.2 particle diameters, suggests that the physical interpretation of these parameters provided by Pouliquen does not apply to the data collected on the 400 grit underlayment. Pouliquen’s model, however, was derived from data collected with a bulk flow made of monodisperse spherical particles and a rough base made of glued spheres. However, sand grains are nonspherical and polydisperse (Fig. 4), which is also true for the particles embedded in the sandpaper underlayment on which they flowed. It is surprising that this model describes our data well, at least for the 40 grit underlayment, considering the polydisperse nature of sand, along with the irregular shape of the particles.
Figure 4. Microscope image of a sample of sand grains used in this experiment. The particles have some roundness, but they are not spherical. 4. Conclusions and Future Work In this paper, we have examined how the deposit function of a granular chute flow depends on the roughness of the underlayment upon which that flow occurs and the angle of inclination of the chute. We conclude that the coarser 40 grit sandpaper supports a steady flow that eventually stops for a significantly larger angular range than what is observed for the 400 grit sandpaper. We are able to use a model originally derived for spherical particles to fit our data, though the goodness of fit is different between the two underlayments. In the case of the 40 grit underlayment, the fit is exceptionally good, which may mean that particle shape is not an important factor in the deposit function. Furthermore, D changes in sign between the two underlayments, which suggests that there may be an intermediate underlayment roughness where its value is zero and a graph of hstop(θ) versus θ yields a straight line. Comparing the fits for the 40 grit and 400 grit underlayments, we see that the former does a significantly better job of modeling its corresponding data set. This may be due to the avalanching behavior that occurred in some trials on the 400 grit sandpaper, but not in those on 40 grit. Future research should be conducted on the effects of particle shape and polydispersity on hstop(θ), as well as on the possibility of employing manmade roughening techniques to alter the amount of material that remains on a sloped surface. Our findings suggest that in places where hazardous dense granular flows like landslides may occur, there may be benefits to artificially changing the roughness of a surface to control the amount of material retained. 5. Acknowledgements The author would like to acknowledge the assistance and mentorship of Dr. Jonathan Bennett of NCSSM, Dr. Karen Daniels of NCSU, and Dr. Duane Deardorff of UNCCH, without whom this project would not have been possible.
92 | 2021-2022 | Broad Street Scientific
PHYSICS
6. References Andreotti, B., Forterre, Y., and Pouliquen, O. 2013. Granular media: between fluid and solid. Cambridge University Press, Cambridge. Groupement de Recherche Milieux Divisés. 2004. On dense granular flows. European Physical Journal E 14: 341-365. Jaeger, H., Nagel, S., and Behringer, R. 1996. The Physics of Granular Materials. Physics Today 49(4): 32. Jaeger, H. and Nagel, S. 1992. Physics of the Granular State. Science 255(5051): 1523-1531. Pouliquen, O. 1999. Scaling laws in granular flows down rough inclined planes. Physics of Fluids 11: 542. Silbert, L. et al. 2001. Granular flow down an inclined plane: Bagnold scaling and rheology. Physical Review E 64: 051302.
PHYSICS
Broad Street Scientific | 2021-2022 | 93
AN INTERVIEW WITH DR. AMAY BANDODKAR
From left to right, top to bottom: Dr. Jonathan Bennett, BSS Faculty Advisor; Melody Lee, 2022 BSS Essay Contest Winner; Vish Ravichandran, BSS Editor-In-Chief; Lucia Wang, BSS Publication Editor-In-Chief; Dr. Amay Bandodkar, Assistant Professor in the Electrical and Computer Engineering Department at North Carolina State University; and Hrishika Roychoudhury, BSS Editor-In-Chief. What are you currently working on in your research? Right now we are focusing on two aspects: wearables and implantables. In the field of wearables, we all know about Apple Watch and all these wearable devices that are out there. Almost all of them monitor parameters, like say heart rate, how many steps you have taken, and approximately calculate how many calories you have burned. But in addition to this, there are a whole host of chemicals that are important to monitor. For example, can these wearable systems monitor the glucose levels, or can they monitor the stress levels that a particular person may be experiencing? To understand human physiology in a more comprehensive way, one would also want to monitor biochemicals that the body produces. What we are trying to do is develop the next generation of wearable devices that will also integrate chemical sensors so we can understand the human body in a more holistic manner. For example, trying to develop sensors that can monitor glucose or can monitor other chemicals like lactate (because it is a good indicator of physical stress). In addition to that, we're trying to monitor small proteins that are present, which could be an indication of inflammation or, say if a person is taking medication, what is the effect of that medication on the human body. If we can monitor these in real time, in a continuous manner, it will be quite useful to understand the human body.
94 | 2021-2022 | Broad Street Scientific
In implantables, presently we are focused more on the neuroscience applications. So we are trying to develop these really tiny devices that can go inside the brain of an animal to study how the brain circuitry functions. It is kind of really amazing that we have made so much progress in the field of biomedical sciences and understanding the human body, but when it comes to the brain we barely know how the brain functions. So there are a lot of question marks when it comes to understanding how the brain reacts to particular conditions, and why it sends certain signals under certain conditions. These are all questions that neuroscience people have, without the tools to answer these questions. We are trying to develop tools that can help us answer these kinds of questions. The key aspect of being wearable or implantable is, how can we make these systems as small as possible. For brain implantable systems, we have to make it small so that it can actually go inside the brain. For wearables you could say, even if the system is big, who cares, you can just put it on the body. But nobody likes to put big, bulky things on the body. We want to make it so small that you won't even realize that the device is on your body. So that's that's the ultimate goal; how do we play with the materials of the device, how do we play with the electronics that are involved, and then the designs that are involved to make them interface with soft tissues in an intimate fashion. Because if you look at conventional electronics, they are made of rigid materials, like silicon, metals, all these things are rigid materials and two-dimensional. On one hand you FEATURED ARTICLE
have these rigid materials, but the human body — be it the skin, the brain, or any organ – it is really soft, delicate and three dimensional. How do you take two dimensional, rigid materials and the soft material like tissues closer without causing any damage to the tissue? So, we play with the mechanical properties of these devices and the material properties to make them soft and stretchable, such that you can easily wrap them around the human body or any organ, without causing any damage. What do you think the potential impact of wearable technologies is and where do you think these technologies are heading? So far for general applications, it could be that I want to monitor my overall health status, and what kind of a quality of life I'm leading. But if you look from more of a medical applications point of view, let's consider the case of a patient in the ICU. Right now, a lot of instruments are connected to the body to measure oxygen levels or heart rate and all these things. But if they want to do some blood analysis, they have to take the blood sample, send it to the labs, and do the analysis. It may take several hours, and depending on which country you're in, it can take several days. But if you can have small patches that stick on the body and they are monitoring your vital signs as well as numerous biochemicals, then that will be very useful because you are going to be getting information in real time, as compared to doing studies and monitoring them once or twice a day. This will be extremely important for critically ill patients. These devices can be used for sports applications as well, for athletes trying to see if they are improving their performance. For people with high stress levels in school or at work, if you want to measure the stress levels and you want to change your life schedule accordingly, these systems can be used for such applications as well. The applications are numerous. What first got you interested in the fields of electrical and computer engineering and working with biosensors? Was there something that specifically sparked your interest, or did you always know that you wanted to pursue these fields? My story is not as romantic as it might be for some people. In my case, what happened was I was doing my undergraduate in India, and my department happened to organize a conference on biosensors. I really wanted to get into research. I did not have any preference of, “Oh, this is what I will do with that”, it was just a coincidence that they had the conference. I attended the entire two day conference and I was really amazed by the importance of biosensors, and the kind of innovation that people were doing in that field. You can develop the world's best medication. But to give that medication to a particular patient, you first need FEATURED ARTICLE
to know that the patient has a particular condition. The biosensors are the first set of devices that one requires to get that information. So this is basically what got me interested in biosensors. While working on sensors as an undergraduate, I started realizing how important wearable biosensors would be, because the conventional systems were always, “Oh you take the sample, and then you may have a handheld device like a blood glucose meter or something.” But a wearable system would really solve a lot of these problems where you have just discrete data points. That's where I got really interested in whether we can take the concepts of conventional biosensors and apply them for wearable applications. To do that, you would need to bring in expertise from electrical engineering, materials science, chemical engineering, and biology. Thus it was a pretty interdisciplinary kind of research, and that's something that really excites me because almost on a daily basis, I talk to people with expertise in electrical engineering, then later I talk to people with expertise in biology. It's really amazing when you get that opportunity to work at that interface of multiple fields. You mentioned speaking with people from all these different fields on a regular basis. What is a typical day like for you at NC State? As a new faculty with a lab that has been evolving, I will say a good 20-30% of my time is in the lab with students, seeing how they're doing, and trying to work with them to find a solution. And then the remaining time is basically split between teaching courses and also writing proposals and grants, and then networking with other professors to try to find new ideas to explore. One thing that I want to point out is that in the field of research, a lot of the time, nobody knows the solution. For example, when I go into my classroom to teach and give my students a set of problems, I know the solutions to those problems. If the students don't know, they come and ask me: “Hey, I tried this but I could not solve it, can you give me the solution?” But when those same students come in and are working in the lab with me, they have the tendency to think, “Oh, if I don't know what to do, the professor should know the solution.” However a lot of the time, even I don't know the solution. And that's the beauty of research, right? Because we are trying to explore new things. So there are a lot of failures, and this is something that a lot of the students who are new to research find really unique there. There's no textbook solutions to the problems, because if there is, then then that's not research, because then people already know how to do it. That's something that really excites me, just thinking about, “Oh, what could be the solution, and how can we come up with a solution?”
Broad Street Scientific | 2021-2022 | 95
What are some questions or challenges that you face in your research and how do you overcome them? The challenges are always dependent on the project, they are pretty unique. As I said, my research is pretty interdisciplinary. Perhaps you're trying to build an electronic system to understand the brain. So we need an electronics engineer and a neuroscientist to work together. There are certain challenges on the neuroscience part and then there are certain challenges on the electronics part. If you want to build a system, you need to find a solution that will somehow navigate both the problems, and sometimes the problems can be really in opposite directions. I can make the best electronic system, but the biology will not be able to accommodate that kind of system. If I try to make it compatible with the biology, the electronic system properties may be really terrible. So how do we find that middle ground? Another problem is sometimes when these two completely different people with completely different expertise try to talk to each other, communications can be a problem. For example, the neuroscientist/biologist may see a problem, but the electronics person may not be able to fully appreciate the severity of the problem, because for them, the biology is a complete black box, and for the biology person, electronics is a complete black box. So it takes a good amount of time to understand what are the challenges in the other field and try to solve problems accordingly. When it comes to solutions, as I said, we sometimes just have to do trial and error. We go, “We tried it and it did not work, why did it not work?” I always ask my students to look at it in a broad way, don't just narrow your approach to a narrow range of solutions, because a lot of time you have to go in a completely different field to find a solution. So for example, we're trying to develop a wearable sensor for a particular chemical and we're trying to figure out what would be the best way to make that sensor. We read hundreds of papers in the field of sensors to find potential solutions to our problem, but we could not. Instead, we were able to find a solution from some of the papers that were actually for drug delivery, which is completely different. We were able to get some inspiration from these drug delivery papers, and we were able to modify it for sensing applications. That is something that I always tell my students: do not restrict yourself to, “Oh, I am a sensor guy so I'm only going to read papers on sensors.” Sometimes you have to go in a completely different field to bring in some knowledge and solutions. What is your favorite part about your research? I would say exploring new things, things that people have not done. And also, the fact that almost every day, you’re 96 | 2021-2022 | Broad Street Scientific
faced with new, unknown challenges. So, my life is never boring. It never seems monotonous to me because every day is a new challenge, so I know there’s always something new to think about. And since my lab works on multiple projects, if I sometimes get overwhelmed by one problem, then I can give it some rest. Let me think about the other problem that is there in the other project. That is something that really keeps me excited. And also the fact that being in academia, I’m always surrounded by young, enthusiastic, self-motivated students. I’m much older than you guys, but in the academia field, I’m still considered as a “young faculty”. So I’m super enthusiastic right now, but ten years, twenty years down the line, I may feel like “Oh, I’ve been a professor a long time and its novelty has worn out”. But that’s not something that’s going to be the case because you’re always surrounded by these super enthusiastic students who want to try new things. And if I kind of get bored sometimes, just being in this enthusiastic environment boosts up my energy. So that’s also something that keeps me excited about my work. What areas in technology and software do you think are changing the most right now? What will be different in the next ten years or so? That’s a good question. I wish I had a crystal ball to see how things would evolve. In research, things evolve pretty fast, so it’s kind of hard to say how my lab is going to look ten years down the line. I can say this: ten years ago, I was doing research that was not even close to what I’m doing right now. So, I am pretty confident that ten years from now, I’ll probably be doing something that is completely different, but building up on whatever experience I’ve had. The way to explore new things is to build on your expertise. Whatever expertise that I’m building right now, I’ll be applying it to different fields. And I can say that even presently, we are doing work on wearables and implantables, but that’s not the only thing I’m planning to do. I’m trying to explore new things as well. For example, there’s a lot of interest in stem cells, of growing tissues, like brains, in petri dishes. People are actually trying to do that using stem cells. We are collaborating with groups that do this really fancy work, and we are trying to develop sensors to understand how the brain actually grows from an individual cell to a complete organ. So we are trying to develop sensors for these types of applications as well. And this is something that we have never done in the past. We know that this is a new thing, and we want to see if we can build on our experience in wearables and implantables for this kind of stem cell research. We are already kind of going in new fields and exploring new things, and I’m pretty sure five years, ten years down the line, it will be completely different from what we are doing presently.
FEATURED ARTICLE
How has COVID-19 specifically affected your work and your research goals as of now? It has certainly affected quite a lot just because of the restrictions that were imposed on the university and how many people can be in the lab. We really had to make sure that we are dividing the time to make sure that people are doing social distancing and the minimum number of students were there in the lab at a given time. It was challenging, but at the same time, we tried to make the most of it. COVID actually kind of made us more efficient because we knew that we had a short amount of time to spend in the lab, so instead of complaining, why can't we make the most of it? Let's try to make our processes more efficient. This is something that's going to help my lab a lot, because we now know that we can work in a lot more efficient way. So, we are trying to see the silver lining in the restrictions that are imposed on us. What do you like to do outside of work? Is there anything in particular that has been fascinating you recently or that you’ve gotten into? If I'm not in the lab or working, I like to go hiking. I'm really glad that I'm in the Raleigh area so there are a lot of opportunities to go hiking. That also helps me come up with new solutions and new ideas to the problems that I face in my work. A lot of the times when I'm on these hikes or walks or biking or whatever I'm doing, in the back of my mind I'm still thinking about problems and solutions. Sometimes I’m suddenly stricken with a potential solution to the problem. Then I’ll send an email to my students and I’ll say “Hey, I was thinking about this and, you know, maybe we should try this. It might work”. And a lot of times that actually helps, so I always encourage my students to recognize that you spend time in the lab, but sometimes you need to get out of the lab to find solutions to the problems in the lab. That's really important. If you were to go back to your high school years or your early undergraduate years, what would you like to change about them? I'm not sure, but I do have some suggestions to young students based on my experience. I did my undergraduate in India, and the opportunities that are there in India and other developing countries are very limited. So, as a high school student, as an undergraduate student, I used to read about all this innovative stuff that's happening by researchers in India and outside in the world. But I never had that opportunity to spend in the lab, to get hands on experience because there were so few opportunities, and so few labs that could actually give these kinds of avenues for such young students. And even when I was an undergraduate, I remember I used to go to a completely different FEATURED ARTICLE
city during summer breaks and winter breaks. Instead of going home, I'm going to go to the lab, which is a thousand miles from my place, just so that I could spend some time to actually get some hands on experience. I feel, in the US, it's relatively easier to get those kinds of experiences. And sometimes, I feel students don't take the full advantage of the opportunities that students have in the US. So, I would say: make the most of it. I always say that to my undergraduate students and high school students, “If only I had these kinds of opportunities when I was in high school or an undergraduate student. I lacked those kinds of opportunities. But you guys have that, so make the most of it.” In the end, you may hate it, but at least you will know that you hate it because it's only through experience that you know whether you like something or don't like something. The worst thing is, in the future, regretting not having that opportunity and saying “I wish I would have tried that when I was in my high school or in my undergraduate, at least I would have known whether I liked it or not." Having had a few years of experience in your research career, how has your career potentially been different than how you might have imagined it to have been when initially going into your career? A lot. I mean, as a professor, it’s a lot more challenging than what I had expected it to be because from day one, people expect things from you. And as a faculty who is mentoring students, I feel responsible for how their life is going to turn out. When I was a student, I was like “Okay yeah, whatever decisions I make, I'm responsible for my own future”. But now, as a mentor of students who work in my lab, all the experiences that they have, I feel like I'm playing some role in how their life is shaping so that makes me a little bit nervous sometimes. I want to make sure that I give them the best opportunities that they can have to make the most of their time in the lab. And also, in writing proposals to get the funding to do research, that's always challenging. So, yes, it's challenging for sure, but so far so good, I'm enjoying it. If you could tell your past self or a student interested in the pursuit of research or STEM something, what would you say? I would say: time is precious. Know what your goals are and pursue them. Do not be shy of failing. You’ll fail multiple times, but in every failure, you learn things. Make the most of whatever time you have because everyone has the same amount of time in a given day. Be focused, try different things, and don't be afraid of failure.
Broad Street Scientific | 2021-2022 | 97