Automating Biology: Scalable Web Infrastructure for the DAMP Lab Workflow Manager Alden Carter Senior Thesis | 2026
1
Automating Biology: Scalable Web Infrastructure for the DAMP Lab Workflow Manager Alden Carter and Kelsey Liu March 2026
1
Introduction
In an era where engineers can simulate rocket launches, airflow dynamics, and entire virtual environments from a computer, biology experiments continue to require manual operations one at a time. A biologist conducting an experiment to sequence a genetic circuit is constrained by hands-on benchwork. Despite advanced practices in biotechnology, sequencing and DNA assembly can still require researchers and lab technicians to manually pipette miniscule volumes. These limitations represent the bottleneck in the design-test-build cycle in which scientists and engineers have yet to reach the full potential of intersecting hardware, software, and wetware [1]. In recent years, efforts in biodesign automation have worked to close this gap. Biology research is still slow, manual, and labor-intensive, relying on trial-and-error rather than predictable engineering workflows. The root of this inefficiency lies in a fundamental disconnection: software, lab hardware, and biological designs operate in isolation, making experiments hard to automate, scale, or reliably reproduce. Oliveira and Densmore characterize this fragmentation as a barrier to the codesign environment necessary for synthetic biology [1], while Miles and Lee emphasize that without closed-loop automation and standardized protocols, even well-designed experiments suffer from reproducibility challenges [2]. As a result, the design-build-test cycle is bottlenecked, limiting the speed and
2 reliability of innovation in synthetic biology. To address these challenges, facilities like the DAMP Lab have emerged as "cloud labs" which is basically a lab that is accessible entirely online. Instead of manually performing each step, a client submits requests for services such as DNA assembly, cloning, and PCR reactions through a web interface, and the DAMP Lab handles execution automatically. This approach excels at handling routine tasks with consistent procedures, bringing unprecedented standardization to biological workflows, though Miles and Lee note it remains less suitable for exploratory research requiring frequent adjustments. The underlying philosophy draws from bio-design automation and modularity, where orders are created using standardized biological "modules" and automatically carried out. Densmore and Bhatia describe this convergence of software, biology, and robotics as essential for enabling more predictable, reproducible, and scalable experimentation [3]. One of the limitations of these automated labs is the barrier to its adoption; many biologists have little experience with integrating automation systems with their experiments. Platforms such as DAMP Lab are being developed to allow scientists to design experiments in these automated and scalable ways without needing to learn how to code [2]. Our project is at this interface. Among the many projects supported by SAIL, the DAMP Lab represented a compelling introduction to the intersection of software engineering and biology. The engineering principles of modularity, abstraction, and automation link directly to these limitations in biology experiments that the lab infrastructure aims to address. The prototype of the original workflow ordering website lacked several critical features for a scalable cloud lab. Although the website displayed services of a biology experiment and workflow creation, many of these features were hardcoded directly into the interface,
3 meaning they are not flexible to user input and are therefore unsuitable for customization and updates. And the backend had no authentication safeguards to protect sensitive operations. Thus, these limitations motivated the central question of this work: how can software engineering principles help with scientific discoveries by improving speed, reliability, and automation of the biological experiment processes? To investigate this learning and research question, our contributions to the DAMP Lab cloud workflow manager focused on extending and refactoring its codebase into a secure, modular, and scalable system using modern full-stack practices. The methods section that follows outlines the technical foundation of this work.
2
Methods
To contextualize the technician foundation of this project, this section first outlines a modern full-stack web application. A typical stack consists of three interdependent layers (Fig. 1), each responsible for a different stage of interaction between the user, website, and long-term data storage [4].
Fig. 1: Layers in a full-stack application [4]
4 The frontend is responsible for handling every component on the website itself that the user interacts with, such as layouts, buttons, forms, and animations. These interactions are kept track of using “states,” representing the current status of the interface. A client program handles these state-based interface updates, commonly built using HTML, CSS, and JavaScript frameworks such as React. The main challenge in this layer is balancing visual design clarity and performance. The developers’ job is to figure out ways to keep interfaces responsive, efficient, and accessible across different devices and browsers [4]. The backend manages the application’s logic, which controls the request-response cycle, authentication, and communication between the user interface and database where long-term data is stored. This server layer, functionally an intermediary, ensures that a user only interacts with data they are permitted to access, whether intentionally or unintentionally. This includes preventing unauthorized access, verifying that incoming requests are legitimate, and protecting against code-injection attacks to these communication points between the server and client [5]. Backend architecture is achieved through Application Programming Interfaces (APIs), which allow and define structured ways for different software applications to communicate and interact with each other [6]. Runtime environment tools (the run code language outside of the website itself) such as Node.js are the tools that implement APIs. With the existence of many such frameworks, choosing the backend runtime environment depends on considerations such as data-handling, scalability, and security [7]. For instance in our project, we used Node.js for its consistent Javascript in both frontend and backend, as well as asynchronous handling of requests such that instead of waiting for operations to complete one at a time, it executes other tasks until a callback [7]. Security authentication forms one of the most crucial
5 responsibilities of the backend. Default security methods implement encrypted passwords through hashing their stored values and token-based authentication (such as JSON Web Tokens). These features prevent unauthorized users from viewing or altering stored confidential data. At the same time, backend development revolves around scalability, in ways that ensure the system behind a website connected to various servers can handle large volumes of simultaneous requests securely, all while being able to be easily augmented and transformed into new system design specifications. The primary purpose of a website is to retrieve and store information that a user wants from the service, such as logins, updating interactions forums, and placing a product order when shopping online. The database organizes information such as users’ information, passwords, and history in the system’s long-term memory. Each piece of data is stored in a structured format such that different types of information can be retrieved uniformly [4]. Thus, the primary challenge in this layer is maintaining data with consistency and integrity across the entire application. This involves validating data before it is saved, enforcing access rules, and storing information of related components of the system in a standardized way. Before initiating any work on the DAMP Lab project, we started with a learning phase to ensure that we had a working understanding of the technologies in the full stack. Our initial step was to complete an introductory tutorial on building a MERN-stack web application. The MERN stack consists of MongoDB as the document database, Express as the server-side application framework, React as the client interface, and Node.js as the runtime environment that executes JavaScript on the server [4]. Although this tutorial did not use NestJS, it established the core concepts of the frontend, backend, and client–server interaction
6 that the later system relies on. After completing this preliminary exercise, we transitioned to a second exploratory project: the NestJS “cats website” tutorial. This stage mirrors the full technology stack used in the DAMP Lab platform and introduces framework-specific components such as controllers, providers, data transfer objects (DTOs), and dependency injection. We additionally practiced building a basic frontend interface without relying on Tailwind CSS, since the DAMP Lab platform uses plain React for its user interface. Tailwind CSS is a styling framework that provides simple utility classes for layout and design, but for this project we focused on writing our own components and styles directly in React to match the lab’s existing conventions [8]. We then moved beyond the basic tutorial and began learning how to integrate authentication into our system by using NestJS’s AuthGuard to implement secure login and account-creation workflows. This introduced us to concepts such as protected routes, JWT-based authorization (JSON Web Token), and role-restricted access within the application [5]. In addition, we incorporated Mailgun, an email API that enables automated message delivery, to send verification emails and other notifications to users. Finally, we deployed our applications to Amazon Web Services (AWS), which involved configuring EC2 instances, managing security groups, and setting up environments that mirrored real-world deployment conditions. Version control remains integral to our workflow throughout development, and a significant part of our preparation involved learning how to use Git effectively. Git provides a complete record of code changes, enabling us to compare earlier states of the application and revert breaking modifications when needed. Its branching and merging model allows multiple developers to work simultaneously without overwriting each other’s work, which is
7 essential in a collaborative research environment [9]. Branch isolation also further supports experimental development, enabling us to prototype ideas safely and integrate them only after they are fully tested. We also adopted a workflow in which new features were submitted through pull requests, allowing mentors to review code, provide feedback, and ensure quality before changes were merged into the main branch. In addition, we learned how to configure. secure environment variables so sensitive information – such as API and database keys – remained local and was not exposed in the public GitHub repository.
3
Results and Discussion
3.1
UI States and User-centered Design
Users routinely dismiss warning dialogues without reading them, especially long dense texts such as conditions and agreements. A redesigned user interface was implemented to improve the overall site’s visual coherence, reduce interaction ambiguities, and communicate action severity more effectively to users. This redesign forced on design principles regarding color usage, as well as user convenience with layout consistency and visual hierarchy. In the original interface, warnings and high-impact actions were not consistently distinguished from routine interactions. As a result, critical messages could be missed, and irreversible actions could be triggered without sufficient visual emphasis. The revised interface introduced a standardized primary-secondary-tertiary-accent color system using the React framework’s theme configurations. Primary interface elements were recolored to match the project’s logo palette, establishing a
8 consistent visual identity as well as brand presence across pages (Fig. 2). This change improved recognizability and reduced cognitive load when navigating different sections of the application. Importantly, visual emphasis is used to slow user interaction at appropriate times in order to draw attention to consequences (Fig. 3). Red is commonly reserved for destructive actions across web applications. This convention was intentionally adopted in the redesigned interface to reduce accidental activation of irreversible operations, such as deleting entire intervals of a user’s work [10].
Fig. 2: Streamlined visual consistency and coloring attention
9 Fig. 3: Visual hierarchy and areas of cognitive load concentrations in a web application [11] From a system perspective, the frontend redesign functions as a critical communication layer without affecting the control layer. No additional backend validation logic was introduced as a part of this change. These surface-level design decisions demonstrate how visuals contribute to system reliability. By improving clarity with consistency, the interface reduces the likelihood of user error while maintaining flexibility and responsiveness. Checkout Pages The checkout page revamp redesigned the order submission pages to create a more intuitive user experience and added features. The updated checkout interface now dynamically pulls service prices from the backend rather than displaying hardcoded values. Underneath the summary it also includes a prominent disclaimer clarifying that displayed costs are estimates only, with final pricing communicated via email after lab personnel review the order (Fig. 4).
Fig. 4: Final checkout page before submitting the job
10 The final checkout page received significant improvements including better styling for visual consistency to match the previous checkout page, navigation protection to prevent users from accessing the submission page without properly checking out, and the removal of redundant payment processing fields since transactions are handled externally through email. Personal information fields were also eliminated because this data is already captured during account creation. In their place, new functionality was added allowing users to title their jobs, view a summary of their account contact information, and provide additional notes with their orders. This feature addresses critical usability issues that existed in the previous checkout system while streamlining the entire order creation process. The original checkout page required users to repeatedly enter personal information that was already stored in their accounts, creating unnecessary work and potential data inconsistency. More significantly, although a job submitted summary page existed in the codebase, it was previously inaccessible through normal user workflows. The redesign integrates this page into the checkout process, automatically navigating users to it after job submission. The page displays created job details, submission timestamps, loading status, and validation warnings for invalid IDs (Fig. 5). Enhanced with a complete order document and PDF download capability, the system provides clients with immediate tracking information similar to commercial platforms like Amazon. This transparency gives clients referenceable documentation for communicating with lab personnel while improving overall order processing clarity.
11
Fig. 5: Job submitted summary page Several improvements could further refine the checkout and order management experience. Implementing automated email receipts functionality would ensure users receive a copy of order details immediately upon submission. A dedicated user dashboard displaying all past and pending orders would give clients a centralized view of their order history rather than requiring them to scroll through their emails. Additionally, introducing order editing capabilities within a limited timeframe before lab processing begins would allow users to correct mistakes or update requirements without needing to cancel and resubmit entire orders, reducing friction for both clients and lab personnel. 3.2
Authentication Flows
API Endpoints Every interaction with a website begins with a request sent to a URL. Web applications such as the DAMP Lab tool regularly
12 handle thousands, up to millions of concurrent requests, naturally leading to some developers have not expected. When a user clicks buttons, submits forms, etc. their browser sends a structured message to a specific address on the internet [13]. From the perspective of the server, there is no difference between a legitimate user and a malicious attacker. Beyond standard user interactions, hackers have a variety of methods to gain unauthorized access to application permissions, such as manually injecting, modifying, or replaying requests to overload or crash a system [14]. API endpoints are the points of exposure between the system and incoming traffic, motivating defensive API design to filter the legitimate ones.
Fig. 6: AuthGuard protocol from frontend to server [12]
13
Fig. 7: Unguarded endpoints through URL access before protecting The application is composed of multiple subpages and service routes, which sensitive data, such as login passwords, flows through (Fig. 6). To protect data in transit, all routes are served over HTTPS (secure HTTP), rejecting requests made without encryption, preventing potential malicious injections or modifications of transmitted javascript data to the server. Browser warnings such as those displayed by Cloudflare warn websurfers of increased risks due to the absence of HTTPS or guarantee of encrypted data. Direct access to protected routes, such as admin-only features, was restricted using an authentication guard (AuthGuard). Before implementing this, any demo user without special permissions would have been able to access pages only intended for admin use. Fig. 7 shows how URLs are discoverable by design, meaning they are not limited by traversing legitimately through a button. By giving a user a time-expiry token embedded with different access permissions upon logging in, the AuthGuard checks incoming requests, such as visiting a subpage URL, for valid tokens before reaching the backend [15]. The formatted requests are implemented through GraphQL rather
14 than REST API, a type of architectural style that uses HTTP to transfer requests, because the schema specifies available data types, queries, and mutations. The tradeoff considered was that, compared to REST, GraphQL requires explicit maintenance in order to reduce over-fetching by returning only requested data fields [6]. In practice, this constraint limits the amount of information revealed in each response and reduces the effectiveness of malicious probing. A practical observation emerged from this: interface-level restrictions are futile, since any request visible in a browser’s network inspector can be manually replicated. For this reason, the API layer enforces all access control decisions instead of depending on the data’s visibility on the frontend. Announcements Feature The announcements feature provides administrators with a streamlined way to communicate important information to all lab users directly through the website's home page. The implementation includes an admin-only button that navigates to a dedicated announcement page where administrators can view the current announcement, create new announcements using a text box, and see a note on supported Markdown formatting options including bold text, italics, and headings (Fig. 8). Once submitted, announcements appear on the home page for all users with a timestamp indicating when they were posted. Administrators also have the ability to hide the current announcement when it is no longer relevant.
15
Fig. 8: Announcement page feature The backend implementation involved creating an Announcement model to define the database schema and structure, along with GraphQL operations including a createAnnouncement mutation for posting new announcements and an announcements query for retrieving existing messages. On the frontend, the system integrates ReactMarkdown for rendering formatted text and establishes the necessary queries and mutations (CREATE_ANNOUNCEMENT, GET_ANNOUNCEMENT, UPDATE_ANNOUNCEMENT) to communicate with the backend, all accessible through a new "/edit_announcements" route. This feature significantly improves the user experience by enabling direct, organized communication between the lab and its users. Previously, no formal announcement system existed on the
16 website, creating a gap in how the lab could share time-sensitive information or updates with the community. Because the organizational structure of the codebase was already well-established, implementing this feature proved relatively straightforward – adding the new route, page, queries, and mutations integrated smoothly into the existing architecture. One key design decision was to retain all announcements in the database rather than overwriting previous messages, which preserves a complete history of lab communications. However, this historical data currently has no interface for access within the website itself, representing an opportunity for future improvements. Another potential addition would be adding a preview function that allows administrators to see how their formatted announcement will appear before publishing it to all users. 3.3
Data Integrity and Scalability
Services and Workflows The service pricing feature replaced a rigid, hardcoded pricing system with a flexible, administrator-controlled approach. Previously, all services were hardcoded – meaning their prices were fixed directly in the source code at $100 regardless of the actual cost of each service. This implementation added a new price column to the admin edit page, allowing administrators to set individual prices for each service based on their actual costs. The backend modifications involved updating the service model to include an optional price field with a float data type, which accommodates both services with known approximate costs and those with variable (unknown) prices. On the frontend, the price field was integrated into several GraphQL operations including
17 CREATE_SERVICE, GET_SERVICES, and GET_JOB_BY_ID mutations and queries, establishing the connection between the database-stored prices and what administrators conFig. through the interface. The EditServicesTable component received a new price column, and the price information was also incorporated into the CanvasType definitions for both services and nodes within the workflow visualization (Fig. 9).
Fig. 9: Admin edit page showing the service table with the newly added price column This addition addresses critical needs for generalizability and scalability in the lab's operations. While services already existed in the system, updating their model and corresponding frontend queries enabled more accurate representation of actual service costs rather than relying on placeholder values. The design decision to make pricing optional reflects practical operational considerations: customers do not complete payment transactions directly through the website, and some services have variable costs that depend on specific project details provided during consultation. This flexibility aligns with biological modularity
18 principles described by Densmore and Bhatia, where modular, adaptable components can be conFig.d and recombined for different applications [3]. By doing this improvement, the system becomes more maintainable and better suited to accommodate the lab's evolving service offerings and pricing structures without requiring code modifications for each change. Bundles and canvas overhaul Previously, workflows and related services employed hardcoded logic; adding or modifying a workflow required changes across multiple backend and frontend files, increasing the risk of producing compilation error and limiting scalability. To address this, bundles were restructured as explicit data objects that encode relationships between services in a custom user-designed workflow in the Canvas builder (Fig. 10). The advantage of users being able to create their own bundles is plenty. Critically, bundling services enables reproducibility across experiments by becoming a ready-made cloud workflow for other researchers to modify or replicate while eliminating inconsistencies posed by human error [3]. Each bundle contains metadata linking its constituent services, which enables consistent propagation of operations in the experimental workflow. In the database layer, overhauling this feature allows storage, queries, validation in future updates to reduce developer overhead and human error as needs scale.
19
Fig. 10: Bundles relationship to Services and Workflows [16] The Canvas interface visualizes how the services are connected. Each bundle is rendered as a distinct grouping, which supports both linear and tree encoded storage structures. A bundle with a linear workflow is shown (Fig. 11). Linear workflows represent sequences of services intended to execute in order, and are suited for simple and fixed procedures.
Fig. 11: Restricted Linear Data Structure of a Bundle
20 On the other hand, if an experiment requires services that lead to conditional or parallel experimental steps (Fig. 12), a tree workflow is important for achieving this option despite consuming more memory and being more difficult to encode. With this flexibility, the system can support future automation, reporting, and validation without altering existing code.
Fig. 12: Tree Data Structure of a Bundle While the bundle and canvas system improved the Workflow Manager’s organization and flexibility, it reveals several limitations in scalability. Linear workflows are predicted to scale predictably, but tree-encoded workflows require additional computation to propagate across branches, which exposes inefficiencies in rendering and storage. Further, some parts are still rigid and repetitive because of the rendering limitations for each node and edge of a bundle. If redesigned, the system could separate layout computation from data logic entirely using a generalized graph-processing approach to support arbitrary workflow structures. Future systems should store workflows
21 consistent with the bundle structure in the database from the start to reduce the translation complexity between these conceptually uniform data objects.
4
Conclusions
This work demonstrates how software engineering principles can directly accelerate scientific discovery by improving the speed, reliability, and automation of biological experimentation. For the DAMP Lab, the contributions outlined in this project transform a prototype workflow manager into an almost production-ready cloud lab platform. The announcements feature establishes clear communication channels between lab personnel and users, the flexible service pricing system enables accurate cost representation without code modifications, and the redesigned checkout process provides clients with professional order tracking and documentation. These improvements address the fundamental challenge introduced earlier: bridging the gap between software and biological experimentation to create the integrated codesign environment necessary for scalable synthetic biology. Beyond the technical contributions to DAMP Lab, this experience provided invaluable professional development in collaborative software engineering, including practical experience with modern full-stack technologies (MERN stack, Nest.js, GraphQL), version control workflows through GitHub, and the communication skills essential for team-based development – expertise directly applicable to future internships and careers at the intersection of computation and biology. The DAMP lab represents just one open application of software
22 engineering in biological research, but the broader trend is clear: biology is rapidly becoming data-driven, creating expanding opportunities for computational experience. Carbonell, Radivojevic, and Garcia Martin identify three critical areas where software engineers are essential to advancing synthetic biology [17]. First, high-volume data from modern sequencing and screening needs computational pipelines for analysis. Second, large-scale automation relies on software to coordinate robotics, sensors, and workflows, requiring engineers who understand both biology and technical systems. Third, combining synthetic biology, machine learning, and automation allows closed-loop experimentation where software analyzes results and automatically designs follow-up experiments, speeding up the design-build-test cycle [2]. In short, software has become central to biological research. Future work on the DAMP Lab could focus on compatibility with other cloud laboratories. Integrating APIs from platforms like Benchling would enable cross-platform automation workflows, allowing researchers to design experiments in one system and execute them in another. This capability is particularly important for open-source workflow managers like DAMP Lab, as it would create a network effect where multiple labs contribute to and benefit from shared automation protocols and standardized biological modules. Such integration represents a step towards the fully realized biodesign automation ecosystem envisioned by Densmore [1]. Lastly, we would like to thank Zoe for her patient guidance and countless hours spent debugging alongside us, Asad for organizing our projects and providing detailed specifications that kept our work focused, Dr. Tomlinson for creating the opportunity to learn and contribute at SAIL, and Dr. Karnaukh for expertly coordinating the program!
23
References [1] S. M. D. Oliveira and D. Densmore, “Hardware, Software, and Wetware Codesign Environment for Synthetic Biology,” vol. 2022. in BioDesign Research, no. 2022, vol. 2022. Elsevier BV, Sept. 24, 2025. doi: 10.34133/2022/9794510 [2] B. Miles and P. L. Lee, “Achieving Reproducibility and Closed-Loop Automation in Biological Experimentation with an IoT-Enabled Lab of the Future,” vol. 23, no. 5, pp. 432–439, Oct. 2018, doi: 10.1177/2472630318784506 [3] D. M. Densmore and S. Bhatia, “Bio-design automation: software + biology + robots,” vol. 32, no. 3, pp. 111–113, Mar. 2014, doi: 10.1016/j.tibtech.2013.10.005 [4] MongoDB, “MERN Stack Explained,” MongoDB. Available: https://www.mongodb.com/resources/languages/mern-stack [5] Auth0, “JSON Web Token Introduction - jwt.io,” JSON Web Tokens - jwt.io, Nov. 30, 2024. Available: https://www.jwt.io/introduction#when-to-use-json-web-tokens [6] “GraphQL: A query language for APIs.,” graphql.org. Available: https://graphql.org/learn/ [7] “Documentation | NestJS - A progressive Node.js framework,” Documentation | NestJS - A progressive Node.js framework. Available: https://docs.nestjs.com/ [8]
React, “Quick Start,” react.dev, 2024. Available:
24 https://react.dev/learn [9] “Git - gittutorial Documentation,” Git-scm.com, 2019. Available: https://git-scm.com/docs/gittutorial [10] Naveen Chikkanayakanahalli Ramachandrappa, “SOLID Design Principles in Software Engineering,” International Journal of Computer Trends and Technology, vol. 72, no. 9, pp. 18–23, Sep. 2024, doi: https://doi.org/10.14445/22312803/ijctt-v72i9p104. [11] “A guide to heat maps for website and mobile app analytics,” Smartlook. https://www.smartlook.com/heatmaps-guide/ [12] K. Carpenter, “Angular: Route Authentication and Guards,” Everything Full Stack, Apr. 24, 2017. [13] MDN Contributors, “HTTP,” MDN Web Docs, Aug. 03, 2019. https://developer.mozilla.org/en-US/docs/Web/HTTP [14] T. Sasi, A. H. Lashkari, R. Lu, P. Xiong, and S. Iqbal, “A Comprehensive Survey on IoT Attacks: Taxonomy, Detection Mechanisms and Challenges,” Journal of Information and Intelligence, vol. 2, no. 6, Dec. 2023, doi: https://doi.org/10.1016/j.jiixd.2023.12.001. [15] “Server Administration Guide,” www.keycloak.org. Available: https://www.keycloak.org/docs/latest/server_admin/index.html [16] J. Grabis and K. Sandkuhl, “Value-Based and Context-Aware Selection of Software-Service Bundles: A Capability Based Method,” Complex Systems Informatics and
25 Modeling Quarterly, no. 10, pp. 21–37, Apr. 2017, doi: https://doi.org/10.7250/csimq.2017-10.02. [17] P. Carbonell, T. Radivojevic, and H. Garcia Martin, “Opportunities at the Intersection of Synthetic Biology, Machine Learning, and Automation,” vol. 8, no. 7, pp. 1474–1477, July 2019, doi: 10.1021/acssynbio.8b00540