I, Science Issue 51: The Rest of the World (Spring 2022)

Page 28

biology & medicine

Preserving the world’s heritage with DNA THOMAS HEINIS , SENIOR LECTURER IN COMPUTING

O

ur cultural heritage is defined by texts such as Euclid’s books, Shakespeare’s works and the paintings of old masters. Preserving this heritage is key as it defines our history and shapes who we are as individuals and as a society. Many artefacts we produce today are digital in nature already. Others, such as Newton’s papers, are scanned and digitised for public view. The preservation of this digital data, however, is at risk since the price to store all artefacts is simply too high – a particular problem in the developing world. Ongoing research at Imperial College is working on a solution to this problem by looking inside ourselves – what better way to store data than the molecule that encodes life itself? What if the answer lied in our DNA? Today’s storage technology revolves around hard disks and magnetic tapes which are ill-suited to preserve important data over decades. Disks and magnetic tapes are shortlived, lasting only 15 to 20 years. This means that if we are to preserve data for future generations, the data must be copied between storage media in a costly process every few years. In theory, magnetic tape could store data for over 20 years, but the incompatibility between tape generations means that migration is necessary every six to 10 years. Our data is, in fact, held hostage in magnetic tape by a progressing technology. What’s more, the process of regular data migration also produces substantial electronic waste. Our current storage technologies are not only obsolete, but also require a lot of storage space and energy. A digital data library that contains an entire culture’s history would require significant and stable infrastructure. Apart from the land mass required to set up a 24/7 functional facility to house all these tapes and disks, a constant supply of power is needed to continuously cool and maintain them, even when data is neither written nor read. In addition, hard drives must be spun occasionally to keep them operational, and magnetic tape requires spooling. The need for strict environmental conditions and upkeep adds to the cost and leads to considerable carbon emissions. A promising solution to this problem is the use of synthetic DNA as storage media. Arbitrary data is encoded in a sequence of nucleotides (the building blocks of a DNA

28

molecule) through a process called DNA synthesis and read back through DNA sequencing. What makes synthetic DNA attractive as a long-term storage media is its longevity: DNA can last for several hundreds of years even under minimal environmental conditions. Since it is so durable, minimal cooling is needed to preserve the data. Its longevity also means that no migration between storage media is needed, thereby reducing cost and waste. The building blocks for DNA are also abundantly available and we already know how to read it back. The process of reading data back from DNA has eternal relevance since we will always be able to read DNA, which is especially important

"All these are important for the existence of data libraries in places where infrastructure conditions are unstable, such as in developing countries." when recovering culturally significant data in the future. Finally, DNA promises an incredible density of information, with only ~15 atoms per bit. This means that a digital library would require less land mass and would be easier to physically move. It could also survive on a limited amount of power for a longer time. All these are important for the existence of data libraries in places where infrastructure conditions are unstable, such as in developing countries. At Imperial College, we are working on all aspects of DNA data storage. These include mapping 0s and 1s into DNA sequences (similar to audio and video codecs); the writing of it (synthesis); the way it will be han-

dled and curated; its sequencing; and finally mapping it back to binary. The synthesis of DNA was until recently the domain of pharmaceutical and life sciences. As such, the cost of synthesis has been a major roadblock for the adoption of synthetic DNA as a storage medium. To reduce the price of synthetic DNA for data storage, we need to re-engineer DNA synthesis. By taking a co-design approach, we can speed up the process and lower the cost of DNA synthesis while taking advantage of the fact that we can manipulate how data is mapped to DNA sequences. We experiment with different methods to fit as much data as possible into a DNA sequence and work on error correction methods to enable the retrieval of data with high fidelity. The team at Imperial College also explores how to store the data in the long term. More specifically, we study how to keep DNA stable, even at room temperature and with minimal demands regarding humidity, light and air quality. Doing so is crucial as it means little or no errors are introduced due to decay. We also improve the accuracy of the reading process through sequencing. For example, we improve machine learning methods to enhance their accuracy. Our work at Imperial College will make storing culturally important data in DNA a reality much sooner. Our work is not only scientifically important but also has immediate practical impact. Specifically, we work with multiple organisations and institutions which need to preserve culturally important data in the long term. For example, we work with the Danish National Archives to encode and store culturally important paintings by King Christian IV dating from between 1583-1591, to preserve them for posterity. We also work with the Tamil Nadu Archives in India to preserve some of the most important scriptures for generations to come. This is of particular relevance as preserving the cultural heritage of the developing world is a major challenge. Creating a storage media that is cheaper and more stable in adverse conditions will enable us to successfully write and preserve our history, ensuring that we are all winners. ■ Thomas Heinis is a Senior Lecturer in Computing at Imperial College, where he leads the SCALE Lab.


Articles inside

Science behind the bars

8min
pages 38-39

Salam and his legacy for the Rest of the World

5min
pages 36-37

The Chinese scientists that cracked the universe’s mirror

7min
pages 34-35

How Hong Kong’s cultural identity shaped its pandemic response

5min
pages 32-33

Ayurveda: A brief introduction to the science of life

5min
pages 30-31

Preserving the world’s heritage with DNA

6min
pages 28-29

Art of the Issue

2min
pages 26-27

Wait, Mother Nature Has a Few Words for You Too

1min
page 25

Indigenous knowledge in the Amazon

5min
page 24

A recipe for disaster: tackling toxic e-waste

6min
pages 22-23

Apoqnmatulti’k (We help each other)

5min
pages 20-21

The Other Side of the Manifold

5min
pages 18-19

Strange bedfellows: Why art and science go hand in hand

4min
pages 16-17

The hidden history of Islamic science

5min
pages 14-15

Graphic Science: How comic books created real heroes

6min
pages 12-13

The Ig Nobel Prizes

5min
pages 10-11

Here and There

4min
pages 8-9

News

8min
pages 4-5

Lost in translation

7min
pages 6-7
Issuu converts static files into: digital portfolios, online yearbooks, online catalogs, digital photo albums and more. Sign up and create your flipbook.