
Researchers from Microsoft and the University of Washington have demonstrated the first fully automated data storage system in artificially created DNA with reading capabilities. This is a key step towards transitioning new technology from research laboratories to commercial data centers.
The developers confirmed the concept with a simple test: they successfully encoded the word «hello» into fragments of synthetic DNA molecules and converted it back into digital data using a fully automated end-to-end system described in , published on March 21 in Nature Scientific Reports.
Digital information can be stored in DNA molecules at very high density, meaning in physical space that is orders of magnitude smaller than that occupied by modern data centers. This is one of the promising solutions for storing the vast amounts of data generated by the world every day, from business records and cute animal videos to medical images and pictures from space.
Microsoft is exploring ways to overcome the potential gap between and wish to store, and our ability to hold it. Some of these methods include the development of algorithms and molecular computing technologies for . This could allow all the information currently stored in a large modern data center to fit into a space roughly the size of a few dice.
«Our main goal is to launch a system that will look almost like any other cloud storage system for the end user: data is sent to the data center and stored there, then simply appears when needed by the client,» says Microsoft senior researcher Karin Straus. «For this, we needed to prove that it makes practical sense in terms of automation».
Information is stored in synthetic DNA molecules created in the laboratory, rather than in the DNA of humans or other living organisms, and can be encrypted before being sent to the system. Although complex machines, such as synthesizers and sequencers, already perform key parts of the process, many of the intermediate steps still required manual labor in the research laboratory. "This is not suitable for commercial use," said Chris Takahashi, a senior researcher at the Paul G. Allen School of Computer Science and Engineering at the University of Washington.).
"People can't run around a data center with pipettes; this approach has too high a risk of human error, is too costly, and takes up too much space," explained Takahashi.

For this method of data storage to make sense commercially, costs must be reduced both for DNA synthesis—creating the fundamental building blocks with meaningful sequences—and for the sequencing process necessary to read the stored information. Researchers say that significant advancements are being made in this direction. .
According to researchers from Microsoft, automation is another key part of this puzzle, enabling data storage to be organized at a commercial scale and making it more accessible.
Under certain conditions, DNA can last much longer than modern archival storage media, which degrade within decades. Some DNA has been found to remain preserved in less than ideal conditions for tens of thousands of years—in mammoth tusks and the bones of early humans. This means data can be stored in this way as long as humanity exists.
The automated DNA storage system utilizes software developed by specialists from Microsoft and the University of Washington (UW). It converts digital data units of ones and zeros into sequences of nucleotides (A, T, C, and G), which represent the 'building blocks' of DNA. The system then employs affordable, primarily standard laboratory equipment to deliver the necessary liquids and reagents to a synthesizer that assembles the produced DNA fragments and places them in storage containers.
When the system needs to retrieve information, it adds other chemicals for proper DNA preparation and uses microfluidic pumps to push liquids into the sections of the system that read DNA molecule sequences and convert them back into information understandable by a computer. Researchers state that the objective of the project was not to prove that the system can operate quickly or cheaply, but rather to simply demonstrate that automation is possible.
One of the most obvious advantages of the automated DNA storage system is that it frees scientists to tackle complex problems, allowing them not to waste time searching for reagent bottles or monotonously adding drops of liquid to test tubes.
"Having an automated system to perform repetitive tasks allows laboratory staff to focus on research, developing new strategies to implement innovations faster," said Microsoft researcher Bihlin Nguyen.
The team from the Molecular Information Systems Lab (MISL) has already demonstrated that it can store cat pictures, remarkable literary works, and archival records in DNA and retrieve these files without errors. To date, they have managed to store 1 gigabyte of data in DNA, breaking .
Researchers have also developed methods for , such as searching and retrieving only images that contain an apple or a green bicycle, using the molecules themselves without converting the files back to digital format.
"We can confidently say that we are witnessing the birth of a new type of computer system, one that uses molecules for data storage and electronics for control and processing. This combination opens up very interesting possibilities for the future," said a professor from the Allen School at the University of Washington. .
Unlike computational systems based on silicon components, DNA-based storage and computing systems must use fluids to move molecules. However, fluids differ from electrons by their nature and require entirely new technical solutions.
The University of Washington team, in collaboration with Microsoft, is also developing a programmable system that automates laboratory experiments by using the properties of electricity and water to move droplets on an electrode grid. The complete software and hardware suite, named , can mix, separate, heat, or cool various liquids and execute laboratory protocols.
The goal is to automate laboratory experiments that are currently performed manually or by expensive fluid-handling robots, as well as to reduce costs.
The next steps for the MISL team include integrating a simple end-to-end automated system with technologies like Purple Drop, as well as with other technologies that allow for searching DNA molecules. Researchers specifically designed their automated system to be modular so it can evolve as new technologies for DNA synthesis, sequencing, and processing emerge.
"One of the advantages of this system is that if we want to replace one of the components with something newer, better, or faster, we can simply plug in the new part," said Nguyen. "This gives us great flexibility for the future."
Top image: Researchers from Microsoft and the University of Washington recorded and read the word "hello", using the first fully automated DNA data storage system. This is a key step in transitioning the new technology from laboratories to commercial data centers.
Source: habr.com
