We talk about the people of the future who decode organic big data. Over the past two decades, the amount of biological data available for analysis has increased dramatically due to the decoding of the human genome. Before this, we could hardly imagine that information literally stored in our blood could determine our origins, check how our body would react to certain medications, and even alter our biological inheritance.
This and other articles are published first on our website. Enjoy your reading.
The attributes of a typical bioinformatician are similar to those of a programmer — red eyes, a hunched posture, and coffee cup rings on the desk. However, the work at this desk is not about abstract algorithms and commands but about the very code of nature, which can tell us a lot about ourselves and the world around us.
Specialists in this field deal with immense volumes of data (for example, the sequencing results of one person's genome take about 100 gigabytes). Therefore, processing such an amount of information requires Data Science approaches and tools. It is logical that a successful bioinformatician must understand not only biology and chemistry but also data analysis methods, statistics, and mathematics — this makes their profession quite rare and in demand. Such specialists are particularly needed in areas of innovative medicine and drug development. Tech giants like IBM and Intel , dedicated to the study of bioinformatics.
What do you need to become a bioinformatician?
- Biology and chemistry (university level);
- Mathematical statistics, linear algebra, probability theory;
- Programming languages (Python and R, C++ is also frequently used);
- For structural bioinformatics: understanding mathematical analysis and differential equations theory.
You can enter the field of bioinformatics either with a biological background or with knowledge of programming and mathematics. Those with a biological background can work with existing bioinformatics programs, while those with programming expertise are suited for a more algorithmic profile of the specialty.

What do bioinformaticians do?
Modern bioinformatics is divided into two main branches — structural bioinformatics and sequence bioinformatics. In the former case, we see a person sitting in front of a computer running programs that help study biological objects (such as DNA or proteins) in 3D visualizations. They build computer models that predict how a drug molecule will interact with a protein, what the spatial structure of a protein looks like in the cell, what properties of molecules explain their interactions with cellular structures, and so on.
Methods of structural bioinformatics are actively used both in academic science and in the industry: it is hard to imagine a pharmaceutical company that operates without such specialists. In recent years, computational methods have significantly simplified the process of searching for potential drugs, making pharmaceutical development a much faster and cheaper process.

RNA-dependent RNA polymerase of SARS-CoV-2 (on the left), as well as its connection with the RNA duplex.
What is a genome?
A genome is all the information about the structure of an organism's heredity. Nearly all living beings have DNA as the carrier of their genome, but some organisms convey their hereditary information in the form of RNA. The genome is passed from parents to children, and during this transfer process, errors can occur — mutations.

The interaction of the drug remdesivir with the RNA-dependent RNA polymerase of the SARS-CoV-2 virus.
Sequence bioinformatics operates at a higher level of organization of living matter — starting from individual nucleotides, DNA, and genes, and ending with whole genomes and their comparisons with one another.
Imagine a person who sees a set of letters from the alphabet (not just any letters, but those of genetic or amino acid nature) and searches for patterns within them, explaining and confirming their findings statistically using computational methods. Sequence bioinformatics clarifies which mutation is associated with a particular disease or why harmful substances accumulate in a patient's blood. In addition to medical data, sequence bioinformatics studies the patterns of organisms' distribution on Earth, population differences among animal groups, and the roles and functions of specific genes. This field allows for testing the efficacy of drugs and exploring the biological mechanisms that explain their action.
For instance, due to bioinformatics analysis, mutations leading to the development of cystic fibrosis—a monogenic disease caused by a defect in one of the chloride channel genes—have been identified and described. Moreover, we now have a much clearer understanding of who our closest biological relatives are and how our ancestors spread across the planet. Furthermore, by reading their genome, individuals can discover their ancestral origins and which ethnic group they belong to. Numerous foreign (, ) and Russian (, ) services offer this service for a relatively low price (around 20,000 rubles).

Results of the DNA test on ancestry and population affiliation from MyHeritage.

Results of the DNA test on population affiliation from 23andMe.
How is the genome read?
Today, genome sequencing is a routine procedure that will cost anyone interested approximately thousand rubles (including in Russia). To read your genome, simply provide a blood sample from your vein in a specialized laboratory: in about two weeks, you will receive your results with a detailed description of your genetic traits. In addition to your genome, you can also analyze the genomes of your gut microbiota: you will learn about the bacteria inhabiting your digestive system and receive a consultation from a professional dietitian.
The genome can be read using various methods, with one of the primary ones being the so-called "next-generation sequencing". To carry out this procedure, biological samples must first be obtained. In every cell of the organism, the genome is the same, so blood is often taken for genome reading (as it is the simplest). Afterward, the cells are broken down, and DNA is separated from everything else. Then, the extracted DNA is chopped into many small pieces, and special adapters—artificially synthesized known sequences of nucleotides—are "sewn" onto each of them. Following this, the DNA strands are separated, and single-stranded sequences are attached to a special plate using adapters, on which sequencing takes place. During sequencing, complementary fluorescently labeled nucleotides are added to the DNA sequence. Each labeled nucleotide emits a beam of light of a specific wavelength upon incorporation, which is recorded by the computer. This allows the computer to read short sequences of the original DNA, which are then assembled into the original genome using special algorithms.

An example of data that bioinformaticians work with is the alignment of amino acid sequences.
Where do bioinformaticians work and how much do they earn?
The path of a bioinformatician is traditionally divided into two main areas: industry and science. A scientist-bioinformatician's career typically begins with a graduate position at one of the major institutions. Initially, bioinformaticians receive a base salary that depends on their institution, the number of grants they participate in, and the number of affiliations—places where they are officially employed. Over time, the number of grants and affiliations increases, and after a couple of years in the academic environment, a bioinformatician can easily earn an average salary (70-80 thousand rubles); however, much depends on diligence and hard work. The most experienced bioinformaticians eventually establish their own laboratories in their areas of expertise.

Where can one study to become a bioinformatician?
- Moscow State University — Faculty of Bioengineering and Bioinformatics
- Higher School of Economics — Data Analysis in Biology and Medicine (Master's Program)
- MIPT — Department of Bioinformatics
- Institute of Bioinformatics (NCO)
Unlike in academia, in the industry, no one will spend their time training employees in the necessary skills, making it generally harder to get in. The career path of a bioinformatician in the industry varies greatly depending on specialization and workplace. On average, salaries in this field range thousand rubles, depending on experience and specialization.
Famous Bioinformaticians
The history of bioinformatics can be traced back to Frederick Sanger, an English scientist who received the Nobel Prize in Chemistry in 1980 for discovering a method to read DNA sequences. Since then, each year, methods for reading sequences have improved, but the 'Sanger sequencing' method laid the foundation for all subsequent research in this area.

Interestingly, many programs developed by Russian scientists are now widely used around the world — for example, the genome assembler , — St. Petersburg genome assembler, created at the Saint Petersburg Institute, helps scientists around the globe assemble short DNA sequences into larger sequences to reconstruct the original genomes of organisms.
Discoveries and Achievements in Bioinformatics
Today, bioinformaticians make numerous useful discoveries. The development of drugs for coronavirus would be unimaginable without decoding its genome and the complex bioinformatics analysis of processes occurring during the disease. An international group of scientists, using comparative genomics and machine learning methods, has been able to understand what coronaviruses have in common with other pathogens.
It turned out that one such feature is the strengthening of nuclear localization signals (NLS) in pathogenic viruses during evolution. This research could assist in studying virus strains that may pose potential risks to humans in the future and possibly initiate preventive drug development.
Moreover, bioinformaticians have played a key role in developing new genome editing methods, particularly the CRISPR/Cas9 system (a technology based on the immune system ). Through bioinformatics analysis of protein data structures and their evolutionary development, the accuracy and efficiency of this system have dramatically increased in recent years, allowing for targeted editing of the genomes of many organisms (including humans).
You can acquire a sought-after profession from scratch or level up your skills and salary by taking online courses at SkillFactory:
More Courses
Source: habr.com
