
During the ongoing COVID-19 pandemic, many issues arose that hackers eagerly capitalized on. From face shields printed on 3D printers and homemade medical masks to substituting full-fledged mechanical ventilators – this flow of ideas inspired and delighted many. At the same time, there were attempts to advance in another area: research aimed at combating the virus itself.
It seems that the greatest potential for stopping the current pandemic and preventing all future ones lies in an approach that tries to get to the very root of the problem. This 'know your enemy' approach is embraced by the computing project Folding@Home. Millions of people have registered for the project and donate part of their CPU and GPU computing power, thereby creating the largest distributed supercomputer in history.
But what exactly are all these exaflops used for? Why allocate such computing power to ? Какая тут работает биохимия, зачем вообще белкам нужно укладываться? Вот краткий обзор фолдинга белков: что это, как он происходит и в чём его важность.
First and foremost: why are proteins important?
Proteins are essential structures. They not only provide the building blocks for cells but also act as catalysts for almost all biochemical reactions. Proteins, whether they are or , consist of long chains of , arranged in a specific sequence. The functions of proteins are determined by which amino acids are located in specific places on the protein. For example, if a protein needs to bind to a positively charged molecule, the binding site must be filled with negatively charged amino acids.
To understand how proteins achieve the structure that defines their function, one must run through the basics of molecular biology and the information flow within a cell.
The production, or of proteins begins with the process of During transcription, the double helix of DNA, containing the genetic information of the cell, partially unravels, allowing access to the nitrogenous bases of DNA for an enzyme known as The task of RNA polymerase is to make an RNA copy, or transcription, of the gene. This copy of the gene is called (mRNA) is a single-stranded molecule perfectly suited for controlling intracellular protein factories, , which are responsible for the production or of proteins.
Ribosomes behave like assembly devices – they capture the mRNA template and match it with other small pieces of RNA, (tRNA). Each tRNA has two active regions – a section of three bases called an , which must match the corresponding codons of the mRNA, and a site for binding the amino acid specific to that During translation, tRNA molecules in the ribosome attempt to randomly bind to the mRNA using anticodons. Upon success, the tRNA molecule attaches its amino acid to the previous one, forming the next link in the amino acid chain coded by the mRNA.
This sequence of amino acids is the first level of the protein's structural hierarchy, and is therefore called its The entire three-dimensional structure of a protein and its functions directly arise from the primary structure and depend on the various properties of each amino acid and their interactions with one another. Without these chemical properties and interactions of amino acids, would remain linear sequences without a three-dimensional structure. This can be observed every time food is cooked – in this process, thermal of the three-dimensional structures of proteins occurs.
Long-range interactions between parts of proteins
The next level of three-dimensional structure, extending beyond the primary, has been cleverly named It includes hydrogen bonds between amino acids of relatively close action. The essence of these stabilizing interactions boils down to two things: and An alpha-helix forms a tightly wound section of the polypeptide, while a beta-sheet creates a smooth and broad area. Both formations have structural as well as functional properties that depend on the characteristics of the amino acids that compose them. For example, if an alpha-helix predominantly consists of hydrophilic amino acids such as or it is likely to participate in aqueous reactions.

Alpha helices and beta sheets in proteins. Hydrogen bonds form during protein expression.
These two structures and their combinations form the next level of protein structure — . Unlike simple fragments of secondary structure, tertiary structure is primarily influenced by hydrophobicity. At the centers of most proteins, amino acids with high hydrophobicity, such as or , are present, and water is excluded due to the "fatty" nature of the radicals. These structures often appear in transmembrane proteins embedded in the lipid bilayer surrounding cells. The hydrophobic parts of proteins remain thermodynamically stable within the lipid core of the membrane, while the hydrophilic parts of the protein are exposed to the aqueous environment on both sides.
The stability of tertiary structures is also ensured by long-range interactions between amino acids. A classical example of such interactions is the , often occurring between two cysteine radicals. If you noticed a smell resembling rotten eggs while getting a perm at a hair salon, it was due to partial denaturation of the tertiary structure of keratin in the hair, caused by the reduction of disulfide bonds by sulfur-containing mixtures.

The tertiary structure is stabilized by long-range interactions, such as hydrophobicity or disulfide bonds.
Disulfide bonds can occur between radicals in one polypeptide chain, or between cysteines from different complete chains. Interactions between different chains form the level of protein structure. A prime example of quaternary structure is in your blood. Each hemoglobin molecule consists of four identical globins, parts of a protein, each maintained in a specific position within the polypeptide by disulfide bonds, and also associated with a heme molecule containing iron. All four globins are linked by intermolecular disulfide bridges, and the entire molecule can bind with several air molecules at once, up to four, and release them as needed.
Modeling Structures in the Search for Treatment of Disease
Polypeptide chains begin to fold into their final shape during translation, when the growing chain exits the ribosome – somewhat like a segment of shape memory alloy can take complex forms when heated. However, as is often the case in biology, it’s not that straightforward.
In many cells, the transcribed genes undergo significant editing before translation, greatly altering the basic protein structure compared to the pure gene base sequence. In this process, translational mechanisms often enlist the help of molecular chaperones, proteins that temporarily associate with the emerging polypeptide chain and prevent it from adopting any intermediate shape from which it cannot transition to the final form.
This all leads to the point that predicting a protein's final form is not a trivial task. For decades, the only way to study protein structures was through physical methods such as X-ray crystallography. Only in the late 1960s did biophysical chemists begin to build computational models of protein folding, focusing mainly on modeling secondary structure. These methods and their descendants require vast amounts of input data in addition to the primary structure – for instance, tables of amino acid bonding angles, lists of hydrophobicity, charged states, and even the preservation of structure and function across evolutionary time scales – all in order to guess how the final protein will look.
Today's computational methods for predicting secondary structure, particularly those in the Folding@Home network, operate with approximately 80% accuracy — quite impressive given the complexity of the task. The data obtained from predictive models for proteins such as the spike protein of SARS-CoV-2 will be compared with data from physical studies of the virus. Ultimately, this will allow us to achieve an accurate protein structure and possibly understand how the virus attaches to receptors. in humans, located in the airways leading into the body. If we can understand this structure, we will likely be able to find drugs that block binding and prevent infection.
Protein folding studies are at the heart of our understanding of a multitude of diseases and infections. Even when we, through the Folding@Home network, figure out how to defeat COVID-19, which we are currently witnessing a surge of, this network will not remain idle for long. It is a research tool excellently suited for studying protein models underlying dozens of diseases associated with improper protein folding — for example, Alzheimer's disease or a type of Creutzfeldt–Jakob disease, often incorrectly referred to as mad cow disease. And when another virus inevitably appears, we will already be ready to start fighting it again.
Source: habr.com
