Improvising cells and the new biology
The living cell, once thought to be a precise molecular factory, is turning out to be more like an improvising jazz ensemble. The old dogma—one gene, one protein, one function—has collapsed. Molecular biologist Ewa Grzybowska argues that recent discoveries show that proteins can switch folds, shift shapes, or even remain gloriously unstructured, improvising their roles as they go. Genes are not blueprints but texts, open to continuous interpretation by cells. Life, it turns out, is not built like a machine but is instead fluid, improvisational, and brimming with creative possibility.
The old paradigms in biology
When Watson and Crick decoded DNA in 1953 and the mechanism of protein-making was discovered, we obtained an extraordinary tool to explain the inner workings of life. The basic principle of making one protein from one DNA template (gene) with the assistance of one messenger RNA (mRNA) and several well-defined amino-acid-transporting RNAs (tRNAs) was so successful that it has been enshrined in millions of textbooks and not questioned for a long time.
Ripples on this smooth surface appeared with the discovery of introns (non-coding parts, or sequences within genes that are not expressed in RNA, and so apparently don’t contribute to determining the protein) and splicing (basically, cutting-and-pasting the relevant parts of the genetic code to construct working mRNA and leaving behind introns). Later, alternative splicing was discovered, which directly contradicts the basic principle of one protein per gene, because it involves making several different working mRNAs from one gene. Several good explanations have been proposed for the role of splicing in general, and alternative splicing in particular, mainly focusing on regulatory functions, but a big question remained: why do all organisms, and especially more complex organisms, have this vast amount of non-coding DNA polluting their genomes?
We disregard what we do not understand, so people called this mysterious genetic material “junk DNA.” However, it should probably have been clear from the beginning that natural selection works at the level of the living cell and is both frighteningly effective and brutal. There is no place in the cellular economy for tolerating molecules that take a lot of energy to make and are of no use whatsoever. The functionality of this genetic material should have been considered axiomatic; we just had no idea what this elusive function might be.
Proteins were considered the main output—the most important product (if not the only one)—of the genetic code. It was assumed that proteins are well ordered, have domains composed of defined elements of secondary structure (alpha-helices and beta-sheets), and that only these well-organized parts have important cellular functions. The poster example was always an enzyme: well structured, fitting to the substrate like a lock fits a key (the “Lock and Key Model” was proposed by Emil Fischer in 1894). It was also assumed that proteins have only one conformation (that is, one shape or structure), associated with one function. This is the one gene → one structure → one function paradigm.
We disregard what we do not understand, so people called this mysterious genetic material “junk DNA”.
Even more important—indeed, revolutionary—changes in the overall perception of how proteins work came with the discovery of intrinsically disordered proteins: proteins that lack a stable three-dimensional structure and yet are fully functional.
This massive paradigm shift in RNA functionality was accompanied by new realizations about the dependence of protein structure and function. The overall three-dimensional shape of proteins (their “tertiary structure”) is thought to be determined by their amino acid sequence—this is a rule called Anfinsen’s Dogma. This view is based on the observation that some proteins can regain their structure after it has been chemically destroyed, even when other cellular components are absent. The idea is that, in the absence of other cellular components, only the amino acid sequence could be directing the process. This is of course true in principle, but as with, for example, Mendelian inheritance, it can be observed only in some instances, and many other times the picture is more complicated. Indeed, folding does not take place in a vacuum, and is assisted by chaperones. However, chaperones mostly ensure proper folding and prevent aggregation; hence, they do not enforce some alternative structure or shape, and therefore do not contradict Anfinsen’s Dogma. More challenging for the principle is the existence of other classes of proteins, including moonlighting and fold-switching proteins.
More studies have led to the puzzling realization that, in fact, most of the human genome (around 85%) is transcribed (converted to RNA), despite the fact that only 1–2% of the genome encodes proteins. This means that cells produce a lot of non-coding RNA. At the same time, different classes of non-coding RNA were attracting a lot of attention. Besides mRNA and known housekeeping RNAs (for example, tRNA and ribosomal RNA [rRNA]), there is a vast pool of functional but non-coding RNAs. These new classes of RNA (miRNA, snRNA, lnRNA, piRNA, cicrRNA) were found to be involved in the regulation of gene expression at both the transcriptional and post-transcriptional levels, providing completely new and vast planes of regulatory activity.
The Human Genome Project was one of the most extraordinary feats in the history of science. Launched in 1990 and completed around 2003, the organized effort of many scientific teams provided fundamental information about our complete genetic code, the so-called “blueprint” for our bodies. Great hopes were placed in its completion, based on a misguided deterministic view that genes (defined as protein-coding units) and their specific configuration determine every aspect of life, and all other levels of protein-making and functioning represent only a natural consequence of this genetic input. The big surprise that came out of the Human Genome Project was that we have a relatively small number of genes: around 20,000, much less than in many other species. But surely we are highly evolved? What is the key to our complexity, then?
For some time, it was known that some proteins with well-defined functions can occasionally moonlight, performing different tasks, which is often a consequence of a change in the protein’s tertiary structure, imposed by some change in local conditions, for example, iron deficiency. Several hundred moonlighting proteins have been characterized so far, including metabolic enzymes, chaperones, secreted cytokines, transcription factors, DNA stabilizers, and components of the cytoskeleton or proteasome subunits. Moonlighting already defies the paradigm of one gene → one structure → one function, but fold-switching proteins diverge even more radically from it. These proteins (also called metamorphic proteins) can acquire an alternative not only tertiary, but also secondary structure, which means that they can fold differently using the same amino acid sequence. Moreover, they may switch reversibly between two differently folded states depending on changing environmental conditions.
Even more important—indeed, revolutionary—changes in the overall perception of how proteins work came with the discovery of intrinsically disordered proteins: proteins that lack a stable three-dimensional structure and yet are fully functional. Instead of one structure, they display a whole range of dynamically changing conformations, which is crucial to their functionality.
Intrinsically disordered proteins are characterized by weak, multivalent interactions that allow them to quickly change binding partners and thereby adapt to changes in the local microenvironment. This flexibility and promiscuity allow them to function as regulatory hubs, integrating various signals and coordinating signaling pathways, gene expression, and cell cycle. Intrinsically disordered proteins represent the entire spectrum, from fully disordered to partially structured. In fact, partial disorder is present in most proteins, often in the form of disordered, flexible linkers between protein domains (which are distinct units within a protein that can fold and function independently of each other) or disordered tails (protein termini). Furthermore, it has been demonstrated that even well-folded proteins require temporary local disorder and conformational switching to fine-tune binding affinities, suggesting that disorder is much more widespread and functionally important than we think.
Moreover, the disorder is evolutionarily conserved, which additionally underlines its importance. As with well-ordered structures, disorder is encoded in the amino acid sequence and can be predicted from it, using specific algorithms.
Another interesting feature of intrinsically disordered proteins is that they are able to acquire an ordered structure after binding to some other macromolecules. This disorder-to-order transition shows that folding can be induced, depending on the conditions. This realization raises further questions concerning the inner workings of the protein-making machinery. It seems that the amino acid sequence, while still the most important factor, only partially determines folding, and the other contributing factors include local physicochemical conditions and time. Time can be an important factor, because a part of a polypeptide chain may have different energy minima and acquire a different structure by itself, compared to the structure it would have acquired as a part of a whole polypeptide.
Thus, if some factor causes a pause in the translation after a substantial part of the protein has already been produced, this may result in a different structure compared to the case in which the entire polypeptide chain is produced all at once. In living cells, folding occurs co-translationally, with the N-terminus starting to fold upon exiting the ribosome, while the C-terminus is still synthesized on the ribosome. This allows for the regulation by a proper timing of the translation process. Recent high-throughput analyses of protein-binding regions on the entire transcriptome revealed that a large portion of this RNA-binding occurs in the coding regions of the transcripts. This was surprising because protein binding to the coding region would interfere with translation. But maybe this is the point? This binding may work as a timer, allowing a part of the protein to fold but withholding the folding of the other part, which ultimately impacts the final structure.
These new discoveries and interpretations produce an emerging big picture that is drastically different from the situation in which DNA-encoded information determines everything, and represents the only source of variability. Instead, we are confronted with layers upon layers of regulation at almost every possible step of the functioning of the cellular machinery.
Conclusions and the future
In the last two decades, the textbook rules and explanations of molecular biology have crumbled and been transformed. Genetic determinism was replaced by a much more elastic view, according to which the genetic code can be interpreted in different ways, producing different outcomes, depending on local conditions, availability of various factors, and time. There are still many unanswered questions concerning the true functionality of non-coding regions, both in DNA and RNA molecules; so far, we have only had a glimpse of the possibilities. Similarly, the role of structural change and disorder in protein functioning still needs clarification. The picture emerging from these new developments is much more dynamic and complicated, but also much closer to the truth. This is a normal scientific process, in which we have gained a better understanding of the complexity of the living cell. What is new is the increasing speed with which science progresses. The exceptional effectiveness of AlphaFold (an AI system that predicts a protein’s 3D structure from its amino acid sequence) and other AI systems suggests that this process will further accelerate and expand, with more and more AI-generated input.
Originally published iai Presents