The Cosmos is a 'mind-bogglingly' big space and it's unlikely we shall ever explore it. 'Polypeptide Space' is vast too, but there's a better chance we might get to explore those parts of it that are important to us - Michael Geisow
Polypeptides: molecules with a 3.5 billion year pedigree
Note: in my article I use the term 'polypeptide' where I'm talking about a single molecular entity and the emphasis is on chemistry. I use 'protein' where function is the focus and it may be composed of multiple polypeptide chains. There's no real difference in chemistry here. No-one knows for sure when polypeptides appeared on earth or how they arose, but there's broad agreement that nucleic acids (RNAs) were their precursors. It's an area ripe with conjecture and thought experiment! It's notable that the key molecules of life today are all polymers (polypeptides, polynucleotides and polysaccharides). Small molecules just don't cut it when it comes to information storage or functions: properties needed to 'be alive'. All life today is underpinned by two of these biopolymers: DNA which carries the plans of life and proteins which carry them out. RNA acts as a kind of go-between.
Nature and evolutionary selection has produced polypeptides which enable life and from where most of our present understanding of their structure and function come. Man-made 'designer' polypeptide science is an emergent field which has implications both for medicine, new materials and applications about which we can only speculate at present.
Polypeptides make life, but what exactly is life?
Taking the cell as the basic unit of life, let's see where proteins play their part. It's worth recalling the 8 properties currently used to define the state of 'being alive'. After all, with our push toward machine 'intelligence' and our exploration of the Cosmos, one day we may need to amend our present definition.
The accepted definition of life (and the natural roles of proteins)
Order Proteins ensure the temporal integrity of the cell and maintain its structure, shape and motility.
Response to stimulus All life is 'aware of' and dependent upon its environment. Proteins provide both internal and external sensors.
Reproduction Protein machinery operates the mechanics of DNA replication and cell division.
Growth and development No life forms are static. Their dynamic nature is operated by proteins.
Management Proteins provide oversight of all cell processes. There's no overall boss: successful businesses that they are, cells run on 'matrix management'!
Energy Energy is essential to life. The key energy generators of all life forms are protein complexes.
Homeostasis 'Keeping on keeping on' through defence and repair relies absolutely upon proteins.
Evolution Providing cells with the capacity to adapt to environmental change (and passing this capability on) is intrinsic to the way proteins are structured. Built up with domains (functional subunits) these domains can be duplicated. mutated and passed between organisms to endow them with new functions
Which proteins are responsible for the key properties of life?
Enzymes are the most diverse and important type of proteins. They provide our energy. They make and break chemical bonds and rapidly catalyse chemical reactions at ambient temperature to produce molecules for cells, some of which still elude synthesis by organic chemists. We also have:
Hormones and receptor proteins regulate our internal ecology
Motor proteins give us movement.
Architectural proteins create the structures of our cells and bodies.
Transport proteins facilitate both extracellular and intracellular transport of small molecules
Regulatory proteins control gene expression
Receptor proteins communicate with the outside world.
Sensor proteins underpin our perception of that world.
Neuronal proteins help us make sense of it.
The proteome and the interactome
The complement of proteins within cells or tissues (we've around 20,000 which could potentially be expressed) is referred to as the proteome. Protein interactions between both large and small molecules are what fundamentally creates and supports life. So important is this concept, molecular biologists refer to the totality of protein structural and functional contacts as the interactome.
Where did proteins come from & what kicked it all off?
It was once an RNA World We are the inheritors of biopolymers forged - perhaps accidentally - in deep time. Nucleic acids started it and (we believe) proteins were relative latecomers to the game of life. Where these nucleic acids came from (on planet earth or as those voting for an extra-terrestrial origin panspermia believe) no-one really knows. Against entropy, nucleic acid molecules grew in complexity, increasing their polymer lengths, acquiring a range of catalytic and structural functions and the capability to reproduce themselves. Eventually the information about these properties became more stably stored in DNA and the capacity 'do more interesting things' led to the invention of proteins. But, the founding member of corporate life - we assume - was RNA, which in those distant days, in the immortal words of Canadian philosopher Marshall McLuhan: 'the medium was the message'. In truth, this is a classic 'which came first, the chicken or the egg' question. The answer is buried in deep time.
Even today, RNA has retained the critical controls RNA collaborates with protein 'acolytes' in the manufacture of proteins in the molecular machines called ribosomes. RNA has retained the critical step of forming peptide bonds. And like card-driven Jacquard looms, ribosomes turn out proteins to the specifications written in messenger RNA. I like to think that early RNA 'delegated' these specifications (genes) to be stored and curated within DNA!
Genome manuals specify many more instructions than products To fulfil the 8 characteristics of life listed above, a lot of proteins are required. But (in us at least) only 1-2% of our DNA stores the information which specifies them. A much larger percentage of our DNA holds the information about which proteins will be produced (expressed) and how, when and where this will happen. The management and mechanics of this information transfer (transcription) once again is a critical collaboration between RNA and proteins themselves.
Editing and Revision Control of protein manufacture occurs at many stages. Genes (DNA) specifying proteins are subject to multiple controls, many still being uncovered. Gene transcripts (RNAs) are tightly regulated both in nature and quantity. Like sentences in the writing of a novel, RNA transcripts are subject to editing (and even redaction!) in producing the messengers (mRNAs) which will be translated into proteins. Following translation into proteins in the ribosomes, more editing occurs. Newly synthesised proteins can be subject to post-translational modification and even proteolytic processing in which their polypeptide chains are cut and peptides removed. Certain of this processing relates to whereabouts proteins are to be sent. Proteins come with peptide 'addresses' which direct them to cell compartments or to the environment external to their host cells.
Compartments for Protein Chemistry
Throughout our cells, proteins self-associate to form the critical machinery which operates most of the characteristics of life listed above. Machinery even exists to repair, recycle or remove proteins themselves. Waste recycling is as critical to cells as it is to the health of our own societies. This machinery itself is of course composed of proteins.
This constant molecular processing requires specialised zones (as in human factories) where conditions can be controlled. Research in molecular cell biology has made us very aware of the critical importance of compartments in our cells. We've long recognised the organelles (nucleus, mitochondrion etc.) but it has been more recently we have realised the proteins themselves form compartments where different types of processing take place.
I will finish my philosophising in this section with an intriguing example (although there are many others I could have selected which will be referred to elsewhere in this text). Protein vaults first noticed in electron microscopy of cells are illustrated on the right as the closed form and one half of the vault assembly is illustrated below at higher resolution.
We've moved on from the early days of protein science to consider the structures, function, regulation and most recently the design and synthesis of macromolecular protein assemblies
THE ENIGMATIC PROTEIN VAULTS
A vault: a poorly understood component of human cells. Discovered by Nancy Kedersha and Leonard Rome in 1986, these structures resemble the arches of a cathedrals vaulted ceiling and have a 39-fold symmetry. These vault structures are huge: composed of about 33,000 amino acids. Present in nearly all eukaryotic cells, vaults have an outer shell made of 78 copies of the Major Vault Protein (MVP). The image shows one half of a vault assembly. Two copies of the illustrated structure connect at the wider end to form the complete vaults assembly which was visualised in the electron microscope. A single MVP (blue) is shown on the right and its location in the complete vault assembly (left). The MVP is made up of a long helical region and around 9 repeating subdomains. Vaults are colossal, about three times the size of a ribosome, and weigh a mammoth 13 MDa with around 33,000 amino acids! They are among the largest naturally occurring particles in human cells, with each cell housing around 10,000 of these structures. In some immune cells, this number can soar to 100,000! Despite their abundance, their exact role remains elusive. However, they are believed to be involved in numerous cellular functions, including nuclear-cytoplasmic transport, mRNA localization, drug resistance, cell signalling, nuclear pore assembly, and innate immunity. Credit: My thanks to Soutick Saha, Wolfram, Purdue University, Canada for this image.
Timeline of protein science: from Berzelius to BoltzGEN
1948 when I was born, there was only a relative handful of proteins which had been isolated in pure form - mainly blood plasma proteins and enzymes - and little detail was known about their structure and function. By the I started a degree in chemistry in 1967, biochemistry had become an established field and the first protein structures including the globins and egg white lysozyme had been determined. We had final Uni. year lectures from Prof Tony North who had been one of the post-docs with Max Perutz working on haemoglobin 3D structure, later with David Phillips at the Royal Institution in London solving the structure and action of hen egg white lysozyme. I had been doing a wet lab project at Warwick and not really getting anywhere with understanding enzyme mechanism. A lot of protein chemistry then involved working in the dark! X-ray crystallography seemed to hold out the real possibility of making significant scientific discoveries in protein science. It was this that drew me away from organic chemistry to study for a D.Phil at Oxford in the Laboratory of Molecular Biophysics. A change of research field made by many pure chemists and even physicists who believed life science held out more opportunities than their more 'mature' fields. We didn't regret that decision. Years later it has inspired me to look back at the roots of protein science and some of the key developments and discoveries that were behind protein science today.
1925 It's important to note here that until the opening up of the fields of genomics, epigenetics and high resolution whole cell ultrastructure, that early protein science focused upon the structure and action of isolated molecules. Today we realise the importance (and feasibility) of studying proteins acting as complexes with multiple molecular partners 'nanomachines' in different cell contexts at different stages of cell differentiation, stimulus and response. It is now feasible to design novel proteins 'from scratch' and also targeted 'binder' molecules. Protein science is heading towards an integrated perspective of cell biology with previously unimaginable means to manipulate it. This exciting time could only have been reached through the developments reported here.
Recognition of proteins as discrete molecules
1838 It was the contribution of chemists - especially from Germany and Sweden, that founded our present understanding of molecular biology. Swedish chemist Jöns Jacob Berzelius recognised the distinction between organic and inorganic chemistry and coined the term protein (Gr. prota: of prime importance!)
1840 Friedrich Ludwig Hünefeld discovered crystalline material in earthworm blood, leading to the identification of haemoglobin (Hb). Otto Funke showed how to produce human Hb in 1851. Along with 'albumins' from blood and egg white, the crystallisation of these proteins proved they were distinct molecules, not just mixtures: many chemists thought they were colloids.
1902 Emil Fisher & Franz Hofmeister both discover the peptide bond.
1926 James B. Sumner crystallized urease, proving that enzymes are proteins.
1934 John Bernal and Dorothy Hodgkin discovered that protein crystals gave distinct X-ray diffraction patterns, which proved that, although clearly large and very complex, protein molecules were ordered just as small organic molecules are. This led them to speculating that, some day, protein 3D structures could be solved by this technique.
Swiss chemist Jöns Berzelius. Originator of the term: protein
The first 3D protein structures
1958 John Kendrew published a low-resolution (6Å) crystal structure for myoglobin -- the first folded protein 3D structure. The following year, Max Perutz reported the 3D structure of haemoglobin at 5Å resolution. Being mostly α helical it was possible to trace the polypeptide backbone. One of my later mentors at Oxford University (Prof A.C.T. North) was a co-author on the seminal Nature paper. Tony North went on with Sir David Phillips to report the structure and mechanism of action of hen egg white lysozyme - the first 3D enzyme structure solved.
1969. When seeking a lab to work for my PhD, I had interviews with Max Perutz and also David Blow who had just published the structure of chymotrypsin, but decided that I'd prefer to work on an unknown protein structure, rather than haemoglobin, So I chose to try and solve the 3D structure of the human protein, transthyretin (a plasma carrier of thyroxine and retinol at the Laboratory of Molecular Biophysics where David Phillips was the Prof.
The early protein crystallographers, a number of whom I was lucky enough to meet or work with, formed a 'diaspora' in the 60s and 70s, taking protein X-ray crystallography know-how to research labs all over the world.
The low resolution structure of haemoglobin. The 4 subunits (2 α, 2 β) are coloured separately and the location of the 4 oxygen binding iron chelating haem groups indicated. Haemoglobins from many creatures have now been solved experimentally.
Extraction and purification
The extraction and purification of proteins for in vitro studies was a challenge faced by many protein researchers (who were by then calling themselves biochemists). Proteins of interest were generally present within complex mixtures and often at very low concentration. Very many scientists: PhD graduates, post-docs and established scientists contributed to present day purification techniques. These techniques exploited the size and shape of proteins, solvent solubility, surface properties including net charge and their specific affinity for small or large ligands (now popularly called 'binders'!)
At first, fairly crude purification techniques were employed guided by assays to follow fractions containing the desired activity. Differential precipitation using inorganics like ammonium sulphate (salting out) was useful especially with bulk tissue extracts of plasma. Another early method was ultracentrifugation which separated protein molecules by size and shape. Thermostable proteins were sometimes recoverable by heating extraction media to precipitate thermally labile contaminants. In the 60s and 70s the substantial and pricey hardback 'tomes' of the Methods in Enzymology series were the 'go-to bibles' for purification of many protein molecules. This was, of course, before the democratisation of methods brought in by the World Wide Web!
Column chromatography
1950 Gel chromatography was introduced by Lathe, Ruthven and others. This was comparable to the 'molecular sieves' already in use by organic chemists to separate small organics. Essentially a 'size exclusion' method, it exploited the differential partition of proteins passing through immobile media with pores sizes comparable with proteins of interest.
The first commercial product, Sephadex® was introduced by Pharmacia in 1959. Size exclusion chromatography was rapidly followed in the '50s by ion exchange chromatography exploiting protein surface charge differences and hydrophobic interaction chromatography which connected with exposed non-polar and aromatic sites. Further developments included affinity and immunoaffinity chromatography which were designed to isolate a protein of interest in one or just a few steps.
Many readers of this article may recall setting up columns in cold rooms and shivering for hours to check that their fraction collector was behaving. In 2025 the emphasis on laboratory robotics and automation has (mostly) confined such experiences to history.
Gel filtration column. Pharmacia introduced Sephadex™ as a matrix for size exclusion chromatography and Sepharose™ as a medium for covalently conjugating affinity ligands for protein binding
Affinity separation of biomolecules
1968 saw the introduction of affinity chromatography by Pedro Cuatrecasas and colleagues. This used the simple expedient of covalently linking the specific ligand for a protein of interest to gel beads.
I used this method back in the 70's to purify sialidase (neuraminidase) from bacterial culture media using the sialic acid rich carbohydrates present in ' bird's nest soup extract' covalently attached to Sepharose beads. Sometimes affinity chromatography leads to a single step protein purification. The adsorbed protein is usually recovered using a change of perfusing media composition or the free specific ligand.
Protein isolation sometimes involves sheer luck. We found we could rapidly isolate annexins from 'acetone whole tissue precipitates' sold by Sigma® on phenyl Sepharose columns in Ca2+ eluants, then releasing bound annexins with EDTA in the eluant.
The principal of size exclusion chromatography (SEC) for protein separation Image credit: EMBL, Heidelberg
The huge impact of UV / Visible light spectroscopy
1935 Arnold Beckman created the National Technologies Laboratories—later named Beckman Instruments. Triggered by the American government’s interest in measuring vitamin content in soldier's rations using UV and Vis light, research led to commercialisation of UV-Vis spectrophotometers in the early 1940s. Of these, the Beckman DU spectrophotometer—first sold in 1941—distinguished itself from competing products by delivering more accurate results and reducing analysis time from hours, or even weeks, to minutes.
My old Head of Department, Gilbert Beavan at MRC National Institute for Medical Research UK was a consultant in the development of the Beckman DU spectrometer, applying this to the measurement of proteins. Light intensity is measured from UV-Vis source lamps before and after the light passes through a sample contained in quartz cuvettes or capillaries. The amount of light absorbed corresponds to the molecular concentration in the sample. The instruments are now ubiquitous for measurements and detection of pharmaceuticals and biopolymers.
Nobel laureate Bruce Merrifield (developer of solid phase peptide synthesis) referred to the UV-Vis spectrophotometer as “probably the most important instrument ever developed toward the advancement of bioscience.”
The later development of fluorescence spectroscopy - in essence the emitted light from UV/Vis illumination of biopolymers - has similarly become indispensable to biochemists.
Fluorescence spectroscopy and microscopy
1910. Fluorescence microscopy was developed by German physicists Otto Heimstaedt and Heinrich Lehmann and brought to the commercial market by Carl Zeiss and Carl Reichert. Since that time both fluorescence microscopy and spectroscopy were used to answer a myriad of biochemical questions. Fluorescent chemicals such as fluorescein and rhodamine have enabled a huge amount of knowledge on individual protein structure and action, as well as whole cell and tissue systems.
There cannot be many molecular biologists reading this article who haven't made use of fluorescent labels in both protein and whole cell analysis. I've used fluorescently labelled proteins and fluorescent ligands to investigate protein structure and conformational changes (where fluorescent ligands act as proximity probes by energy transfer from aromatic amino acids to bound fluorophores).
In experiments to follow endocytosis I perfused cell monolayers on coverslips. Fluorescein emission spectrum is sensitive to the pH environment in cell compartments and this was used to determine pH changes in endocytic vesicles and phagosome - lysosome fusion in leucocytes.
Today fluorescent probes are being used in single cell studies and through the exquisite sensitivity of modern detection systems, even single protein investigations.
In the section below, the use of fluorescent reagents is described which enabled the first sensitive amino acid sequence analyses.
Upper image: the flow cell adapted to fit a standard 1cm cuvette in a spectrofluorometer . Lower image: fluorescein labelled yeast cells within phagolysosomes.
Circular dichroism spectroscopy
circa 1850 CD was discovered by Jean-Baptiste Biot, Augustin Fresnel, and Aimé Cotton in the first half of the 19th century. Circular Dichroism (CD) spectra arise from the differential absorption of left- and right-circularly polarized light by optically active (chiral) molecules such as proteins. This has been applied to quantify secondary structures, as seen in the characteristic CD spectral signatures of the α-helices and β-sheets of proteins and the double helices of nucleic acids.
CD spectroscopy offers a rapid, probe-free means of detecting structural changes and stability shifts. Its ability to provide complementary data makes it an essential component of the modern biophysical toolbox, with applications spanning virtually every field of biomolecular research. CD spectroscopy is widely used in academia and the biopharmaceutical industry to study biomolecules, particularly proteins and peptides. CD spectra provide valuable insights into both secondary and tertiary structure.
1977 Illustrated here are spectra I made of native HSA and large fragments containing one or more of the 3 α helical HSA domains. From intact HSA (prior to the 3D structure being solved) a composition of 44% α helix; 10% β sheet and 45% aperiodic polypeptide chain was estimated.
Unlike X-Ray diffraction and to some extent cryoelectron microscopy, CD spectra are obtained from proteins in free solution where their dynamic characteristics are expressed. Many important ligands are themselves chiral and binding studies followed by CD are informative. I used CD to monitor the binding of bilirubin (a breakdown product of haem groups in plasma) to HSA.
CD spectra of human serum albumin (HSA) and purified peptic fragments isolating HSA domains. The bottom spectrum (a) is freshly prepared human protein. All the spectra show the typical double dip characteristic of α helix secondary structure.
Stopped flow method for following protein interactions
1951 Stopped flow as a experimental method for measuring kinetics was introduced by Britton Chance and developed later by many others for following biopolymer interactions on the order of 1 millisecond. Solutions of two reactants are rapidly mixed by being forced through a mixing chamber, on emerging from which the mixed fluid passes through an optical observation cell. At some point in time, the flow is suddenly stopped, and the reaction is monitored using a suitable spectroscopic probe, such as absorbance, fluorescence or fluorescence polarization. The change in spectroscopic signal as a function of time is recorded, and the rate constants that define the reaction kinetics can then be obtained by fitting the data using a suitable model.
Temperature jump was also used to study fast protein kinetics
The general arrangement for following the interaction kinetics of proteins with ligands and other biopolymers
Caged molecules for measuring real time response
Further knowledge about protein function, especially enzyme kinetics and conformational change has been obtained using caged ligands - especially ATP. After a set up containing a protein of interest has reached equilibrium, the caged reactant molecule is released by a laser light pulse. Fast detector output can be measure using spectroscopic systems or high brilliance X-Rays from a source like the Diamond light synchrotron.
Exciting recent applications of caged molecules have explored conformational states of membrane proteins trapped in transition stages using CryoEM to resolve 3D structures. The mechanism of action of receptors and membrane channels can be dissected in this manner.
Assessment of purity and characterisation
Gel electrophoresis
1937 Swede Arne Tiselius developed electrophoresis for the separation of molecules. This gave a sharper assessment of purity than obtainable by passing proteins through gel exclusion columns and monitoring eluants by UV absorption. Later, gels, first starch then polyacrylamide (PAGE) were used as semi-solid separation media, where proteins were segregated according to their molecular weight and net charge. Discs or tubes were used to hold the separation media. Denaturants such as the detergent SDS were incorporated into the gel matrices to dissociate non-covalently bound proteins pre-solubilised in SDS buffers (SDS-PAGE).
SDS-PAGE allowed an approximate estimate of molecular weight (Mwt) using marker proteins of known Mwt run concurrently in separate tubes. At first researchers polymerised acrylamide and the polymer cross-linker bis-acrylamide in glass tubes (I analysed peptic HSA fragments like this in 1970). Proteins were detected by soaking extruded gel in dyes (Coomassie Brilliant Blue was the dye of choice). Comparisons and estimates of Mwt were made by lining up gels extruded from the tubes along with protein markers of known Mwt such as BSA, ribonuclease etc.
1970 Ulrich Laemmli introduced the use of PAGE polymerised as slabs between removable glass plates. This was an enormous blessing to protein scientists as many protein separations and Mwt markers could be run in the same gel. The gel run illustrated is from recombinant HSAs (rHSAs) produced in yeast which were deliberately truncated for various assessments (more on this in my web page on HSA).
SDS-PAGE Separation of native HSA and recombinant, truncated rHSA and recombinant plasminogen activator 2 (rPAI-2)
The convenience of slab gels led to the development of gradient PAGE which could separate proteins and even peptides) of widely varying size in the same gel. Many other types of gel electrophoretic separations were developed (including, much later, long gels to separate polynucleotides in genome sequence analysis). One of the most valuable developments was 2D PAGE where proteins were separated in one dimension by approximate Mwt and in a second dimension according to their net charge, where they would migrate until they reached the region of gel corresponding to their isoelectric point (pI).
Proteins separated within slab gel matrix could be electrophoretically driven onto nitrocellulose membrane where they could be identified by antibody or other biomarker staining (Western Blotting).
I used all these purity assessment and identification methods working with recombinant HSA, and rHSA peptides. The figure illustrates 2D PAGE of adrenal medulla annexins. Western blotting of these gels using antisera obtained from an analogous protein ' calelectrin' present in the electroplax of the ray Torpedo marmorata indicated strong cross reactivity with a subset of bovine annexins. It's also interesting that several annexins were isolated as the partially phosphorylated forms of Anx2, 4 and 6. This is evident from the close splitting of some of the spots in the pI dimension.
2D gel electrophoresis of bovine annexins. MT separation (vertical axis); Isoelectric point focussing (horizontal axis). Splitting of some of the spots is due to partial phosphorylation. Western blots of 1 and 2D gels similar to this indicated 3 to 4 annexins cross-reacting with an antiserum specific for a Ca2+ binding protein from an electric ray. This was our first clue that all these proteins contained homologous amino acid sequences. We later confirmed this by gas phase sequence analysis.
HPLC and peptide mapping
1967 Waters Associates commercialised High-Performance Liquid Chromatography (HPLC), originally developed by Yale researchers as a type of fast liquid chromatography using high pressure pumping (up to 6000 psi). This led in many reports to the alternative term 'High Pressure Liquid Chromatography'. Microparticles of silica could be derivatised with different amphipathic or ionic chemical adducts. Adsorbed molecules were eluted sequentially using acetonitrile of other gradients formed by suitable inline mixing. The game changer for biochemists was the production of alkyl - linked (e.g. C8 or C18-silica) stainless steel HPLC columns.
These were able to resolve peptides in enzyme digests of proteins, which could subsequently be sequenced by mass spectrometry or using a gas phase instrument as described below.
I used this technique for peptide mapping pharmaceutical grade (ex human plasma) rHSA and other proteins.
Single amino acid changes or post-translational modifications ( Ser/Thr/Tyr phosphate groups, glycans etc) could readily be located and identified. HPLC using C18 derivatised silica was also used to identify each amino acid PTH derivative in Edman sequence analysis, also described below.
A comparison of complete tryptic peptide digests of native human and (secreted) recombinant HSA produced in transformed S.cerevisiae. The only difference found in peptide mapping analyses was the presence of an N-acetyl-Methionine group in rHSA produced intracellularly and not cleaved from the protein.
Electrospray Mass spectrometry
1918 Francis Aston developed the first practical mass spectrometer (MS). Molecular ions from an inlet source were accelerated within a vacuum towards an anode. The 'mass spectrum' of the sample is scanned by applying transverse magnetic or electrostatic fields across the molecular beam as it flies towards a detector. What was measured was actually the ratio of mass to charge (m/z). Between 1970 and 1980 MS was increasingly used in biochemistry, especially in peptide mapping but at first there was a limit to the molecular weight of peptides which could be measured in conventional mass spectrometers. This related mainly to the difficulty of introducing such large, polar molecules in the necessary high vacuum.
A major advance occurred when John Fenn and colleagues developed an electrospray source (ESMS). This in essence transferred even large molecules like proteins into the vapour phase within a low vacuum entry port. Here the ionized proteins were accelerated into the mass spectrometer, colliding with neutral molecules which stripped away water and other non-covalently associated molecules prior to m/z measurement.
As protein samples carried many different ionization states, ESMS spectra have a range of many distinct m/z peaks. When these are deconvoluted a highly accurate molecular weight can be determined. Using a prototype ESMS instrument in 1999, I was able to obtain a molecular weight value close to the theoretical value (the sum of the amino acid Mwts) for recombinant human serum albumin. This demonstrated that the rHSA had been correctly expressed and even that all the disulphide bridges had been formed (had they not been, then the molecular weight would have been higher by 2 proton masses per unformed disulphide bond).
Matrix-assisted laser desorption mass spectrometry
A subsequent advance in 1987 by Koichi Tanaka came from 'Matrix Assisted laser desorption / ionisation time of flight (MALDI) spectrometry. In this, large biomolecules embedded in a solid matrix were ablated and ionised by laser pulse. The most widely used instrumentation determined the time of flight of molecular ions from the source to the detector, from which the m/z values can be calculated (MALDI-TOF). In 1988 researchers used a nicotinic acid matrix entraining HSA, obtaining intact molecular ions in MALDI experiments. MALDI-TOF instruments are now the most popular MS spectrometers in research labs.
Identification of protein modifications and mutations
The resolution available through MS ( less than one atomic mass unit) has enabled the detection of many forms of protein post-translational modification and the identification of covalent attachments. Such analysis isn't possible from genome or mRNA sequencing. Also single amino acid substitutions can easily be determined as in the example illustrated here where α and β chain masses are measured in two Hb mutations.
(a) ESMS spectrum of rHSA expressed in an purified from yeast, compared with (b) pharmaceutical grade HSA. In this experiment, the rHSA mass was within 3Da of the calculated mass: Mr=66437 Da. With present instrumentation, the Mr is within 1 Da of theoretical and could easily distinguish isotopic variants. Note the broad peak shape of the pharmaceutical material (b). This HSA will have been from pooled blood plasma and probably represents the presence of different covalent adducts. ESMS of infant (top) and maternal (bottom) haemoglobins with the Montreal-Chori mutation in the β chain. The child (heterozygote) has sickle cell syndrome and shows both ß sickle and ß Montreal-Chori mutations, The mother has both normal and mutant β globins: she is the Montreal-Chori carrier. The α globins are normal.
Amino acid sequence determination
1955 Prior to this time, chemists had determined the amino acid composition of proteins. The technique used was 'force majeure' : heating protein in concentrated HCl at 100oC to break the peptide bonds. The released amino acids were then separated and quantitated by ion exchange chromatography using the amino group reaction with ninhydrin to yield a blue adduct. But the sequence of amino acids was only determined in short peptides using a variety of chemical and enzymatic methods. The first protein sequenced was insulin by the Cambridge chemist Fred Sanger. A small protein, but still an heroic achievement especially as insulin is made up of two polypeptide chains linked by disulphide bridges.
Sanger used the 'Sanger Reagent' fluorodintrobenzene (FDNB) to react with free amino groups of insulin - this gives a yellow reaction product. The labelled insulin was then cleaved into peptides by proteases or HCl. The labelled peptide amino acid compositions were identified by the ion exchange method described. Peptides were treated again with FDNB and once again hydrolysed. The labelled N-terminal amino acids were separated and identified by paper chromatography. The process was repeated many times to build up a complete picture of the sequence. This approach was also used to obtain the complete amino acid sequences of human and bovine serum albumin (over 500 residues!) by James Brown and co-workers in the 70s.
Manual sequencing methods were increased in sensitivity by the introduction of dansyl chloride to label the amino termini of proteins or peptides. The resulting dansyl derivative is highly fluorescent and can be released by acid hydrolysis and detected by 2D thin layer chromatography (TLC). Later a coloured reagent, dabsyl-isothiocyanate (DABITC) was introduced which replaced PITC in the Edman degradation described below, releasing intensely red coloured thiohydantoin derivatives which were again identified by 2D TLC. There is an excellent description of manual protein sequence analysis by Yarwood (Dept of botany, Durham University) in our Protein Sequencing book illustrated below.
There was a long period of manual protein sequence analysis before the automated protein sequencing described below. This was accompanied by the ingenious efforts of 'card carrying' chemists to synthesise and apply novel reagents (like FDNB, Dansyl chloride and Dabsyl chloride) together with labelled amino acid chromatography and identification. A large number of peptides and even large proteins were completely sequenced by these laborious techniques. I can only applaud the ingenuity and sheer labour of all involved. Methods now gone, but should never be forgotten. Once again, it was the organic chemists who gave us the innovative groundwork that lies behind modern molecular biology. I salute them all!
The Edman sequence method
1950 Another Swedish chemist rescued the formerly slow process of protein sequencing. Pehr Edman developed a method of polypeptide sequence analysis which sequentially reported the identity of amino acids from the free amino terminal residue, working toward the carboxy terminus. Briefly, he introduced a reagent - phenyl isothiocyanate (PITC) - which formed a cyclised derivative, breaking the adjacent peptide bond and releasing the amino terminal residue as a phenyl thiohydantoin (PTH) amino acid derivative.
The remaining peptide bonds in the polypeptide were not affected. The PTH forming reaction could then be repeated to release the next amino acid as a PTH derivative. PTH amino acids were extracted into an organic solvent and identified by chromatography. The Edman Degradation, as it was called, was only limited by the availability of the polypeptide under analysis, its loss each cycle and build-up of impurities.
1982 saw American biologist Leroy Hood, together with associate Mike Hunkapiller overcome two of the major inefficiencies in Edman sequencing ( sample size and peptide loss) in the development and commercialisation of the gas phase protein sequencer. The key to automating this form of Edman sequencing was to miniaturise the reaction chamber as a sandwich between cylindrical glass blocks. Peptide samples (less than a few µgms) were loaded onto a coated porous glass disc/ Reagents were applied mainly in the gas phase or were non-polar solvents which minimised loss of peptide from the peptide in the sample chamber.
On a good day, up to 30 amino acids could be determined in polypeptide samples adsorbed to the sample carrier disc. This was a superb and enabling instrument. My lab at the National Institute for Medical Research in London installed the second UK instrument; Mike Waterfield at the Institute for Cancer Research in London obtained the first. Of course there was a huge queue of scientists waiting for instrument time both in my lab and at the ICRF).
Mike was able to obtain sequence from tiny amounts of peptide from the epidermal growth factor receptor (EGFR). This enabled US GenenTech company scientists to clone the EGF receptor. Its (gene) sequence was published in a ground-breaking Nature article. I was able to sequence the first annexin peptides which established these as members of a new Ca2+regulated protein family.
The gas protein sequencer was just one of many scientific instruments appearing from the 1980s which were to prove really enabling to both protein science and genetics.
Edman Degradation: Peptides or intact protein: in this case with N-terminal sequence NH2-Asp-Gly- is reacted with PITC. The peptide bond Asp-Gly is cleaved, the Asp identified as the PTH amino acid by chromatography and the new N-terminal sequence NH2-Gly- is subject to further rounds of Edman degradation. Image credit: Michael Geisow The Applied Biosystems commercial protein sequencer in my lab at Delta Biotech. A breakthrough in sensitivity of protein sequence analysis by the Edman technique. The plastic molecular model on the instrument is one subunit of human transthyretin!
HPLC C18 separation of 100 pmol PTH amino acids indicating the resolution obtainable Image credit: M J Geisow & A Aitken Protein Sequencing - a practical approach IRL Press 1989.
With John Findlay (Leeds Uni) we published what was in 1989 a state-of-the-art lab manual for protein sequence analysis. Still in print I think
Mass spectrometric amino acid sequence analysis
Mass spectrometry was an early method for amino acid sequence analysis. In most cases, proteins were digested into peptides and these were introduced into the MS instrument. Fragmentation of the molecular ions within the instrument produced spectra (MS/MS) from which sequence could be inferred. This was famously used by Howard Morris at ICL to obtain the sequences of two opiate-active 'enkephalins' from vanishingly small samples from brain. This is further described on this website in the page 'The discovery of the opiate peptides and their receptors'. A graphic example of the MS spectra of sequentially truncated recombinant protein is illustrated here.
Mass spectrometry is now an established technique for natural and recombinant sequence analysis, detecting and identifying proteins and posttranslational modification like glycosylation. Sequences inferred from genomics cannot determine such modified forms and direct experimentation is essential.
An exquisitely detailed guide to sequencing by MS was published in our Practical approach book (illustrated above) by Klaus Biemann (Dept of Chemistry MIT, Cambridge MA)
ESMS of C-terminal truncation products from recombinant expression.
Development of instrumentation specific to the molecular biology market
Around the 1980s scientific instrument companies began to and commercialise analytical techniques which were developed in the 'science base' (academic and other public sector science institutions. In essence, they saw the size of the potential market. Certain instrumentation already in use in industry and research organisation - mainly for the characterisation of small organics - was adapted for work with macromolecules. This affected NMR, X-ray diffractometers, optical and electron microscopes. Novel instruments sometimes by start-up businesses spun out of University research departments included protein and DNA sequencers, cell microinjection equipment and miniaturised PAGE in kit form. Around this time the demand (especially in Big Pharma) was for high throughput analysis and laboratory automation. These trends continue to the present day. The cost are high, but it is beyond doubt that the new and accessible instrumentation has greatly accelerated knowledge in protein science since the end of the last millennium.
Protein structure determination
X-ray crystallography
Following the ground-breaking work on the globin structure by Perutz and Kendrew X-ray crystallography became the de facto technology for protein structure determination. The size (Mwt) of proteins analysed increased quite rapidly, which demanded more powerful X-ray sources.
1915 William Coolidge developed the first rotating anode X-ray source, which allowed the generating electron beam to focus on a constantly renewed target surface reducing radiation damage. This source was produced commercially by Philips in 1929. Rotating anode X-ray sources appeared in molecular biology labs from the late 60s and were in use when I worked for my DPhil at the Molecular Biophysics lab at Oxford Uni.
2007 saw the UK Diamond Light Synchrotron Source become available for protein structure determination. The very high intensity X-radiation allowed very rapid data collection which is now used by research visitors worldwide. A new laboratory at the synchrotron complex specialises in data collection from membrane proteins
NMR studies on proteins and their interactions
In 1950 pioneering research by Richard Ernst and Kurt Wüthrich at the ETH in Zurich led to the use of Nuclear Magnetic Resonance (NMR) for investigating protein structure. The 2002 Nobel Prize in Chemistry recognizes accomplishment of the vision, first elucidated over 20 years ago, that solution NMR spectroscopy could be developed as a technique for atomic resolution determination of the structure and dynamics of biological macromolecules in solution.
NMR has since become a key technology for investigation of protein structure, dynamics and binder (ligand) interactions. Unlike X-ray crystallographic structure determination, proteins in NMR samples are subject to free diffusion. This means that NMR provides a different, more dynamic picture of proteins.
Protein structure determination is based on chemical shifts to predict both secondary and tertiary structural regions of proteins. Together with other information including residual dipolar couplings (RDCs) and interproton distances (NOEs) and prior knowledge from other technologies, the size (Mwt) of proteins which could be studies increased year on year in the 70s and 80s. A good background review on this can be found by Cavalli, Salvatella, Dobson & Vendruscolo.
Oxford University was a global leader in the application of NMR to biomolecular structure. In 1971 one of the leading NMR researchers - Raymond Dweck - looked over my shoulder in the Molecular Biophysics Lab and announced (unhelpfully) that "protein X-ray crystallography is dead". How wrong was to be proved to be . .
Solution Structure of Proteinase Inhibitor IIA Superposition of five structures calculated from NMR-derived constraints; drawn from Protein Data Bank ID code 1BUS (top). Ribbon diagram of the energy-minimized structure; drawn from 2BUS (bottom).
1932 German scientists Ernst Ruska & Max Knoll constructed the first electron microscope, Ruska was awarded a Nobel Prize. A critical tool in cell biology ever since, the development of cryogenic electron microscopy (Cryo-EM) in the 70s and 80s. In principle the high resolution (2-3 Å) determination of protein and other biomolecular structures was feasible , but radiation damage due to the observing electron beams limited its application.
In 1984 scientists at the European Molecular Biology Labs were able to determine the structures of a number of viruses vitrified at cryogenic temperatures. Later developments led to breaking the sample resolution barrier of 2-3 Å resolution; high enough to resolve amino acid position and orientation.
Cryo-EM was applied to determine the 3D structure of the light collecting bacterial membrane protein, rhodopsin, gaining Jacques Dubochet, Joachin Frank and Richard Henderson the 2017 Chemistry Nobel Prize. I was delighted to learn of this. I was able to obtain UK Medical Research Council funding for Richard back in 2000 when I managed the LINK protein engineering programme. Richard had been working for years on the difficult task of extracting membrane proteins for structural analysis.
Cryo-EM broke through this barrier. Since then, Cryo-EM has taken over where X-ray crystallography and NMR have to stop: the analysis of intrinsic membrane proteins (which depend on a supporting lipid environment for their 3D structures), and very large proteins and protein assemblies, such as whole viruses, bacterial and eukaryotic protein "nanomachines"
Ribbon representation of bacteriorhodopsin membrane protein structure solved by cryo-electron microscopy top view looking down on the plane of the membrane. Lower view seen in transverse section within the bacterial membrane. The light trapping retinal ligands (purple) a covalently linked to lysine residues. Source: PDB 1BRD
It's beyond doubt that cryoelectron microscopy will solve many questions in cell biology. A considerable number of these concern the structure and action of membrane proteins which enable cell communication with the outside world.
For example, all opioid receptors are 7-transmembrane G-protein coupled receptors with an extracellular binding domain and an intracellular signalling domain. Opioid binding to the extracellular domain activates cytoplasmic heterotrimeric G proteins by opening the Gα α-helical domain (AHD - see image) to enable GDP–GTP exchange (the signal). There has been an unsolved question about how the opiate antagonist naloxone rapidly mitigates opiate signalling. Naloxone is an acute drug for reversing opiate poisoning, yet its mechanism of action had r\until recently remained unknown. Opiate drugs act through activating the μ -opioid receptor (MOR) and its coupled G protein. Saif Khan et al. ( 2025 Nature648 755) used CryoEM to capture snapshots of intermediate conformations of MOR - G protein with bound opioid agonist or naloxone. Their work provided insights into G-protein coupled receptor pharmacology.
The data supports a model in which efficacy is primarily driven by the transition from a signalling-silent ‘latent’ state to an ‘engaged’ state. This transition results in a hallmark ‘active-like’ MOR conformation and an extended α5 helix, despite the AHD remaining closed. The engaged state is structurally primed for subsequent AHD opening, which is essential for GDP release. Stabilization of the latent state probably reduces nucleotide exchange rates, approximating basal activity seen in receptor-free GDP–Gαi1β1γ2, whereas progression through downstream states accelerates GDP release, thereby enhancing G protein activation. Taken from Khan et al. with the author's permission.
Protein structure prediction, interactions and design
From the early days of X-ray crystallographic protein structure determination scientists all over the world attempted to predict the 3D folding of proteins of unknown structure. As the database of experimentally determined structures grew, this looked feasible.
Around 1970 Researchers Peter Chou and Gerald Fasman developed an algorithm to predict runs of secondary structure (α helix and β sheet) using indices calculated from the frequency of appearance of residues in experimentally reported protein structures. They termed these indices propensities. For instance, the amino acids Ala, leu and ile would appear often in α helix, whereas glycine and proline would appear more frequently at the termini of helices. This approach was modestly successful and I developed a further slight improvement in the predictive accuracy of amino acid propensities in 1980 using the then increased number of known structures and splitting them into α or β class proteins.
Secondary structure could be approximately located in sequences and homologies with known protein structures were encouraging, but tertiary structure determination remained elusive. Attempts to fold proteins ab initio using molecular dynamics approaches ran into computation time limitations. Yet, logically, proteins fold rapidly following biosynthesis and also rapidly following experimental unfolding by denaturants in the lab. What did nature know that we didn't?
Critical Assessment of Techniques for Protein Structure Prediction
2018 saw AlphaFold generated protein structure predictions enter the international CASP (CASP13) competition. The results led founder and organiser, John Moult from the University of Maryland, USA, to pronounce that the long time protein folding prediction had finally been cracked.
Timeline: CASP 13 AlphaFold1 led in the most difficult category, where no homologous protein structures were available which could act as templates. CASP14 (2020) AlphaFold2 results led the organisers to declare that the protein folding problem had been solved for single chain proteins.
Since AlphaFold’s breakthrough in 2020, the landscape of protein structure prediction has diversified rapidly. Although the management of the originating commercial business, Deep mind, might have restricted access, they have made the programme open source in the form of the AlphaFold server and created the AlphaFold database of millions of predicted protein structures. In the spirit of this openness, follow on AI supported software is similarly available through online access. DeepRosettaFold, ESMFold, OmegaFold, AlphaFold3, and BETA are the most notable entrants, each pushing the boundaries of speed, accuracy, and biological relevance.It has come to cover protein and nucleic acid structures and interactions. These tools extend or compete with AlphaFold’s deep learning approach, each with unique strengths in speed, accuracy, or integration with other biological data.
Less validated on large benchmark sets compared to AlphaFold
AlphaFold3
DeepMind (2025)
Expanded deep learning framework
Improved accuracy, predicts complexes and interactions, broader biological scope
Computationally intensive, still evolving in accessibility
BETA
Launched alongside AlphaFold3 (2025)
Specialized deep learning
Focuses on flexible/disordered proteins, challenging cases
Early stage, less widely adopted yet
Predicting protein folding and interactions
2025 has seen the development of the programme Boltzgen. This approach isn't another structure prediction tool like AlphaFold or RosettaFold. Boltzgen has excited the fundamental and applied biomolecular science communities. BoltzGen is a single all-atom model for designing miniproteins, peptides, nanobodies. It represents a new generation of generative biology models that unify protein structure prediction with binder (a.k.a. protein partner or ligand) design.
Whilst AlphaFold and its successors focus on predicting how natural proteins fold, BoltzGen is designed to create entirely new proteins and peptides that can bind to specific biological targets, while simultaneously predicting their folded structures in the bound states. The image right, taken from the paper of Hannes Stark et al. shows a cyclic peptide (a frequently synthesised type of peptide drug) designed to fit the extremely high affinity biotin binding protein streptavidin from Streptomyces avidinii
BoltzGen design of a cyclic peptide to bind streptavidin. The target protein is rendered blue, the designed binder is rendered yellow. The brown peptide in the target is predicted to change conformation on binding (flexible). Image credit: BoltzGen: Toward Universal Binder Design Hannes Stark et al..
Where is protein science headed today?
As little as 10 years ago it would have been guesswork, wishful thinking even, to make predictions about the future of protein science. But we can see from the technologies that have been and are still being developed, that we can now ask deeper questions about the functioning of living organisms which results from their proteomes. In the 1950s-60s the focus of biochemists was on individual proteins, studied mainly in isolation in vitro. Increasingly we are studying protein and other biopolymer assemblies, both in vitro and in vivo. We are now able to identify and define cell interactomes. These are giving deep insights into the details of protein function in health and pathology.
It's notable how the disciplines (techniques and knowledge) of geneticists and protein scientists have merged to give the present day integrated molecular biology that we may now confidently predict will give us a far more complete understanding of the cell. And on the basis of that, of the whole organisms that the cells make up.
So what are these remarkable polypeptides?
Polypeptides are unbranched biopolymers of amino acid monomers chemically linked in linear chains by the peptide bond (see illustration). In natural polypeptides there are 20 amino acids (biochemists often refer to these as 'residues'. These are essentially the same throughout the living world. This must mean that the evolutionary 'decisions' behind their selection must have been before the time of the last common ancestor (LCA) of all life on earth. Why 20 amino acids? Why not more or less? What evolutionary pressures went into their selection? I have included more discussion on this in the section on Adventures in polypeptide space. Of course we cannot presently answer such questions.
The existing set of amino acids have chemical side chain characteristics (the R group in the graphic) which include: hydrophobicity and hydrophilicity (water hating and water loving); acid and base character (polarity); aromatic and aliphatic groups and intra-chain links (disulphide bridges) formed by the sulphur containing side chain of cysteine. In natural proteins a rationale for each of these characteristics can be seen from detailed studies of polypeptide structures and functions.
Polypeptides with defined functions are generally classed as proteins, noting that many proteins are composed of more than one polypeptide chain. Most proteins comprise polypeptides longer than an arbitrary 50 amino acids. Size matters! Proteins are important molecules. Massively so, since they give structure, function and animation to all known life on earth: microorganisms, plants and animals. But as we will see later, natural proteins represent a tiny fraction of theoretical polypeptide space.
The peptide bond upon which we all depend (highlighted in green) links the amino acids in proteins. In the above schematic, two generic amino acids are shown:R represents any one of the 20 natural amino acids.
Amino acid sequences determine polypeptide 3D folding and function
The folding of polypeptide chains is dependant upon their amino acid sequences. The dominant forces which structure polypeptides in 3D are: hydrogen bonds mainly between the atoms which make up the polypeptide chain (often called the backbone in the literature); Hydrophobic interactions between non-polar amino acid side chains which are generally buried away from solution; Salt bridges between polar side chains and Disulphide bridges formed between cysteine amino acid side chains. There are other interactions too called π - π between buried aromatic amino acid side chains. Extensive disulphide bridging can be seen in the serum albumin molecule illustrated below.
Studies of proteins from organisms which live in very hot or high salt environments (extremophiles) generally have more extensive hydrogen bonds, disulphide links or salt bridging between polar side chains. But all these interactions are absolutely dependant upon an aqueous environment. Proteins won't fold in organic solvents. You just end up with an aggregated, structureless mess. It's this universal observation which underlies the (expensive!) search for water on cosmological bodies. Without water, life, at least in the way we have defined it above, just couldn't exist.
Some proteins are exceptionally highly conserved.
Such proteins are considered universal because they have changed very little in amino acid sequence and 3D structure throughout nature. They're especially conserved from microbes to man because their properties and functions are fundamental to life. They are key players fulfilling the eight criteria for life listed in the introduction to this article. Evolution has 'tinkered' extensively with many proteins, changing their amino acid sequences and shape and duplicating or deleting segments of polypeptide chains. But not these proteins: the penalty for even single amino acid changes through mutation in these is usually death! A particularly appropriate maxim here being 'If it ain't broke, don't (even think to) fix it'.
Copying and synthesising DNA (genome replication) and operating cell division (mitosis)
Energy harvesting and conversion
Synthesis and maintenance of cell walls
Synthesis of or import of essential essential chemicals from the environment (a.k.a nutrition)
The most highly conserved proteins
The proteins which carry out the functions in the list on the left were invented before eukaryotic life evolved. Nature had already found biochemical solutions to some of the trickiest chemistry on earth!
Human serum albumin (HSA) is made up of 584 amino acids (top). A globular protein, very soluble in buffered solution, it's very stable; 17 disulphide bridges (S-S) contribute to its stability and probably contribute to its relatively long half life in our blood plasma. HSA illustrates several of the points made in this article. Its surface offers many opportunities to bind (and transport) molecules including fatty acids, steroids and several pharmaceuticals. HSA's polypeptide is folded into 3 similar domains (discussed further below). Before HSA 3D structure was determined, I was able to separate complete or partial domains using enzymes. The separated large fragments of HSA retained both their fold (as far as could be determined in the 1970s!) and their binding functions.
Protein surfaces have evolved for interactions
It's generally the surface of proteins which underlies their functions.
Amino acid side chains fully or partially exposed to aqueous environment interact, usually reversibly, with biopolymers such as proteins, nucleic acids and polysaccharides. Interactions are also made with small molecules like cell lipids, neurotransmitters, hormones, drugs, ions including Ca2+, many metals, vitamins and other co-factors). Interaction sites (also called binding sites) vary from small regions of proteins to large areas (which is generally the case in those critical proteins in the list above).
The gross shapes of proteins was inferred at an early stage from their physical properties. These were classically determined by methods including comparative measurement of migration rate through gel matrices or sedimentation speed during ultracentrifugation. Globular proteins (a class to which the serum albumin illustrated and many others belong) are roughly spherical or elliptical. There appears to be a limit to the physical dimensions of globular proteins. This probably relates to unfavourable energetics burying extensive segments of polypeptide chain and still end up with a stably folded structure.
Instead, large globular proteins generally fold into a series of compact regions called domains. These bury modest regions of polypeptide chain away from surrounding media. The polypeptide chain of HSA (584 amino acids) folds into 3 similarly structured (homologous) domains. Whilst the architecture of HSA domains is essentially the same, different binding sites arise from dissimilar amino acids within HSA domains.
The giant protein titin, which gives elasticity to skeletal muscle, has over 34,000 amino acids. These are incorporated into several hundred homologous domains which share their folding with domains present in antibody proteins. It's a feature of natural proteins that homologous amino acid sequences (and their corresponding 3D folds) have been identified in the proteomes of all lifeforms. This is often, but not always, found to correspond with similar properties or functions. This characteristic has usefully enabled protein functions to be inferred from genome sequencing alone.
The cell membrane binding proteins known as annexins (described below) are folded into either 4 or 8 very similar domains. It has been assumed that such homologous domains arose from gene duplication events. Evidence of the 'repurposing' of analogous folding and function can be found throughout nature.
No size limits apply to the elongated, fibrous class of proteins like collagen, keratin and elastin, which are generally made up of repeating short polypeptide sequences. These fibrous proteins associate side by side to form microfilaments. Large surface areas of such proteins are exposed to the aqueous media. Generally these microfilaments assemble into bundles through lateral interactions such as hydrogen bonds. Where globular proteins take up extended forms - as in the shape and motility conferring cytoskeleton of cells - reversible self-association occurs. Examples include actin which self - associates to give filamentous actin F-actin and tubulin which yields microtubules. This class of microfilaments, composed of globular protein units, just like the 'professional' fibrous proteins, can also associate laterally to form filament bundles.
Knowing polypeptide 3D structures is essential to explain their properties and functions
To understand what makes life 'tick' it's necessary to determine protein structure and function. Knowing the 3D structure of proteins is critical for enabling applications like the development of drugs. The recognition that proteins have distinct folding of their polypeptide chains was first proved for myoglobin and haemoglobin by Max Perutz and John Kendrew at Cambridge in the 1960s. It's worth noting that the isolation and purification of proteins at that time was still a considerable challenge. With present day techniques and gene cloning, just about any protein can be prepared in pure form for structural studies. 3D protein structures can then be determined by X-ray crystallography or cryoelectron microscopy. Remarkably accurate computer protein structure prediction methods have been developed in the last few years as described later.
Core architectural elements of protein folding
A huge range of protein 3D structures are now known (upwards of 123,465 structures are described in the Protein Data Base (PDB)! This has allowed key common features of folded polypeptide chains to be established. These are the well known two secondary structures:. Alpha helix - in which hydrogen bonding between successive peptide bonds pulls the polypeptide chain into a right-handed spiral. One of the earliest stages in folding a new protein molecule, even while this is still being synthesised on the ribosome, is nucleation of alpha helix. Beta (pleated) sheet in which hydrogen bonding between adjacent polypeptide chains, pulls them into a sheet structure. A beta turn (beta bend or reverse turn) is the shortest length of polypeptide which can bridge elements of the beta sheet. It's generally these secondary structural elements, with amino acids often partially or fully buried away from the aqueous environment, which give shape to the protein's functional surface discussed above.
Are there other regular secondary structures?
Our knowledge of the common regular secondary structures (alpha helix and beta sheet) derives from the data set of natural proteins. If we ever go deeply into general polypeptide space, we might find others!
Other regions of polypeptide chains aren't characterised by such regular structures, but nevertheless adopt similar conformations in homologous proteins. The association of these secondary structural elements to give the folded proteins seen in the PDB is called tertiary structure. A large number of proteins are formed by the further non-covalent association of separate polypeptide chains as in haemoglobin and transthyretin, described below. This association is known as the protein's quaternary structure.
ALPHA HELIX
L-Amino acids linked by the peptide bond form a regular right handed helix. This conformation of polypeptide chain is brought about by regular intrachain hydrogen bonds between the main chain N-H groups and carbonyl O groups.
BETA SHEET
The beta conformation of polypeptide chain is illustrated by hydrogen bonds formed between non-consecutive regions. This form of interchain hydrogen bonding forms sheet like structures. Adjacent sections of polypeptide chain may run parallel or antiparallel to each other relative to the chain termini. This common form of secondary structure is referred to as beta pleated sheet because of the regular kink in the polypeptide chain at the location of the tetrahedral carbon atoms. The notional surface created by such beta pleated sheets is frequently twisted. This can be seen in the transthyretin structure illustrated below
Natural proteins can broadly be classified into structural classes
Globular proteins can be loosely described as alpha class if they're folded almost completely with alpha helix secondary structure as shown for human serum albumin and the annexins. Or they can be regarded as beta class if their polypeptide chain is almost all folded as beta secondary structure as first shown in human transthyretin, whose structure I solved by X-ray crystallography in 1970 (illustrated below). These classes of protein are not particularly distinguished by their amino acid content. It;s their amino acid sequences that lead to the distinct alpha or beta classes. With a huge number of protein structures now determined, it can be seen that these proteins represent two architectural extremes. Many proteins have a mixture of these secondary structures, or sometimes none at all.
Fibrous proteins (like collagen) which frequently have monotonously repeating amino acid sequences with few hydrophobic side chains, tend to form extended alpha helices leading to their description as 'coiled' proteins. Such extended helices often self-associate as dimers or trimers which twist around each other. These assemblies are correspondingly described as coiled coils. Although fibrous proteins are generally thought of as having structural functions, some of these coiled coil assemblies participate in key cell membrane biology. One example are the members of the BAR protein family. These self associate as coiled coil structures which bind to membrane lipids and support the bending force important in pinching off cell membrane as vesicles in cell transport processes.
HUMAN SERUM ALBUMIN
An alpha class protein. Human serum albumin (HSA) which I studied in depth at the National Institute for medical research (London UK). Its polypeptide chain is essentially all folded into alpha helix secondary structure connected by short sections of polypeptide chain. The molecule is folded into 3 domains (shown as coloured sections of alpha helix).
TRANSTHYRETIN DIMER
A beta class protein. Human serum transthyretin dimer structure. I solved this molecular structure for my D.Phil. The protein is actually made up of 4 identical monomers. Half the protein (a dimer) is illustrated here. A pair of such dimers interact with apposed beta sheets forming the tetrameric assembly which constitutes the binding site for thyroxine (see the web page on transthyretin). A number of mutant transthyretins with amino acid changes have been reported. Many of these changes interfere with the protein's ability to form the tetrameric structure illustrated leading to the formation of amyloid plaques in tissues. Small changes in a single polypeptide can lead to big consequences for the human expressing it.
Polypeptide chains folded into the beta secondary structure form the structural 'core' of many proteins (like the transthyretin tetramer illustrated, which is composed by the opposition of a pair of transthyretin dimers illustrated earlier. The obvious pocket between the paired dimers forms a binding site for thyroxine which this protein transports in our blood plasma. Hydrophobic amino acid side chains are usually buried within globular proteins.
But there are many proteins where this isn't the case. In beta proteins embedded in cell membranes, hydrophobic amino acid side chains face into the hydrophobic environment of the membrane lipid bilayer and the beta sheet 'staves' wrap round to form a beta barrel The central beta sheets in the transthyretin tetramer illustrated gives a rough idea of what such a beta barrel looks like.
Beta barrel proteins can have many beta staves and form pores. Amino acid side chains in the central region of such beta barrel proteins are generally hydrophilic.\, lining the pores. Dynamically inserted into membranes by bacteria or components of the blood plasma complement system these can lead to cell death through leakage of cytoplasm. The beta sheets in each transthyretin monomer form an extended sheet within the dimer. Misfolding of this protein can lead to the formation of extensive beta sheet resulting in dangerous amyloid aggregates leading to cardiomyopathies and other pathologies.
Alpha helix bundles which expose hydrophobic amino acid side chains on the protein surface are also present in membranes. The large family of receptor proteins are in this category where a single hydrophobic alpha helix spans the cell membrane with the receptor component facing the extracellular environment and the signalling component facing the cytosol. Exposure to their specific ligands frequently leads this class of receptor to form dimers or trimers in the membrane. Another instance of protein machinery assembling through association.
Membrane proteins may also be formed from helical bundles spanning in the membrane. An early example where the 3D structure was solved in rhodopsin, which through its bound retinol co-factor is sensitive to photons. Other transmembrane helical proteins act as ion pumps, channels and gates and are critical in nerve transmission and energy generation.
TRANSTHYRETIN TETRAMER
Quaternary structure The complete transthyretin tetramer looking down the space occupied by side chains projecting from the two beta sheets. Closely apposed at the centre of the tetramer as they open out towards your viewpoint (through the beta sheet twist) they form binding sites for the hydrophobic thyroxine and similar ligands, which the protein transports in blood plasma.
Common functional domains: evidence of long protein evolutionary past
Inspection of many protein sequences and 3D structures indicates that they have arisen by gene duplication events as large sections of contiguous polypeptide chain adopt very similar 3D folding. This is sometimes (but not always) evident from the amino acid sequences as in HSA illustrated above and annexin A5 illustrated here. Annexins are present in almost all eukaryotic life forms like us and may even be present in some prokaryotes like the bacteria. Some plant annexins are highly homologous in amino acid sequence with annexins in us. Such a degree of conservation suggests annexins evolved at a time before the divergence of plants, animal and the other living taxonomic kingdoms.
The 3 protein examples in this article
Apart from my having worked extensively on the proteins described here (each has a web page on this site), they are now are playing a significant role in medicine. Transthyretin is responsible for certain amyloidoses, Serum albumin has been approved as an excipient in pharmaceutical and stem cell approaches and annexin 5 is in late stage clinical trials for Retinal Vein Occlusion.
ANNEXIN: A 4-DOMAIN MOLECULE
Annexin 5: the molecule is composed of 4 domains made up by 5 alpha helices. I have rendered the polypeptide domains with different colours. In most annexins these domains bind calcium ions and subsequently form complexes with membranes, interacting with acidic phospholipids like phosphatidyl serine.
Aligned amino terminal sequences of the 12 human annexins (hANXAn) and 2 plant annexins (cANXD) for comparison. Shown in red are the amino termini and shown in yellow are the first two alpha helical regions of an annexin core domain. Despite the great evolutionary distance between the plant and animal kingdoms, the amino acid sequences in the annexin cores (yellow) are notably highly conserved, As discussed later, the absence of homology of the amino termini (red) is significant as these regions are disordered (flexible). An arrow points to phosphoglycerate (used to localise one binding site for membrane lipids). The green dots are calcium ions which are important effectors for bringing together annexins and membrane.
Closely similar (homologous) polypeptide domains which evolved for specific functions are frequently found in very different proteins. This is evident in the annexin family of calcium-regulated lipid membrane binding proteins. The image (left) above highlights the four annexin Ca2+ and lipid binding domains (differently coloured for emphasis). In the right hand image annexin amino terminal sequences shown (in red) are completely dissimilar in fold and function from the highly conserved helical core regions (highlighted in yellow). It's now known these amino terminal amino acid sequences are highly flexible and specifically bind different biopolymers. These include proteins, RNA and glycans. Here nature is showing us duplication and repurposing of an important functional module (called the 4 domain annexin core). The critical function of this core is to bind to cell membranes. The very different amino terminal 'domains ' then enable annexins to link other cell components to cell membrane surfaces. The annexin family is described in depth in another section of this website.
The repeating homologous polypeptide loop structure of HSA and the presence of 3 structural (and functional) domains was described earlier. Another very important class of proteins which are built up of distinct domains are the proteins of our immune system which include the antibodies. In this case the antibody domains have similar folds but quite distinct function. There's abundant evidence that polypeptide domains have been shared by 'horizonal gene transfer' both within and between organisms. Bacteria share genes in this manner which is often the basis of transfer of antibiotic resistance. In the past viral infection may have been responsible for such gene transfer in higher (eukaryotic) organisms too.
Polypeptide chains adopt 3D folds but these are dynamic.
Biophysicist: 'What skills did evolution teach you in folding lessons?'
Intrinsically Disordered Protein: 'Reeling & Writhing & fainting in coils' quote - Lewis Carroll: Alice's Adventures in Wonderland
As well as small molecules (ions, drugs) many proteins interact with other biopolymers, most often proteins, but also DNA, RNA, glycans or the lipids of cell membranes. Flexible backbones are required. Polypeptide chain movement between two or more folded states is intrinsic to the functions of many proteins. This specific rearrangement of polypeptide chain is described as conformational change. It may involve movements of the discrete polypeptide chains in multi-subunit proteins like haemoglobin (Hb). This was seen as relative movements of the paired Hb alpha and beta chains on binding oxygen resulting in oxygen affinity to be co-operatively increased. Movement of the polypeptide chain with folded proteins is also common. Some enzymes undergo extreme conformational change during chemical bond making or breaking. The two domains of the glycolytic enzyme phosphoglycerate kinase (solved by my former colleague Phillip Evans at Oxford Uni.) kinks by 56º at a 'hinge' region to bring two active sites into contact to form the critical cell energy molecule ATP. Hinge regions are also present in immune proteins like the antibodies which allows the antigen binding sies (Fabs) to span different targets. The amino terminal regions of some annexins become available for binding to other protein partners on calcium ion binding as described below.
Polypeptide chain movement is as important as fold
Proteins have intrinsically disordered regions or the whole molecule may lack the order seen in the many crystal structure studies. Polypeptide chain movement is crucial for enzyme catalysis binding other protein partners in signal transduction and assembling complex nanomachinery
Polypeptide chain flexibility, especially at the beginning and end of polypeptide chains, is common and is often critical to protein function. Protein structures determined by X-ray crystallography or cryoelectron microscopy (CryoEM), by the very nature of these methods, show 'frozen' or time-averaged views of the polypeptide chain. Proteins in the PDB (Protein Data Bank) are shown with static folds, but that isn't how they exist in cells or researcher's test tubes. Under physiological conditions proteins are essentially dynamic. This includes movement of buried secondary structure. Some regions of polypeptide chain are particularly mobile. These are often the polypeptide chain ends (the amino or carboxy terminus). Chain termini can be in such rapid motion they are invisible in structures determined by X-ray crystallography.
Unlike crystals of small small molecules, protein molecules are surrounded by water in their crystals which permits chain motion as well as the free diffusion of ions and organics. In solution during lab experiments and within cells the chain termini and sometimes the entire protein molecule may lack regular order. These are now referred to intrinsically disordered proteins (IDPs). The amino termini of the three proteins I studied in depth (HSA, transthyretin and annexin) were each 'unstructured' under particular conditions. Each region has an important role to play in these proteins functions.
ANNEXIN IN THE PRESENCE OF CA2+IONS
X-ray structure of Annexin A5 in the presence of Ca2+ions. (green dots). A large section of the amino terminal polypeptide region (arrowed) is not visible as it is in random motion. The small molecule in the upper right of this image is phosphoglycerate. This marks one of the lipid head group binding sites.
ANNEXIN IN THE ABSENCE OF CA2+IONS
At low levels of Ca2+ the amino terminal region folds up into alpha helix conformation and can now be seen in the X-ray structure above (marked by a number of arrows)
Group hugs!
Polypeptides with the intrinsically disordered regions (IDRs) mentioned in the last section are frequently involved in assembling molecular machines and subcellular compartments. These assemblies may comprise a single type or more usually a number of different polypeptides.
Capsids, chaperones and condensates
The earliest type of protein assemblies observed by electron microscopy were the 'capsids' which encase the genomes of many viruses. polypeptides in capsids generally form regular polygons: the symmetry arising from packing together many similar polypeptides. In bacteria, assemblies of polypeptides form the ''outboard motors' which give these single cell organisms motility.
Notable structures formed by polypeptide self association are the chaperonins which form hollow chambers where damaged or nascent proteins can fold or refold. Also proteosomes where proteins enzymically tagged as 'end of use' can be disassembled, cut up by proteases and the amino acids recycled. A more recently identified structure within human and other animal cells is the 'vault' illustrated in the introduction to this article. There are many more instances, known and probably yet-to-be-discovered molecular machines are built up of dissimilar proteins, with different functions. Very well understood examples of molecular machines include the proteins involved in the replication and regulation of DNA transcription and the separation of daughter chromosomes in mitosis.
Proteinaceous compartments
Prokaryotes carry out most of the functions listed in the earlier definition of life within the cell cytoplasm. The more complex eukaryotic cells require compartmentation of functions. This is addressed by many different membrane enclosed organelles (the mitochondrion, nucleus, lysosomes etc.). But some important eukaryotic functions occur within membrane-less regions called biomolecular condensates (the nucleolus is one example). Proteins and RNA within such regions assemble through multivalent association including most of the types listed in the earlier section on protein interactions.
These regions are said to be liquid-liquid phase separated. Diffusion in and out of these regions is not restricted by membranes, which require transport systems for proteins or RNA to cross them. Maybe the biomolecular condensates represent a legacy from the earlier prokaryotes which lack organelles. The intrinsic disorder of polypeptides which are involved in biomolecular condensates is important to their formation. See: A Framework for understanding functions of biomolecular condensates on molecular to cellular scales Lyon A. et al.At least one of the annexins - AnxA11 - participates in biomolecular condensates. AnxA11 has a very extended amino terminus rich in proline residues which leads to a large disordered region (IDR).
Modification of polypeptide amino termini is especially important in biological functions
Intrinsically disordered regions of proteins can be sites of enzymatic addition or removal of phosphate or other chemical groups (known as post-translational modification. By this means cells exert regulatory control on proteins or protein assemblies. The amino termini are frequently subject to such modification. Amino termini of histone proteins which bind and compact segments of DNA within chromosomes are modified by acetylation and other covalent change. This contributes to theepigenetic regulation of gene expression. Amino terminal sequences are also frequently subject to alteration by alternative splicing of the mRNA which specifies them. Addition or removal of sections of polypeptide chain in this manner changes the functions, distribution or interaction with other biopolymers in cells. This is especially seen in the annexins. This family of proteins is highly 'promiscuous' in biopolymer interactions through post-translational modification and alternative splicing!
Intrinsically disordered regions may also be sites of enzyme cleavage of the polypeptide chain. Many proteins are initially synthesised with 'signal peptide sequences' which are recognised as transport addresses directing proteins to subcellular compartments or secretion pathways. These peptides are excised during transport. Other proteins - especially hormones directed to secretion pathways - have 'propeptide sequences', which are similarly subject to enzymatic removal prior to release from the cell.
How many natural proteins has evolution created?
Most natural proteins range in size between 80 and 1000 amino acids (often called residues by biochemists), but there is a noticeable size peak at ~340 residues which corresponds with the sizes of the large family of receptor proteins which sit in our cell membranes. Working out an estimate for the number of different protein sequences is a combinatorial exercise. Taking a polypeptide chain of a modest 200 amino acids, as an example, since 20 different amino acids can be present at each point, there are theoretically 1.6 x 10260 different natural polypeptides.Even given such a vast theoretical number of unique polypeptide sequences, there is further diversity in nature to be accounted. Over 200 'post-translational' chemical modifications have been recognised. These include the attachment of chemical groups like phosphate or sugars.
Prediction of polypeptide and protein 3D folding
The stumbling block in polypeptide and 3D protein structure prediction (for many decades) was to determine how localised secondary structures (alpha helix and beta sheet) and interconnecting peptide chain associated to form the folded protein tertiary structure. It was evident that the amino acid sequence determined the folding of polypeptide chains. Many purified proteins could be unfolded and refolded with full regain of function. A race was on to try to determine the folding of proteins from their amino sequences. Success was seen as 'the holy grail' of protein science.
Prediction of regions of polypeptide chain which were likely to form secondary structures was tackled first with modest success by a number of groups. The approach assigned each amino acid a propensity (a preference) to be present in alpha, beta or no secondary structure. This was essentially a statistical method.
Structure Prediction - a blast from the past!
Following the lead of workers Chou and Fasman, on the 70s, I recalculated these amino acid propensities for all alpha or all beta protein classes of experimentally determined proteins. The values I obtained gave a modestly improved prediction of secondary structure. Much later, with Dr Willie Taylor ( then at Birkbeck College London), We were able to localise the alpha helical regions of the annexins I was working on Our helix locations were later shown to correspond reasonably well with the structure of annexin determined by X-ray crystallography (Image below). Prior to the publication of the structure of and annexin by Robert Huber in Germany, Willie and I then attempted to predict the 3D folding of one of the annexin domains using the latest prediction methods. We used the known 3D structure of the known Ca2+ binding proteins, which also having a bundle of alpha helices as template for our prediction. The Ca2+ binding site was at at turn between tow helices the called an 'EF hand' When the structure of rat Annexin A5 was published we could see that our prediction of a bundle of 5 alpha helices was correct, and or location of the Ca2+ binding site was also correct, but our estimate of the tertiary structure (the overall fold) was incorrect. In fact the annexins were found to belong to a new class of Ca2+ binding, proteins. We were close, but no cigar! As described later in this article, protein secondary and tertiary structure prediction has now become relatively routine.
PREDICTION OF ANNEXIN SECONDARY STRUCTURE
Our early prediction of the location of secondary structure (alpha helix in this instance) from amino acid sequence for the 'core' region of annexins. Several annexin sequences are shown with particularly homologous sequences boxed. A1,2 & 4: AnxA1, AnxA2 & AnxA4. TC: an annexin from the electric ray; Torpedo marmorata Our predicted number and location of alpha helical regions (blue lozenges) is shown against the actual regions located in the structure determined by X-ray crystallography (black lozenges). We had quite good correspondence with the experimentally determined helix location. Our match for the 2nd helix bordering the first Ca2+ binding site was accurate.
OUR ATTEMPTED PREDICTION OF ANNEXIN TERTIARY STRUCTURE
Prediction of the tertiary structure of a single annexin domain (70 amino acids)
Developments in AI have revolutionised protein structure prediction
An international competition CASP was held biennially where researchers submitted predicted 3D protein structures subsequent to publication of experimentally determined ones. Over decades these predictions improved slowly. The big breakthrough came with a computer programme AlphaFold. This approach by researchers from the Alphabet company Deep Mind used Machine Learning and their approach has been applied to structure all proteins from the sequences in public databases. The 3 researchers behind this advance were awarded the 2024 Nobel Prize in chemistry. From the deluge of research papers and open access resources following this, we might describe this AI breakthrough as a tipping point for protein structure determination and de novo protein design.
There are now a number of competitive offerings online, mainly open source, for predicting polypeptide folds from amino acid sequence. All methods are subject to continuous improvement through learning from ever-increasing databases of experimentally determined structures and physical realities of atom-atom interaction. As with the genetic code itself, there's some amino acid redundancy in folded proteins. Why do I claim this? Well in the rapid mutation rate seen in bacteria and viruses, through selection pressure, many changes of individual or even a number of amino acids do not destroy 3D protein structure or function and may even enhance infectivity. Protein engineers and designers have noted considerable tolerance of 3D structure to minor sequence changes.
Computational tools will be vital as we explore polypeptide space. Of course, many 'designer proteins' which might exist stably in vitro might be incapable of folding or otherwise not be viable in vivo where they have to interact with the intracellular gamut of biomolecules. Very many sequences in polypeptide space may also be impractible due to constraints such as insolubility. Most proteins presently have to be synthesised through expression of recombinant DNA in bacteria. Problems in the structure and stability of their nucleic acid templates. or inherant toxicity may also block the production of certain protein sequences. But amongst these 'designer polypeptides' will be molecules that evolution has not had time to express, but have valuable properties for both research and commerce. Some polypeptides are highly toxic to humans. Ricin is one of many examples. The DNA synthesis industry will have to monitor requests for certain RNA or DNA sequences which might have lethally toxic consequences. Commercial suppliers are already taking this threat into account.
From genomic DNA sequences, without the need to isolate proteins, we can infer the amino acid sequences of every life form from microbes to man. The human genome encodes the sequences of around 20,000 proteins, other life forms express even more. Ambitious initiatives to sequence the genomes of all lifeforms on earth are in progress.
Knowing the 3D structures of proteins is critical for understanding function and development of drugs. In just the last few years, we have learnt how to exploit machine learning (ML) to predict the folding of proteins and protein assemblies. We still presently need to confirm the 3D structures of proteins significant for drug development (where knowing precise structure at an atomic level is critical) by X-ray crystallography or cryo-electron microscopy. This has happened (for example) with structure determination of the 'spike protein' of SARS-Coronavirus-19 as part of the global effort to develop therapeutics for COVID.
Changes in amino acids at one or many more points in a polypeptide chain often doesn't alter the overall folding of a polypeptide. We've see many changes through genome mutations in the amino acid sequence of the coronavirus spike protein without loss of structure or function (infectivity). But some single amino acid changes may destroy function without altering the overal 3D folding. Mutation of a residue in an alpha helix to proline can disrupt this secondary structure. Loss of a single 'catalytic' amino acid (often serine or histidine) in proteolytic enzymes like trypsin renders the molecule unable to cut polypeptide chains.
Developments in Machine Learning have opened access to de novo polypeptide &binder design
Designing proteins and peptides for medicine
The earliest attempts at polypeptide design were targetted (unsurprisingly) at developing therapeutics. Right now this has become a huge and diverse endeavour by academic establishments, public sector research organisations, biotech and pharmaceutical industries. Some of the areas of pathology being addressed are listed here:
Cancer
Autoimmunity
Infection and immunity
Antibiotics and microbial resistance
Rare conditions
Dementias
Neurodegenerative disease
Aging
Many (but not all) of these conditions arise from protein dysfunction, inappropriate protein expression or exploitation of normal cell surface proteins for attachment or entry leading to infection by microorganisms and viruses. Certain pathologies previously considered 'undruggable' quot; are now subject to new approaches which include silencing or knockdown of expression of cell proteins through gene editing in vivo. BUT developments are now arriving so quickly would be difficult to provide anything apart from a brief summary here. It's immediately evident from the above list that manipulation of the proteins in the human immune system is one of the most fertile fields for protein and peptide design and the search for inhibitory drugs. The modular form of immunoglobins, especially IgG, has lent itself to the exporation and modification of antibodies through protein engineering.
De novo protein and protein binder design
RFdiffusion (RoseTTAFold diffusion) has been developed alongside breakthrough generative machine learning methods like AlphaFold2 for designing proteins and protein binders. These methods have been fine tuned by several groups to design antibodies or 'minibodies' (the antigen combining domains of antibodies). By fine tuning, this means that the programmes were trained on antibody structures.A recent publication form David Bakers Lab reports on success with creating specific antibody molecules using RF diffusion targeting influenza haemagglutinin and Clostridium difficile toxin B..
David Baker (2024 Nobel Laureate for Protein design), runs The Institute for Protein Design at the University of Washington, USA. The institute applies computational approaches to the development of proteins targetted towards diseases related to dysfunctional proteins which are currently difficult or impractical to cure: termed 'undruggable'.
In protein design, 'diffusion' refers to a class of generative AI models that simulate the gradual transformation of random noise into structured protein shapes or sequences, enabling the creation or prediction of novel proteins and binders.
What Are Diffusion Models?
Diffusion models originated in image generation (e.g., DALL-E, Stable Diffusion). They work by starting with random noise, like static on a TV screen. Learning to reverse the noise process, gradually refining it into a meaningful output such as an image, or in this case, a protein structure. In protein science, this concept is adapted to generate or predict 3D protein structures or binding interfaces by treating atomic coordinates or sequence-structure pairs as data that can be 'denoised' into biologically plausible forms.
Diffusion in Protein Structure Prediction
Models like GeoFlow-V2 and RFdiffusion apply diffusion to atomic coordinates. They begin with noisy or incomplete structural data. Through iterative refinement, they predict realistic protein folds, even for sequences with no known structure. This is especially powerful for de novo design, where no template exists.
Diffusion in Protein Binder Design
The PPDiff model jointly generates sequence and structure of a protein that can bind to a target. It explores the vast space of possible binders by diffusing from random or masked inputs toward high-affinity, structurally compatible designs. It bypasses traditional trial-and-error methods, offering on-demand binder generation. Diffusion models are prized for:
Flexibility: hey can handle masked inputs, partial structures, or arbitrary targets.
Generative capacity: They do not just predict - they invent.
Integration: They unify structure prediction and design in a single framework sculpting proteins from noise.
A highly schematic picture of the diffusion model for design of a protein structure (itself generated by an AI image generative model!)
To design amino acid sequences that will support a target 3D fold, the ProteinMpNN method has been developed. Together these deep learning techniques are being applied to protein and binder design projects with considerable success. In the polypeptide space section of this article, these methods are used with a minimal 'archaeic' amino acid set to map designed amino acid sequence onto chosen polypeptide folds.
Searching biological databases
The latest protein structure prediction, modelling and design technologies rely on the vast databases of molecular biological data as input to AI neural networks (Machine Learning). This approach gave rise to AlphaFold protein structure prediction software which learnt from the databases of experimentally determined protein structure. AI-assisted searching of genome sequence, protein sequence and taxonomic data has led to the development of a search engine: 'Google for DNA' brings order to biology's big data & MetaGraph: Petabase-Scale Search for Genomics & Beyond Making sense of the global, dispersed and variously presented biological data to apply to the discovery and development of potential therapies will require powerful approaches of this kind to bring molecular and macroscopic knowledge together. But a note of caution is required here. As AI generated knowledge is added, biases could easily to be introduced into the biological databases used as the teaching resources for that same AI.
Tools for searching the biological databases are proliferating rapidly and it may be hard for researchers to focus on the best (for them) technologies. As an example of a recently released Natural Language search model I include: BioModelBench which aims to support intelligent searches for computational biology models, tools and biodata sets.
Developing novel molecular interactions: protein co-folding and binder discovery and design
New therapeutics often require the development of small or large molecular binders (a.k.a drugs) which bind to protein targets with nanomole to picomole affinity. This of course has been the basis of the pharmaceutical industry's endeavours from the start. Drug screening in in vitro and in vivo assay has long been the gold standard for such drug development. We have entered an era where the binders can themselves be peptides or polypeptides. Computational approaches employing known 3D structures had formerly relied on molecular dynamics (MD) to fit binders to proteins. Now we have generative protein and peptide design. This can be applied to design protein or peptide binders for a very diverse range of targets including 'minibodies' , peptides and small organic drugs. This is very reliant upon the data from which designer AI can learn. But ligands interact with protein targets in many ways. There are numerous examples of experimentally determined protein structures in the PDB with bound ligands at 'active' (orthosteric) sites, but few examples available of protein structures with ligands bound at 'inhibitor' (allosteric) sites. Researchers often comment that there are vastly more solved protein structures with ligands bound at both orthosteric and allosteric sites in the possesion of 'Big Pharma', but these are presntly not public domain, relating as these do to the core businesses of the pharmaceutical industry.
There are now a large number of reports of designed polypeptides, describing their structures and functions. In many instances, designed polypeptides have been transferred from the dry lab (computers) to the wet lab (expression of recombinant DNA in micro-organisms, isolation of the designed molecules and testing for structure and function). Regular competitions are now run by Nipah and Adaptyv in a manner analogous to the CASP competition for protein structure prediction See: Protein base :The home of protein design data
Just one example illustrated here is taken from the hundreds available in Protein Base is a 26KDa designed binder for the Epidermal Growth Factor Receptor (EGFR). Submitted to the Adaptyv Bio EGFR Binder design competition, it has been expressed and has a strong EGFR binding affinity (KD) of 4.3 x 10-10. Note the extensive use of beta sheet which gives the designed structure considerable stability.
Peptide framework design
Design: Porous peptide framework design. Credit: Ganatra, P. et al.
Much of the future of protein and peptide design lies in building interacting peptides and proteins. It's notable that evolution has produced many polypeptide structures which interact to build nanomachines or to provide containment (I refer here to the vault proteins in the introduction to this protein resource) and delivery vehicles like the viral protein shells which carry their genomes).
Researchers are now successfully designing protein and peptide frameworks which could be used for purposes like drug delivery. In the figure α-helical peptides conjugated to π-stacking residues that readily crystallise into frameworks through a combination of π-π stacking and coiled-coil (helix-helix) interactions. Ref: Ganatra et al.
Searching small molecule space
Known chemical compounds (to say nothing of the unknowns or unsynthesised molecules) is another vast molecular space. The recent surge of commercially available synthetic chemicals in that space provides the opportunity to search for ligands of therapeutic targets among billions of compounds. In the reference given below, emphasis is placed on strategies to explore ultra-large chemical libraries and synergies with emerging machine learning techniques. Screening is becoming virtual to reduce the time and expense of the essential final wet lab screens. Structure-based virtual screening of vast chemical space as a starting point for drug discovery
Extremely broad molecular docking of designer ligands with experimentally determined or AlphaFold predicted protein targets. Credit: Carlsson J. et al. Current Opinion in Structural Biology
How many unique polypeptides might there be in a synthetic (man made) polypeptide space?
Even in the more than 3 billion years of life on earth, it seems unlikely that there's been time for evolution to access more than a minute fraction of potential polypeptide space. This raises the question: what applications might we make given the rapidly increasing number of ways to synthesise polypeptides now available?
Proteins already exist with highly useful thermodynamic, enzymatic and therapeutic properties - the DNA polymerase enzymes used in Polymerase Chain Reaction (PCR) to give just one example. Birds appear to have sensory proteins (cryptochromes) in their eyes which can detect earth's magnetic field to use as aids to navigation. Lifeforms exist which have proteins which generate light or electric currents. But can we use novel proteins and peptides in medicine? Apart from the challenge of immune response to novel entities introduced into our bodies, there is the problem of stability: polypeptides are subject to metabolism including proteolytic attack and rapid clearance by filtration in the kidneys. N-acetylation or cyclisation has been used to minimise peptide ligand loss. What if we built some or all of potential binders with D-amino acids, which would be resistant to proteolytic attack and have very different pharmacokinetics?
The inventory of polypeptide structures and functions is already very long, but may represent a drop in the proteinaceous ocean so far. So how might we enter polypeptide space ourselves and 'boldly go where nature has not been before'?
For general polypeptide design, we aren't restricted to evolution's choice of amino acids!
Going back in deep time - when there were only the ancestors of today's prokaryote micro-organisms, we don't really know when the familiar 20 amino acid building blocks of proteins came on the scene. We don't even know if their precursors were the same: going back that far in time, no chemical evidence remains. With a degree of inference, it has been assumed that a small number of amino acids were available during the early development of life. In an intriguing study, a research team asked if proteins designed with a very limited amino acid 'alphabet' 'Ancient folds' could yeild stable structures and possess functions. They selected a limited 'alphabet' of 10 amino acids, rimarily small, polar, and aliphatic residues such as the canonical Gly, Ala, Asp, Glu, Val, Ser, Ile, Leu, Pro, and Thr. So important residues in modern proteins like aromatic and basic side chains were absent. The minimal amino acid set was encoded in DNA in bacterial expression systems. The authors used the ProteinMpNN and RFdiffusion methods described in the section on protein design to generate sets of amino acid sequences that would fit into template 3D folds. Using a variety of structure assessing methods, the reduced alphabet proteins were found to have stable folds, adopting the designed fold and demonstrating good chemical and thermal stability. A metal binding function was created. Giacobelli et al. The minimal amino acid set would not serve all the functions of modern proteins, but does suggest that an evolutionary path to today's proteins was perfectly possible.
A recent article described proteins that assemble and disassemble on command, using an incorporated light-sensitive amino acid phenylalanine-4′-azobenzene 𝗔𝘇𝗼𝗙. They describe a computational approach to design protein–protein interactions regulated by non-canonical amino acids, focusing on the light-responsive phenylalanine-4′-azobenzene (AzoF). Using this approach, light-responsive cyclic homo-oligomers and heterodimers, which only assemble in AzoF’s trans configuration and disassemble when AzoF photoisomerizes to the cis configuration. Biophysical characterization confirms the light-responsive assembly and disassembly of these complexes, and the crystal structures match the design models with atomic accuracy. These light-responsive proteins can be used in constructing light-responsive hydrogels and engineering synthetic ligand receptors to optocontrol cell signalling in mammalian cells.
If we fully venture into synthetic polypeptide space, around 500 diferent amino acids are known to date. Considering these as potential monomeric units, the number of possible polypeptide molecules would be essentially, infinite. The huge challenge would be to identify small subsets of polypeptides which would teach us about biological function or have meaningful applications. I have no doubt that to access such vast molecular spaces, machine Learning would be essential here.
We might intuitively visualise a vast type of library (described in the literature as the 'Library of Babel') where the books could carry a different polypeptide sequence on every page, differing by one residue from the sequence on the neighbouring page. In this imaginary library, every known natural polypeptide would be present in books on the shelves. Such a library would also carry a vast number of polypeptide sequences which have valuable functions, not present in nature, or still remain undiscovered. But without a 'map' it could take many lifetimes (and enormous funds!) to locate them.
To represent this vast collection of molecules as a space that might be explored and manipulated we have to turn to maths and extremely fast computers. A mathematically interrogatable polypeptide space might be multidimensional with a dimension for at least each of the 20 common amino acids and chain length. First imagine all natural protein sequences (known by experiment or deduced from genome sequences) of up to, say 1000 residues) positioned as points in such a multidimensional (tensor) space. It would be expected that large regions of this space will be unoccupied or very sparcely populated. Homologous amino acid sequences differing in one or a few residues would be closely grouped. Fibrous proteins with highly repetitive sequences (collagen, spiderin) would appear in regions of polypeptide space remote from the bulk of globular proteins. Others might either be closely grouped (leaving vast regions of polypeptide space empty) or randomly distrubuted throughout. Just visualising existing proteins mapped into this space could be revealing about the rules and options of nature to date!
Envoi: Why even contemplate exploring polypeptide space?
Scientists are basically curious. English poet Keats deprecated the splitting of white light into the visible colours that scientist Isaac Newton published in his masterwork: Optiks. Keats rhymed that Newton had 'unwoven the rainbow' and had 'reduced it to a prism' . Keats preferred the notion that some mysteries should remain. His view was that by explaining nature, humans would lose their sense of awe, respect and well, wonder. But our drive to discover how life works will surely enhance, not diminish our sense of wonder at what evolution has acheived (without our help!). And unlike the legend of treasure at the inaccessible foot of the rainbow, tangible and accessible 'crocks of gold' will surely be awaiting us in our future adventures in polypeptide space.
Even the Rainbow isn't large enough to hold the DNA sequences needed to specify all possible proteins. The missing ones are presently 'Somewhere over the Rainbow'!
References
I have provided hyperlinks within this article to selected papers, reviews and resources about proteins. The ones I have included are mostly open source. This is such a vast field it's impracticable to even partly cover the main aspects of it. Here I have (immodestly) given references to my own published research on the proteins I worked with that I have used as examples in this review.