The latest example of a groundless socially constructed triumphal legend spread by materialists is the claim that some software called the AlphaFold software (existing in forms such as AlphaFold, AlphaFold2 and AlphaFold3) solved the long-standing problem of biology called the protein folding problem.
The human body contains 20,000+ protein molecules, and most protein molecules have a sequence of hundreds or thousands of amino acids specified by some particular gene. But a protein molecule is not a mere chain of amino acids. Protein molecules have folded three-dimensional shapes necessary for their functions.
Simplifying things, you can think of it this way:
Gene --> Polypeptide sequence (amino acid chain) --> Protein
How is it that a protein molecule gets the three-dimensional folded shape needed for its function? Does it read instructions on how to make such a shape from a gene? No, it does not. Nowhere in DNA or its genes is there any such thing as a specification for how to make the three-dimensional shape of a protein molecule.
Ignoring possible cases of shape duplication, you can say that within the human body there are very roughly 20,000+ different protein molecule structures assumed by very roughly 20,000+ different types of protein molecules. The problem of how these folded shapes arise is the unsolved problem called the protein folding problem.
We should avoid getting confused between the protein folding problem and the protein folding prediction problem, which are two separate problems. The protein folding prediction problem is the problem of predicting the three-dimensional shape of a protein molecule using that protein's amino acid sequence. The protein folding problem is the very different problem of why proteins assume the three-dimensional shapes that they have. Some progress has been made on the protein folding prediction problem by some AlphaFold software using machine-learning and massive databases storing information on genes and the shapes that proteins have. But that progress is merely progress on the protein folding prediction problem, not progress in solving the protein folding problem. The protein folding problem is still unsolved. Scientists do not understand how proteins are able to form into the three-dimensional shapes that they have, the shapes needed for their function. A scientific paper notes that "AlphaFold2 (AF2) revolutionized protein structure prediction, yet it is often conflated with the protein folding problem," thereby noting the confusion that is occurring between a protein structure prediction problem and a separate and distinct protein folding problem.
Below are some quotes by scientists confessing that the protein folding problem has not been solved.
- "In real time how the chaperones fold the newly synthesized polypeptide sequences into a particular three-dimensional shape within a fraction of second is still a mystery for biologists as well as mathematicians." -- Arun Upadhyay, "Structure of proteins: Evolution with unsolved mysteries," 2019.
- "The problem of protein folding is one of the most important problems of molecular biology. A central problem (the so called Levinthal's paradox) is that the protein is first synthesized as a linear molecule that must reach its native conformation in a short time (on the order of seconds or less). The protein can only perform its functions in this (often single) conformation. The problem, however, is that the number of possible conformational states is exponentially large for a long protein molecule. Despite almost 30 years of attempts to resolve this paradox, a solution has not yet been found." -- Two scientists, "On a generalized Levinthal's paradox," 2018.
- "How proteins fold remains a central unsolved problem in biology. While the idea of a folding code embedded in the amino acid sequence was introduced more than 6 decades ago, this code remains undefined. While we now have powerful predictive tools to predict the final native structure of proteins, we still lack a predictive framework for how [amino acid] sequences dictate folding pathways....Almost seven decades of experimental and theoretical inquiry have not revealed a 'folding code' at the amino acid level, i.e., rules endowed with the generality and predictive power required to connect amino acid sequence to how the protein attains its structure....Machine learning made it possible to identify weak correlations to generate the structure most likely to correspond to a sequence. This tour-de-force effort has largely solved the problem of predicting protein structure from sequence...but with a key limitation: the algorithm that predicts the structure is a complex black box of pattern recognition that casts little light on the process of folding and that tells us nothing about why only some sequences fold, or how physics and evolution are coupled." -- Five scientists in the year 2025 (link).
- "The real challenge—that remains unanswered after more than 50 years of research in the structural biology field—is understanding the mechanisms that lead proteins to fold into their native state. The reason for these difficulties is that the central question of the protein folding problem remains unresolved: specifically, how a sequence of amino acids encodes its folding pathways." -- Scientist Jorge A. Vila, 2025 (link).
- "One of the most puzzling and unsolved challenges in molecular biology is understanding how proteins fold. " -- Scientist Jorge A. Vila, 2026 (link).
- "The origin of functional proteins remains a fundamental biological enigma...The physical principles governing protein genesis itself, from prebiotic condensation to functional protein emergence, remain unresolved." -- Nine scientists, 2026 (link).
A year 2026 paper makes it clear that contrary to boasts in the press, the AlphaFold2 software does not actually solve the protein folding problem, the problem of how protein molecules almost instantly acquire very complicated 3D shapes needed for their function. The year 2026 paper states, "The explanatory scientific understanding of the protein folding problem is thus not directly advanced by AF2 [AlphaFold2]." Later the same paper says, "The protein folding problem remains unsolved."
A paper published in the year 2026 throws some cold water on triumphal boasts about the AlphaFold software, while reiterating that the protein folding prediction problem (a "what" problem) is very different from the protein folding problem (the "how" problem of how proteins fold into the shapes needed for the biological function):
"AlphaFold only works some of the time.... However, true single sequence structure prediction has remained elusive. The cautious old guard who initially responded to AlphaFold by highlighting the difference between protein structure prediction (the what) and protein folding (the how), correctly saying that the second remains an open problem, has surprised nobody by not going on to work on protein folding."
How well does the AlphaFold series of software predict the structure of a protein, given its amino acid sequence? According to one recent paper, the "official" answer to this question is to be found in the 2026 paper "CASP16 Protein Monomer Structure Prediction Assessment" which you can read here. Below I'll call this the CASP16 wrap-up paper. Over many years, there have been a series of CASP competitions, in which competitors using different types of software have tried to predict the 3D shape of a protein molecule, given an input of its amino acid sequence. The latest of these competitions which has published results was the CASP16 competition held in 2024, and its results are described in the paper above. (A CASP17 competition has just recently completed, but it will be quite a few months before we have a scientific paper publishing its results.)
In a previous post I did discussing the CASP14 competition, I noted that there are two different ways of calculating the accuracy of a protein shape prediction: a GDT_TS measure and a more stringent GDT_HA measure. Using the more stringent GDT_HA measure, Figure 2B of the CASP16 wrap-up paper gives us the graph below showing prediction accuracy of the latest versions of the AlphaFold software (AlphaFold2 and AlphaFold3). AF2 refers to AlphaFold2. AF3 refers to AlphaFold3.
Each little dot represents a particular prediction (or maybe a set of predictions). The vertical position of the dot represents how accurate the prediction was. Dots near the the top of the graph (near the 100 mark) are very good predictions. Dots in the blue areas of the graph are not very good predictions, predictions only about 70% correct. Dots below the blue areas of the graph are poor predictions.
Anyone who has read the hype about the AlphaFold series of software may be surprised by the result above. The graph seems to show an average prediction accuracy of only about 70%, with the accuracy ranging from only about 40% to as high as almost 98%. But didn't we read again and again science news articles making it sound like the AlphaFold series of software had mastered the problem of predicting the structure of proteins from their amino acid sequences? Such articles were very misleading.
We read this: "Overall, [predictive] performance on monomer targets in CASP16 showed minimal improvement compared to CASP15." The reference is to the CASP16 competition held in 2024 and the CASP15 competition held in 2022. The statement I just quoted contradicts the impression that we have got from the press of skyrocketing progress in this area.
A key issue in evaluating the predictive effectiveness of the AlphaFold series of software is the issue of target size. The target size refers to the amino acid length of a protein which AlphaFold software attempted to describe using its structure prediction methods. The wrap-up paper mentioned above does a bad job of describing these target sizes. To find the target sizes, you must go to the page here (the page for the CASP16 competition) and then click on the "Target List" link, which takes you to the page here.
The page isn't well-designed in terms of presenting information in a way that the average reader will understand. Under the heading of "Multimers" we have a cryptic column heading of "Res" which gives numbers. Those numbers are the "residue length" of the proteins that were used as prediction targets, and these "residue lengths" were the number of amino acids in the proteins. If you click on the "Res" column, the column will sort by the lengths of the amino acid sequences in the proteins. When you see a number above about 450, you are looking at more difficult prediction targets. When you see a number above 1000, you are looking at the most difficult prediction targets. When you see a number below 400, you are looking at the easiest prediction targets.
Scrolling down the page, I see that the selected prediction targets were mostly not very challenging. Very many types of human proteins have more than 1000 amino acids, and more than 500 types of human proteins have more than 3000 amino acids each. But out of 150 prediction targets of the CASPR16 competition, only 29 had amino acid lengths greater than 1000. About 35 of the 150 or so prediction targets had an amino acid length between 1000 and 600. 38 of the prediction targets had an amino acid length between 600 and 450. About 49 of the 150 or so prediction targets had an amino acid length less than 450, which is less than the average amino acid length of a human protein.
So given this not-very-challenging set of prediction targets that has less-complex-than-average types of proteins about as frequently as more-complex-than-average types of proteins, we should not be too impressed by Figure 2B showing a prediction accuracy of about 70%. It would seem that if the prediction targets had been proteins with above-average complexity as often as targets with average complexity, that the reported prediction accuracy of about 70% would have been something much smaller, such as maybe 60% or 50%.
From the graph above, it is clear that the AlphaFold family of software does not very well predict the structure of protein molecules, using the amino acid sequence as an input. Its performance accuracy should not be described as "very good" but merely as perhaps "fair-to-somewhat-good."
There is a site called the AlphaFold Protein Structure Database which you can reach here. The average person will be puzzled by how you can use this site to check the quality of predictions by the AlphaFold software. I can describe one way. You can get a list a protein names by going to the UnitProt database site here, and typing in a query like this:
This will give you a result of thousands of rows, showing human proteins that have more than 1000 amino acids. You can then type in some of the data from that result set into the search box of the page for the AlphaFold Protein Structure Database which you can reach here.
Doing that, I got some results from that page, and the results were not very impressive. For example, I typed in the phrase "Transcription factor TFIIIB component B'' homolog" using a phrase I had got from the UniProt result set using the query above. I got these results from a few queries using the AlphaFold Protein Structure Database:
Protein Name | Gene | UniProt ID | Amino Acid Length | Global Quality (How Well Alpha Fold Predicts the Struc- ture) |
BDP1 | A6H8Y1-4 | 1372 | 46.75 (Very Low) | |
ARID1A | O14497-2 | 2068 | 48.75 (Very low) | |
CEP290 | O15078 | 2479 | 60.53 (low) | |
ZNF292 | O14497-2 | 2578 | 46.91 (Very Low) | |
PCARE | A6NGG8 | 1288 | 43.78 (Very Low) | |
TAF4 | O00268 | 1085 | 52.78 (Low) |
I did not have to do much work to find these examples of low-quality predictions by the AlphaFold software. I had to only search through a random list of about 15 or 20 proteins with an amino acid length greater than 1000. When I searched for proteins with an amino acid length between 500 and 700 (only somewhat more complex than an average human protein molecule), I quickly found some examples that the AlphaFold software performed poorly on when analyzing. For example, its prediction about the TANK-binding kinase 1-binding protein 1 of 611 amino acids was not very good, being rated as only "63.62 (Low)." And the AlphaFold software's prediction about a Nucleolar protein 4 of 536 amino acids was not very good, rated as only "60.91 (Low)."
I may note that the AlphaFold Protein Structure Database site which you can reach here uses quality-rating adjectives that are way too positive-sounding. The database routinely describes predictions that are only about 70% accurate as being "high" in quality. In few other fields would ratings be so charitable. For example, if a car assembly team assembled only 70% of the car's parts correctly, it would make no sense to claim its assembly skill was "high." And if someone predicting the outcome of mixing particular quantities of particular chemicals had an accuracy rating of only 70%, people would not rate his accuracy as "high," but complain that he was poor at predictions. You would hardly claim that your physician was "high" in proficiency if he prescribed the right medicine only 70% of the time. In field such as physics, any theory of gravitation would be subject to the most scornful derision if it made predictions that were only 90% accurate, and its proponents claimed that this was "high" accuracy. In fields such as physics, a rating of "high" quality tends to go to only predictions that are at least 99% accurate.
An example of the grotesque misinformation being stated in the popular press about the AlphaFold software is this claim in an article:
"DeepMind started developing AlphaFold in 2018. In 2020, it was recognized as a solution to humanity's 50-year-old 'protein folding problem,' which sought to answer how amino acids automatically fold into complex 3D shapes. "
The AlphaFold software did nothing to solve the "50-year-old 'protein folding problem,' which sought to answer how amino acids automatically fold into complex 3D shapes." All that the AlphaFold software did was to make some progress on a different problem, the protein folding prediction problem. And the progress that the AlphaFold software made was very limited. The graph above with the blue areas shows that it is still impossible to accurately predict the structure of most proteins from their amino acid sequences.
It wasn't just the writer quoted above who misinformed us on this topic. It was the organization that runs the CASP competition, which issued a highly misleading press release in 2020. One of endless science press releases with unfounded boasts, the press release quoted a Professor Dame Janet Thornton as saying the following:
" One of biology s biggest mysteries is how proteins fold to create exquisitely unique three-dimensional structures. Every living thing from the smallest bacteria to plants, animals and humans is defined and powered by the proteins that help it function at the molecular level. So far, this mystery remained unsolved, and determining a single protein structure often required years of experimental effort. It s tremendous to see the triumph of human curiosity, endeavour and intelligence in solving this problem. A better understanding of protein structures and the ability to predict them using a computer means a better understanding of life, evolution and, of course, human health and disease."
Here Thornton bungled badly by confusing two different problems: the protein folding prediction problem (the problem of predicting the structure of proteins from their amino acid sequence) and the protein folding problem (the problem of how proteins are able to form their three-dimensional structures from amino acid sequences that do not specify such structures). It was a goof as bad as someone getting all mixed up and conflating and confusing the problem of how to play the card game called bridge, and the problem of how to construct a bridge. At the time this quote was made, and even so today, nothing had occurred to justify this boast about "solving this problem," as nothing had been done to solve the protein folding problem; and no software existed that very accurately and reliably predicted protein structure. And no such software exists even today.
Referring to the AlphaFold3 software using the phrase AF3, a year 2026 paper states this: "We conclude that AF3 is a poor predictor of D-peptide chirality, fold, and binding pose."
We are reminded here that materialism is a giant myth-making machine. Materialism grinds out example after example of unfounded socially constructed triumphal legends, particularly when such legends serve the ideological needs of materialists. The protein folding problem is a great thorn in the side of materialists. Somehow proteins in our bodies are constantly forming into the three-dimensional shapes needed for their biological function, just as if some purposeful intelligence of vast power was acting to achieve such effects. This reality is troubling to materialists, and we can understand why they would wish to construct an analgesic legend to ease their discomfort over this matter. But the claim that AlphaFold family of software programs solved the protein folding problem is as unfounded as the claim that Darwin explained the origin of species.
At about this time (September, 2026) the latest in the CASP series of competitions has completed, a competition called CASP17. It will be quite a few months before there is published a paper that details the predictive accuracy of competitors in this competition. In December 2026 there will be a conference announcing preliminary results of the CASP17 competition. Given the misleading content in the November, 2020 CASP press release, we should be suspicious of boasts of any press release announcing results of the CASP17 competition, and subject such a press release to critical scrutiny.
A major contributor to the social construction of unfounded triumphal legends is the Nobel Prize organization, which sometimes gives undeserved Nobel prizes, or publishes unfounded or misleading claims when awarding Nobel prizes, as I document in my post here. When it awarded a Nobel prize in chemistry to two people who had worked on protein structure prediction, the claim was modest. Officially the award is merely going "for protein structure prediction." But the announcement page links to biographies of the two scientists awarded, and in both cases the pages have remarkably misleading statements. On the pages here and here, we have this very misleading statement:
"In 2020, Demis Hassabis and John Jumper presented an AI model called AlphaFold2. With its help, they have been able to predict the structure of virtually all known proteins."
No such accurate prediction occurred, as my text above documents; and AlphaFold2's accuracy as measured in the graph above is only about 70%. Someone trying to defend this statement quoted above would sound pretty duplicitous. He might say, "We just said predict, we didn't say predict accurately."



No comments:
Post a Comment