Header 1

Our future, our universe, and other weighty topics


Saturday, September 9, 2023

Some Accidentally Unachievable Molecular Machines in Your Body

In order to help perpetuate their dogma that humans are accidents of nature owing our existence to lucky random mutations in the distant past, our biologists are very fond of using what I call shrink-speaking language: language that makes the endless towering cathedrals of biological complexity look like mere crumbs.  This involves many tricks of diminutive representation, such as:

  • Referring to gigantically organized human bodies as "bags of chemicals" or "carbon stuff" or "star stuff."
  • Referring to stratospheric leaps of biological organization and  functional complexity as mere "variants." 
  • Referring to enormous new bonanzas of unprecedented biological engineering (such as the Cambrian Explosion) as mere "diversification." 
Our biological professors have in effect received their marching orders that they are supposed to describe human bodies as something that nature might have accidentally produced. But there are some cases where they encounter so much fine-tuned organized and purposeful functional complexity that professors find it almost impossible to follow such orders. One such case is the case of extremely organized protein complexes that even our reductionist professors have repeatedly described as molecular machinery

Organisms such as ourselves involve hierarchically structured and enormously organized complexity that cannot be credibly explained by appealing to random mutations. What we have in a human body is enormously organized and fine-tuned complexity so immense that it can be called an enormous engineering effect. In his interesting book Cosmological Koans, the physicist Anthony Aquirre tells us about just how complex biological life is. He states the following on page 338:

"On the physical level, biological creatures are so much more complex in a functional way than current artifacts of our technology that there's almost no comparison. The most elaborate and sophisticated human-designed machines, while quite impressive, are utter child's play compared with the workings of a cell: a cell contains on the order of 100 trillion atoms, and probably billions of quite complex molecules working with amazing precision. The most complex engineered machines -- modern jet aircraft, for example -- have several million parts. Thus, perhaps all the jetliners in the world (without people in them, of course) could compete in functional complexity with a lowly bacterium."

So if a lowly bacterium has a functional complexity comparable to a jetliner, what kind of functional complexity does a human body have? Functional complexity so great it can be called an enormously strong engineering effect. The human body includes many types of  molecular machines that are best classified as accidentally unachievable. Something is accidentally unachievable if there are no imaginable unguided accidental events that could cause its origin. 

I can explain the idea of something being accidentally unachievable with a simple example that is easy to understand.  A bridge across a shallow stream is something that is accidentally achievable. We can imagine some accidental arrangement of rocks that might make a kind of bridge across a shallow stream, and we can imagine some lightning bolt causing a tree to fall, making a bridge across a shallow steam. But a bridge across a very wide and deep river is something that is accidentally achievable. There is no conceivable series of accidents or unwilled natural events that could create a bridge over an average section of a wide river such as the Mississippi, which has an average width of one mile. 

Let us look at some of the most impressive cases of molecular machinery in your body, systems requiring so many thousands of well-arranged parts that they are reasonably called accidentally unachievable. Pay attention to the numbers given in the second columns of the tables below, numbers which tell how many amino acids parts have to be well-arranged to get the protein mentioned in each row. If the numbers were very low numbers such as 3 or 5 or 7, it might be excessive to state that the relevant molecular machines are accidentally unachievable. Instead, the actual numbers will be numbers in the hundreds or thousands, meaning that we will be constantly finding that individual protein components of these molecular machines each required hundreds or thousands of well-arranged amino acid parts (with a consequence that the total system required thousands of well-arranged amino acids, equivalent to tens of thousands of well-arranged atoms).  

Example #1: The Apoptosome 

apoptosme

(Image credit:  Wikipedia Commons, derived from Yuan et al. 2010, Structure of an apoptosome-procaspase-9 CARD complex)

Shown above is the apoptosome protein complex involved in programmed cell death. Note the references in the chart to propellers, which remind  us how much the complex resembles a product of engineering. Humans have more than 20,000 types of protein molecules, and the average protein molecule is a very special arrangement of more than 400 different amino acid parts. The arrangement of amino acids in each protein is as hard-to-achieve by chance as 400 accidentally typed characters making a paragraph of grammatical and functional prose. Extremely complex engineering arises in the form of protein complexes, in which different proteins (often useless by themselves) work together as team members to achieve some dramatic functional result. We see that in the visual above, where multiple instances of several different types of protein molecules come together to form an extremely complex structure consisting of thousands of well-arranged amino acid parts, and consisting of a total of tens of thousands of well-arranged atoms. A page describes the action of these individually useless proteins coming together to form a functional protein complex:

"The process of programmed cell death, also known as apoptosis, is highly regulated, and the decision to die is made through the coordinated action of many molecules. The apoptosome plays the role of gatekeeper in one of the major processes, termed the intrinsic pathway. It lies between the molecules that sense a problem and the molecules that disassemble the cell once the choice is made. Normally, the many subunits of the apoptosome are separated and inactive, circulating harmlessly through the cell. When trouble occurs, they assemble into a star-shaped complex, which activates protein-cutting caspases that get apoptosis started."

Another site that includes a 3D rotating animation of the structure shown above says this:

"The apoptosome is revealed as a wheel-like complex with seven spokes. On top of the wheel is a spiral-shaped disk that allows for docking and subsequent activation of proteases, which then target cellular components. When active, the apoptosome is revealed to be a dynamic machine with three to five protease molecules tethered to the wheel at any given time."

The "Apaf 1" part of this complex (APAF_HUMAN ) involves 1248 amino acids.   

Example # 2: The Spliceosome

spliceosome

At the site here, we read this about the human spliceosome: 

"The spliceosome is a complicated and formidable example of a multi-subunit molecular machine, with the pre-catalytic form being the largest spliceosomal complex, containing 5 RNA molecules and 65 proteins, in addition to a substrate mRNA precursor. The arrangement and activities of all of these has to be intricately coordinated, paradoxically to catalyse a rather simple chemical reaction."

The paper here describes the spliceosome as a highly dynamic machine, like some race car that has its parts changed or replaced by a pit crew as the race car stops for pit stops:

"Indeed, ∼45 proteins are recruited to the human spliceosome as part of the spliceosomal snRNPs, whereas non-snRNP proteins comprise the remainder. The composition of the spliceosome is highly dynamic with a remarkable exchange of proteins from one stage of splicing to the next. These changes are also accompanied by extensive remodeling of the snRNPs within the spliceosome."

Below are the number of amino acids involved in these parts, which I looked up using the UniProt online database (you can use the links to check the numbers I have given):

Protein

Number of amino acids

Conment

SF3A1

793

On Chromosome 22

SF3A2

462

On Chromosome 19

SF3A3

501

On Chromosome 1

PRPF3

683

On Chromosome 1

PRPF4

522

On Chromosome 9

PRP8

2335

On Chromosome 17

U5S1

972

On Chromosome 17

PRP31

499

On Chromosome 19

PRP4

522

On Chromosme 9

SNUT1

800

On Chromosome 11

SLU7

586

On Chromosome 5

RBM22

420

On Chromosome 5

FAM32A

112

On Chromosome 19

XAB2

855

On Chromosome 19

CACTIN

758

On Chromosome 19

PLRG1

514

On Chromosome 4

Altogether the structure shown above requires more than 9000 amino acids that have to be arranged in just the right way. The structure shown above is not specified in DNA, which merely specifies which amino acids make up each of the protein parts. The amino acid information needed to make the structure above (insufficient to specify the physical arrangement of the structure) is not at all contiguous in DNA. To assemble the structure above, among other wonders of construction a human body must magically gather genetic information scattered across many different chromosomes in the nucleus, like someone quickly finding just the right 45 loose pages hidden in random books of 46 tall, long bookcases in a public library. The table above shows that at least eight of the 23 human chromosome pairs would need to be accessed: Chromosome 1, Chromosome 4, Chromosome 5, Chromosome 9, Chromosome 11, Chromosome 17, Chromosome 19 and Chromosome 22. 

The CORUM database page here gives details on 39 different proteins involved in just one part of the spliceosome, a part called the spliceosome C complex. The CORUM database page here gives details on 38 different proteins involved in another part of the spliceosome, a part called the spliceosome pre-B complex. The CORUM database page here gives details on 43 different proteins involved in another part of the spliceosome, a part called the spliceosome B complex. The CORUM database page here gives details on 113 different proteins involved in another part of the spliceosome, a part called the spliceosome A complex. The CORUM database page here gives details on 139 different proteins involved in another part of the spliceosome, a part called the spliceosome E complex. The paper here says, "The spliceosome is composed of as many as 300 distinct proteins and five RNAs, making it among the most complex macromolecular machines known."

Example # 3: RNA Polymerase II

The page here discusses RNA polymerase:

"RNA is a versatile molecule. In its most familiar role, RNA acts as an intermediary, carrying genetic information from the DNA to the machinery of protein synthesis. RNA also plays more active roles, performing many of the catalytic and recognition functions normally reserved for proteins. In fact, most of the RNA in cells is found in ribosomes--our protein-synthesizing machines--and the transfer RNA molecules used to add each new amino acid to growing proteins. In addition, countless small RNA molecules are involved in regulating, processing and disposing of the constant traffic of messenger RNA. The enzyme RNA polymerase carries the weighty responsibility of creating all of these different RNA molecules...RNA polymerase is a huge factory with many moving parts".

There are three different versions of RNA Polymerase, RNA Polymerase I, RNA Polymerase II, and RNA Polymerase III. In a previous post I discussed the proteins that make up the RNA Polymerase III complex: more than a dozen proteins with amino acid sequences specified across ten or more different chromosomes, with the complex requiring thousands of well-arranged amino acids. Below are the number of amino acids involved in RNA Polymerase II, which I looked up using the UniProt online database (you can use the links to check the numbers I have given):


Protein

Number of amino acids

Comment

RPB1_HUMAN

1970

On Chromosome 17

RPB4_HUMAN

142

On Chromosome 2

RPB7_HUMAN

172

On Chromosome 11

RPB2_HUMAN

1174

On Chromosome 4

MED21_HUMAN

144

On Chromosome 12

RPB9_HUMAN

125

On Chromosome 19

RPB3_HUMAN

275

On Chromosome 16

RPB11_HUMAN

117

On Chromosome 7

RPAP3_HUMAN

665

On Chromosome 12

RPAB1_HUMAN

210

On Chromosome 19

RPAP1_HUMAN

1393

On Chromosome 15

RPAB5_HUMAN

67

On Chromosome 11

MED20_HUMAN

212

On Chromosome 6

RPAB3_HUMAN

150

On Chromosome 3

A0A3B3IRX3_HUMAN


2222

On Chromosome 12

MD13L_HUMAN

2210

On Chromosome 12

MED12_HUMAN

2177

On Chromosome X

MED13_HUMAN

2174

On Chromosome 17

MD12L_HUMAN

2145

On Chromosome 3

Requiring many more well-arranged amino acids than RNA Polymerase III, the RNA Polymerase II protein complex clearly requires more than 10,000 amino acids that have to be arranged in just the right way. The RNA Polymerase II structure is not specified in DNA, which merely specifies which amino acids make up each of the protein parts. The amino acid information needed to make RNA Polymerase II  is not at all contiguous in DNA. To assemble RNA Polymerase II, among other wonders of construction a human body must magically gather genetic information scattered across many different chromosomes in the nucleus, like someone quickly finding just the right 15+ loose pages hidden in random books of 46 tall, long bookcases in a library. The table above shows that at least eleven of the 23 human chromosome pairs would need to be accessed: Chromosome 2, Chromosome 3, Chromosome 4, Chromosome 6, Chromosome 7,  Chromosome 11, Chromosome 12, Chromosome 15, Chromosome 17, Chromosome 19, and Chromosome X.

Example # 4: Proteasomes 

The wikipedia.org article on proteasomes tells us this:

"Proteasomes are protein complexes which degrade unneeded or damaged proteins by proteolysis, a chemical reaction that breaks peptide bonds...In structure, the proteasome is a cylindrical complex containing a 'core' of four stacked rings forming a central pore. Each ring is composed of seven individual proteins."

A paper on this topic is entitled "Gates, channels, and switches: elements of the proteasome machine." We read this:

"The proteasome has emerged as an intricate machine that has dynamic mechanisms to regulate the timing of its activity, its selection of substrates, and its processivity. The 19-subunit regulatory particle (RP) recognizes ubiquitinated proteins, removes ubiquitin, and injects the target protein into the proteolytic chamber of the core particle (CP) via a narrow channel."

Another paper is entitled "The 26S Proteasome: A Molecular Machine Designed for Controlled Proteolysis." A page on the site of the Theoretical and Computational Group tells us this:

"Recycling of unneeded protein molecules in cells is performed by a molecular machine called 26S proteasome (Figure 1), which cuts these proteins into smaller pieces for reuse as building blocks for new proteins. Proteins that need to be recycled are labeled by tags made of poly-ubiquitin protein chains. The 26S proteasome machine recognizes and binds to these tags, pulls the tagged protein close, then unwinds it, and finally cuts it into pieces. As the cell's recycling machinery, the 26S proteasome is vital for a variety of essential cellular processes, including protein quality control, cell cycle regulation, adaptive immune response, and apoptosis....The 26S proteasome recruits, unfolds, and degrades poly-ubiquitin tagged proteins through a complex interaction clockwork of over 60 known protein subunits that is driven through ATP hydrolysis."

A scientific paper tells us this:

"The 26S proteasome is a multisubunit complex that catalyzes the degradation of ubiquitinated proteins. The proteasome comprises 33 distinct subunits, all of which are essential for its function and structure.

Below is a depiction of the human 26S proteasome structure, one that labels some of its protein parts. We see three different views of the same protein complex, with different protein parts labeled (the Greek letters used stand for alpha and beta parts mentioned in the table below):

26s proteasome

Image credit: Xing Guo et. al, link.

Below are the number of amino acids involved in these parts, which I looked up using the UniProt online database (you can use the links to check the numbers I have given):

Protein

Number of amino acids

Coment

Proteasome subunit beta type-1

241

On Chromosome 6

Proteasome subunit beta type-2

201

On Chromosome 1

Proteasome subunit beta type-3

205

On Chromosome 17

Proteasome subunit beta type-4

264On Chromosome 1

Proteasome subunit beta type-5

263On Chromosome 14

Proteasome subunit beta type-6

239On Chromosome 17

Proteasome subunit beta type-7

248On Chromosome 20

Proteasome subunit alpha type-1

263On Chromosome 11

Proteasome subunit alpha type-2

234On Chromosome 7

Proteasome subunit alpha type-3

255On Chromosome 14

Proteasome subunit alpha type-4

261On Chromosome 15

Proteasome subunit alpha type-5

241On Chromosome 1

Proteasome subunit alpha type-6

243On Chromosome 14

Proteasome subunit alpha type-7

248On Chromosome 20


The structure shown above clearly requires several thousands of amino acids that have to be arranged in just the right way. The structure shown above is not specified in DNA, which merely specifies which amino acids make up each of the protein parts. The amino acid information needed to make the structure above (insufficient to specify the total structure) is not at all contiguous in DNA. To assemble the structure above, among other wonders of construction a human body must magically gather genetic information scattered across many different chromosomes in the nucleus, like someone quickly finding just the right 60 loose pages hidden in random books of 46 tall, long bookcases in a library. The table above shows that at least eight of the 23 human chromosome pairs would need to be accessed: Chromosome 1, Chromosome 6, Chromosome 7, Chromosome 11, Chromosome 14, Chromosome 15, Chromosome 17, and Chromosome 20.

Example # 5: The ATP Synthase Complex

Another example of accidentally unachievable molecular machinery in the human body is the ATP synthase protein complex. It's a very complex molecular motor system described in a paper entitled "ATP Synthase: Motoring to the Finish Line." The paper refers to this complex as a "sophisticated molecular machine."  We read this: "ATP synthase is an unusually efficient rotary motor that synthesizes ATP at rates exceeding 100 molecules per second."  Another scientific page tells us this:

"ATP synthase is one of the wonders of the molecular world. ATP synthase is an enzyme, a molecular motor, an ion pump, and another molecular motor all wrapped together in one amazing nanoscale machine. It plays an indispensable role in our cells, building most of the ATP that powers our cellular processes....Why have two motors connected together? The trick is that one motor can force the other motor to turn, and in this way, change the motor into a generator. "

ATP Synthase
ATP Synthase (image source: link)

Below are some of the components of ATP synthase, as listed in the UniProt database.

Protein

Number of amino acids

Comment

ATP5E_HUMAN

51

On Chromosome 20

ATP8_HUMAN

68From Mitochondrion

ATP6_HUMAN

226From Mitochondrion

ATP5S_HUMAN

215On Chromosome 14

ATP5J_HUMAN

108
On Chromosome 21

ATP68_HUMAN

58On Chromosome 14

ATP23_HUMAN

246On Chromosome 12

ATP51_HUMAN

69On Chromosome 4

ATP5H_HUMAN

161

On Chromosome 17

ATPPK_HUMAN

94

On Chromosome 7

ATPB_HUMAN

529

On Chromosome 12

ATPD_HUMAN

168

On Chromosome 19

ATPA_HUMAN

553

On Chromosome 18

AT5G1_HUMAN

136

On Chromosome 17

ATPG_HUMAN

298

On Chromosome 10

ATP5G2_HUMAN

141

On Chromosome 12

ATPO_HUMAN

213

On Chromosome 21

AT5G3_HUMAN

142

On Chromosome 2

AT5F1_HUMAN

256

On Chromosome 1



ATP Synthase seems to require thousands of amino acid parts arranged in just the right way, which amounts to a special arrangement of tens of thousands of atoms. The arrangement of the parts of the ATP Synthase complex is not specified in DNA, which does not specify which proteins are parts of particular protein complexes. To assemble the structure above, among other wonders of construction a human body must magically gather genetic information scattered across many different chromosomes in the nucleus, like someone quickly finding just the right 19 loose pages hidden in random books of 46 tall, long bookcases in a public library. The table above shows that at least 12  of the 23 human chromosome pairs would need to be accessed: Chromosome 1, Chromosome 2,  Chromosome 4, Chromosome 7, Chromosome 10, Chromosome 12, Chromosome 14,  Chromosome 17, Chromosome 18, Chromosome 19, Chromosome 20 and Chromosome 21.

Example # 6: The Origin Recognition Complex/Replicative Helicase Complex

Mammalian cells are so complicated they have been compared to factories or jet aircraft. The reproduction of most cells in the human body is a miracle of replication beyond the understanding of today's science. Scientists have confessed that they do not know what causes the fantastically complex process of cell reproduction. Scientists merely understand phases of such a process, and what components play a role in the process. 

One of those components is called the origin recognition complex. The wikipedia.org article on this protein complex says this: "The origin recognition complex (ORC) is a highly conserved six subunits protein complex essential for the initiation of the DNA replication in eukaryotic cells." The ORC complex works as a team with a "replicative helicase" complex consisting of the six bottom rows on the table below. So the wikipedia.org article on the ORC complex lists all of the items in the table below as the items in the complete complex. Below are some details of these subunits:

Protein

Number of amino acids

Comment

ORC1

861

On Chromosome 1

ORC2

577

On Chromosome 2

ORC3

711

On Chromosome 6

ORC4

436

On Chromosome 2

ORC5

435

On Chromosome 7

ORC6

252

On Chromosome 16

Cdc6

560

On Chromosome 17


Mcm2

904

On Chromosome 3

Mcm3

808

On Chromosome 6

Mcm4

863

On Chromosome 8

Mcm5

734

On Chromosome 22

Mcm6

821

On Chromosome 2

Mcm7

719

On Chromosome 7


Science writer Amber Dance skillfully describes the operations of the unit above:

"The average dividing cell must copy—perfectly—3.2 billion base pairs of DNA, about once every 24 hours. The cell’s replication machinery does an amazing job of this, copying genetic material at a lickety-split pace of some 50 base pairs per second. Still, that’s much too slow to duplicate the entirety of the human genome. If the cell’s copying machinery started at the tip of each of the 46 chromosomes at the same time, it would finish the longest chromosome—No. 1, at 249 million base pairs—in about two months. 'The way cells get around this, of course, is that they start replication in multiple spots,' says James Berger, a structural biologist...'But that poses its own challenge,' says Berger, 'which is, how do you know where to start, and how do you time everything?' Without precision control, some DNA might get copied twice, causing cellular pandemonium... It takes a tightly coordinated dance involving dozens of proteins for the DNA-copying machinery to start replication at the right point in the cell’s life cycle...Kicking off the process is a cluster of six proteins that sit down at the origins. Called ORC, this cluster is shaped like a double-layer ring with a handy notch that allows it to slide onto the DNA strands, Berger’s team has found...Once ORC has settled onto the DNA, it attracts a second protein complex: one that includes the helicase that will eventually unwind the DNA. Costa and colleagues used electron microscopy to work out how ORC lures in first one helicase, and then another. The helicases are also ring-shaped, and each one opens up to wrap around the double-stranded DNA. Then the two helicases close up again, facing toward each other on the DNA strands, like two beads on a string."

The molecular machinery shown above clearly requires more than five  thousand amino acids that have to be arranged in just the right way. The structure of the molecular machinery described above is not specified in DNA, which merely specifies which amino acids make up each of the protein parts. The amino acid information needed to make the structure above (insufficient to make the 3D structure) is not at all contiguous in DNA. To assemble the structure above, among other wonders of construction a human body must magically gather genetic information scattered across many different chromosomes in the nucleus, like someone quickly finding just the right 14 loose pages hidden in random books of 46 tall, long bookcases in a library. The table above shows that at least nine of the 23 human chromosome pairs would need to be accessed: Chromosome 1, Chromosome 2, Chromosome 3, Chromosome 6, Chromosome 7, Chromosome 8, Chromosome 16, Chromosome 17 and Chromosome 22.

Example # 7: The Nuclear Pore Complex

The nuclear pore complex or NPC is a large protein complex found in the "nuclear envelope" that is the outer boundary of the nucleus inside human cells.  A science research press release tells us this: "

"For structural biologists, the human NPC is a challenging yet exciting 3D puzzle, with around 30 different proteins each present in multiple copies. This amounts to around 1000 puzzle pieces, which form a round core with surrounding flexible parts."

The wikipedia.org article on this complex states that it consists of "456 individual protein molecules, and 34 distinct nucleoporin proteins." So the complex apparently requires 34 types of protein molecules. The article tells us that the "principal function of nuclear pore complexes is to facilitate selective membrane transportation of various molecules across the nuclear envelope." This mean that nuclear pore complexes have the extremely complex job of acting like gatekeepers, letting the right kind of molecules get into the nucleus of the cell, and keeping out the wrong type of molecules.  The article tells us that there are typically about 1000 of the nuclear pore complexes in every cell. We read of some impressive functionality of these nuclear pore complexes:

"Notably, the nuclear pore complex (NPC) can actively mediate up to 1000 translocations per complex per second. While smaller molecules can passively diffuse through the pores, larger molecules are often identified by specific signal sequences and are facilitated by nucleoporins to traverse the nuclear envelope."

The article tells us that a nuclear pore complex has a molecular weight of about 110 megadaltons. A dalton is the mass equal to a twelfth of the mass of a carbon atom. A protein complex of 110 megadaltons would have the mass of about 9 million carbon atoms. Apparently the proteins that make up this complex are particularly complex proteins. Below are the exact numbers (we may assume that there are multiple instances of such proteins in a nuclear pore complex). 


Protein

Number of amino acids

Comment

NUP98_HUMAN

1817

On Chromosome 11

NU153_HUMAN

1475

On Chromosome 6

NUP93_HUMAN

819

On Chromosome 16

NU107_HUMAN

925

On Chromosome 12

NU205_HUMAN

2012

On Chromosome 7

NU160_HUMAN

1436

On Chromosome 11

NU214_HUMAN

2090

On Chromosome 9

NUP85_HUMAN

656

On Chromosome 17

NUP50_HUMAN

468

On Chromosome 22

NUP88_HUMAN

741

On Chromosome 17

NU133_HUMAN

1156

On Chromosome 1

NU155_HUMAN

1391

On Chromosome 5


The molecular machinery shown above clearly requires more than 12,000 amino acids that have to be arranged in just the right way, which amounts to a special arrangement of more than 100,000  atoms. The structure of the molecular machinery described above is not specified in DNA, which merely specifies which amino acids make up each of the protein parts. The amino acid information needed to make the structure above  is not at all contiguous in DNA. To assemble the structure above, among other wonders of construction a human body must magically gather genetic information scattered across many different chromosomes in the nucleus, like someone quickly finding just the right 34 loose pages hidden in random books of 46 tall, long bookcases in a library. The table above shows that at least nine of the 23 human chromosome pairs would need to be accessed: Chromosome 1, Chromosome 5, Chromosome 6, Chromosome 7, Chromosome 11, Chromosome 12, Chromosome 16, Chromosome 17 and Chromosome 22.

nuclear pore complex
The nuclear pore complex (credit: Protein Data Bank, link)

Six Reasons These Molecular Machines And Their Behavior Are Accidentally Unachievable

There are six main reasons why we must regard the molecular machines described above as accidentally unachievable.

Reason #1: Chance processes such as Darwinian evolution could never produce the genes needed to make the proteins that make up such molecular machines (the gene origination problem)To perform the task a particular protein molecule performs, a type of protein molecule typically requires some specific fine-tuned gene, an amino acid sequence with most or nearly all of the protein's actual amino acid sequence, a chain of hundreds or thousands of amino acids specially arranged to produce a functional effect. Evolutionary biologist Richard Lewontin stated"It seems clear that even the smallest change in the sequence of amino acids of proteins usually has a deleterious effect on the physiology and metabolism of organisms." A biology textbook tells us, "Proteins are so precisely built that the change of even a few atoms in one amino acid can sometimes disrupt the structure of the whole molecule so severely that all function is lost." And we read on a science site, "Folded proteins are actually fragile structures, which can easily denature, or unfold." Another science site tells us, "Proteins are fragile molecules that are remarkably sensitive to changes in structure." paper describing a database of protein mutations tells us that "two thirds of mutations within the database are destabilising."  Those who think that functional folded protein molecules could gradually arise (getting longer and longer from a small size) will be dismayed to read this statement in a 900+ page textbook on protein chemistry: "Polypeptides less than about 70 amino acids in length should not fold because they should not be able to bury a large enough number of hydrophobic amino acids to overcome the configurational entropy of their random coils." Folding is required for most functional protein molecules. 

Accordingly, we cannot explain the origin of genes through some gradualism approach that imagines that first there was one tenth of the gene that was useful for one purpose, and then there was two tenths of the gene that were useful for some other purpose, and then finally we got the version of the gene that humans now have.  Human genes with only half of their base pairs or a third of their base pairs are not useful, and their corresponding protein molecules are not useful with half of their amino acids. 

But how hard would it be to get by chance or random mutations an amino acid sequence that would be the core of a useful protein molecule? That depends on the number of amino acids in the protein. Here we run into a simple principle that is the bane of all theories of accidental biological origins: the principle that a simple linear increase in the number of parts that must be well-arranged results in an exponential or geometric increase in the unlikelihood of such an arrangement occurring by chance. A small increase in the number of parts quickly results in what is called a combinatorial explosion, in which the number of possible combinations skyrockets. This is why computer security experts often tell you to use at at least 14-characters for the password of any financial account.  If you change your password from 7-characters to 14 characters, that doesn't make it merely twice as hard for a hacker trying all combinations to break into your account; instead it is is roughly 10,000,000,000 times harder. 

The chart below shows some of the relevant mathematics. If you doubt these numbers, you can verify them using the Large Exponents Calculator here. Since there are 20 different amino acids used in protein, you use 20 in the first row of such a calculator. Numbers such as E+6 refer to powers of ten. So 3.2 E+6 means 3,200,000; 1.024 E+13 means 10,240,000,000,000; and E+26 means 1 followed by 26 zeros. The bottom of the chart is a number of combinations equal to about 1 followed by more than 2600 zeros. 


Number of amino acids in a molecule

Number of possible combinations of the molecule's amino acids

5

3.2 E+6

10

1.024 E+13

20

1.048576 E+26

40

1.099511627 E+52

80

1.208925819 E+104

160

1.461501637 E+208

320

2.135987035 E+416

640

4.562440617 E+832

1280

2.081586438 E+1665

2000

1.148130695 E+2602


We can see from the chart above that the odds become utterly prohibitive once you start to get amino acid lengths much longer than about 160. Even if you very generously assume that a particular protein molecule only needs to have half of its amino acid sequence matching its actual sequence (an assumption too generous because of what we know about the sensitivity of protein molecules to small changes), you still have a case where we should never expect chance processes to produce successful amino acid sequences (corresponding to functional protein molecules) as long as 320 amino acids. 

In most of the protein complexes described above, we have some very complex proteins consisting of very long amino acids chains that we should never expect to have arisen by chance or Darwinian processes, never in the entire visible universe even given billions of years. Specifically:

  • One of the complexes (the spliceosome) had a protein consisting of 2335 well-arranged amino acids.
  • Another of the complexes (the apoptosome) had a protein consisting of 1248 well-arranged amino acids. 
  • The nuclear pore protein complex had one protein requiring 2090 well-arranged amino acids, and another protein requiring 2012 well-arranged amino acids, along with three other types of proteins each requiring more than 1000 well-arranged amino acids.
  • The origin recognition complex/replicative helicase complex required 7 types of proteins that each required more than 700 well-arranged amino acids. 
  • The RNA polymerase II protein complex described above had five types of proteins each requiring more than 2000 well-arranged amino acids, and three other types of proteins each  requiring more than 1000 well-arranged amino acids.
I could say much more about why proteins with amino acid sequences as long as this are not explicable by Darwinian processes, but that would involve repeating too many of the points in a previous post. See my previous post here for quite a long discussion on why it is not credible to suppose that fine-tuned amino acid sequences of this length ever could have arisen through any type of natural selection.  

Reason #2: we lack any explanation as to why very complex proteins would fold correctly. To be functional, proteins have to fold in just the right way, to achieve very complex three-dimensional shapes. But we don't understand how this folding occurs. DNA specifies only the linear sequence of the amino acids that make up a protein, not the complex 3D shape of a protein. Don't be fooled by press accounts claiming that the AlphaFold2 software did something to solve the protein folding problem. Such software did not produce any progress in solving the protein folding problem (the problem of how proteins are able to fold into the complex 3D shapes needed for their function). Such software merely produced progress in a different problem: the protein folding prediction problem, which is the problem of predicting the 3D shape of a protein from its amino acid sequence. 

One maneuver is an appeal to what is called Anfinsen's Dogma, a claim that the 3D shape of a protein is entirely a function of its amino acid sequence. Such an appeal is futile because it is a "rob Peter to pay Paul" affair rather like "solving" your college tuition burden by charging your tuition on your credit card.  If Anfinsen's Dogma were true, then genes would all-the-more-enormously have to be "just right" to allow for a properly folded 3D protein molecule; and in that case the gene origination problem becomes exponentially worsened.  The person who appeals to Anfinsen's Dogma lessens the protein folding problem at the expense of exponentially worsening the gene origination problem  (the problem of how 20,000+ suitable genes ended up in human DNA).    Appealing to Anfinsen's Dogma seem to make Reason #2 of these six reasons seem less convincing, at the expense of making Reason #1 seem enormously and exponentially more convincing. Such an appeal produces no net progress in making molecular machines like those above seem accidentally achievable. 

I may note that there are very good reasons for rejecting Anfinsen's Dogma, such as the very massive reliance of protein folding on helper molecules called chaperone proteins, which show that the 3D shapes of proteins are not a simple function of their amino acids sequences as Anfinsen's Dogma claims. 

Reason #3: we have no credible physical explanation for how  a transcription event could promptly find the right gene to make a particular protein (a "needle from the haystack" type of event). Cells are constantly creating new proteins to replace proteins that disappeared because of the short lifetimes of proteins. The page here has a chart showing the lifetimes of human proteins, and we see a bar graph showing most of the proteins have a half-life between about 10 hours and 70 hours. A muscle protein might live for three weeks, but a liver protein might live for only a few days. To create new proteins, a cell uses a process called gene transcription. In this process a particular gene in DNA will be converted to a messenger RNA molecule that helps to build the new protein. 

Cell transcription occurs quickly. The source here lists a time of ten minutes for a gene to be transcribed by a mammal, but another source lists a speed of only about a minute. The great majority of that is used up by the reading of base pairs from the gene, with typically more than 1000 base pairs being read each time a gene is transcribed. The finding of the correct gene to read in DNA seems to occur in only seconds, not minutes, or at most a few minutes. 

Descriptions of DNA transcription fail to explain a huge issue: how does a cell find the right gene in DNA so quickly? Human DNA contains more than 20,000 genes, each of which is just a section of the DNA. The DNA is like an extremely long necklace of many thousands of beads, and a typical gene is like a group of several hundred of those beads. We should actually imagine multiple such necklaces, because DNA is scattered across 23 different chromosome pairs. Now if genes had gene numbers, and DNA was a set of numbered genes in numerical order, it might be easy to find a particular gene. So if a cell knew that it was trying to find gene number 4,233, it could use a binary search method that would allow it to find that gene pretty quickly. 

But no such method can be used within the human body. Genes do not have gene numbers that can be accessed within the human body, and DNA is not numerically sorted. DNA has no indexes that might allow a cell to find some particular gene that it was trying to find within DNA.  So we have an explanatory "needle in a haystack" problem.  Or we might call it a "needle in the haystacks" problem, because human DNA is scattered across 23 different chromosome pairs, as shown in the diagram below:


A scientific text tells us some information that makes this explanatory problem seem more pressing:

"One might have predicted that the information present in genomes would be arranged in an orderly fashion, resembling a dictionary or a telephone directory. Although the genomes of some bacteria seem fairly well organized, the genomes of most multicellular organisms, such as our Drosophila example, are surprisingly disorderly. Small bits of coding  (that is, DNA that codes for ) are interspersed with large blocks of seemingly meaningless DNA. Some sections of the  contain many genes and others lack genes altogether. Proteins that work closely with one another in the cell often have their genes located on different chromosomes, and adjacent genes typically encode proteins that have little to do with each other in the cell. Decoding genomes is therefore no simple matter. Even with the aid of powerful computers, it is still difficult for researchers to locate definitively the beginning and end of genes in the DNA sequences of  genomes, much less to predict when each  is expressed in the life of the organism. Although the DNA sequence of the human genome is known, it will probably take at least a decade for humans to identify every gene and determine the precise  sequence of the protein it produces. Yet the cells in our body do this thousands of times a second."

We have here a very severe navigation problem. A cell is somehow able to find the right gene in only seconds or a few minutes when a new protein is made, even though DNA and chromosomes seem to have no physical organization that could allow for such blazing fast  access to the right information. In an article on Chemistry World, we read this:

"How does the machinery that turns genes into proteins know which part of the genome to read in any given cell type? ‘To me that is one of the most fundamental questions in biology,’ says biochemist Robert Tjian of the University of California at Berkeley in the US: ‘How does a cell know what it is supposed to be?"

Biochemist Tjian has spoken just as if he had no idea how it is that a cell is able to navigate to the right place to read a particular gene in DNA. Later in the article we read this:

"For one thing, the regulatory machinery ‘is unbelievably complex’, says Tjian, comprising perhaps 60–100 proteins – mostly of a class called transcription factors (TFs) – that have to interact before anything happens. ....As well as promoters, mammalian genes are controlled by DNA segments called enhancers. Some proteins bind to the promoter site, others bind to the enhancer, and they have to communicate. ‘This is where things get bizarre, because the enhancer can sit miles away from the promoter,’ says Tjian – meaning, perhaps, millions of base pairs away, maybe with a whole gene or two in between. And the transcription machinery can’t just track along the DNA until it hits the enhancer, because the track is blocked. In eukaryotes, almost all of the genome is, at any given moment, packaged away by being wrapped around disk-shaped proteins called histones. These, says Tjian, ‘are like big boulders on the track’: you can’t get past them easily.... ‘Even after 40 years of studying this stuff, I don’t think we have a clear idea of how that looping happens,’ says Tjian. Until recently, the general idea was that the TFs and other components all fit together into a kind of jigsaw, via molecular recognition, that will bridge and bind a loop in place while transcription happens. ‘We molecular biologists love to draw nice model schemes of how TFs find their target genes and how enhancers can regulate promoters located millions of base pairs away,’ says Ralph Stadhouders of the Erasmus University Medical Centre in Rotterdam, the Netherlands. ‘But exactly how this is achieved in a timely and highly specific manner is still very much a mystery.’ "

Later in the article Tjian says he was shocked by the speed at which some of the process occurs. He expected it would take hours, but found something much different:

"The residence times of these proteins in vivo was not minutes or hours, but about six seconds!’, he says. ‘I was so shocked that it took me months to come to grips with my own data. How could a low-concentration protein ever get together with all its partners to trigger expression of a gene, when everything is moving at this unbelievably rapid pace?’ "

The rest of the article is just some speculation, which Tjian mostly knocks down, and the article itself calls "hand-wavy." We are left with the impression that no one understands how cells are able to instantly find the right gene.


Reason #4: chance processes would never produce the arrangements of proteins like those found in the molecular machines listed above. What we must never forget is that a protein complex involves three types of organization:
  • The one-dimensional organization of amino acids found in the sequence of amino acids that makes up a protein;
  • the three-dimensional organization of such a sequence to make a complex folded three-dimensional shape needed for a particular protein molecule to function properly;
  • the entirely different three-dimensional organization needed for the proteins of a protein complex to fit together in the right way to make a physical arrangement so complex that it may be called a "molecular machine."
How is it that protein molecules form into protein complexes consisting of multiple protein molecules? Some may guess that DNA is read to determine which type of proteins should team up with other proteins to make particular protein complexes.  But that does not happen. DNA does not specify which proteins belong to particular protein complexes. In fact, the tables above show that the genes corresponding to the proteins that make up the protein complexes are typically found in widely scattered chromosomes. That would seem to be the opposite of what would happen if DNA was specifying that particular protein molecules should team up with other types of protein molecules to make particular kind of protein complexes. 

So what explanation do biologists give for how protein complexes form? Their attempts at explanations consist of little more than hand-waving.  They mainly appeal to chance collisions of molecules floating around in the human body. This is no more credible  than claiming that tornadoes passing through junkyards can create automobiles out of the spare parts that are lying around the junkyards. 

A very important point here is that vast wonders of molecular assembly are happening continuously in the human body. Every week very many of the molecular machines described above (and countless others not described) are being assembled in your body. And as I have shown above, such molecular machines are built using amino acid sequences from very scattered chromosomes So we have an effect no more explainable by chance collisions than tornadoes building cars out of junk scattered in diverse places of a junk yard. And such an effect is constantly occurring in your body, in massive numbers. 

There is no way to explain this by trotting out some Darwinist phrase such as "very lucky things can happen given a million years of chance events."  We are not talking here about some miracle of genetic luck that occurred once in an eon. We are talking about miracles of complex purposeful assembly that are occurring in massive numbers every single day in your body. Darwin doesn't do anything to get the materialist out of this jam. 

The statements below are indications that scientists simply have no credible explanation as to how very complex protein complexes (like those discussed above) can form from their constituent protein parts:

  • "The majority of cellular proteins function as subunits in larger protein complexes. However, very little is known about how protein complexes form in vivo." Duncan and Mata, "Widespread Cotranslational Formation of Protein Complexes," 2011.
  • "While the occurrence of multiprotein assemblies is ubiquitous, the understanding of pathways that dictate the formation of quaternary structure remains enigmatic." -- Two scientists (link). 
  • "A general theoretical framework to understand protein complex formation and usage is still lacking." -- Two scientists, 2019 (link). 
  • "Protein assemblies are at the basis of numerous biological machines by performing actions that none of the individual proteins would be able to do. There are thousands, perhaps millions of different types and states of proteins in a living organism, and the number of possible interactions between them is enormous...The strong synergy within the protein complex makes it irreducible to an incremental process. They are rather to be acknowledged as fine-tuned initial conditions of the constituting protein sequences. These structures are biological examples of nano-engineering that surpass anything human engineers have created. Such systems pose a serious challenge to a Darwinian account of evolution, since irreducibly complex systems have no direct series of selectable intermediates, and in addition, as we saw in Section 4.1, each module (protein) is of low probability by itself." -- Steinar Thorvaldsen and Ola Hössjerm, "Using statistical methods to model the fine-tuning of molecular machines and systems,"  Journal of Theoretical Biology. 

Reason #5: once assembled, such molecular machines act as if they knew where to go, which would not happen by accident . It is not merely the assembly of such molecular machines that defies anything we should expect to occur accidentally. It is also the behavior of such molecular machines, in the sense that they always seem to act exactly as if they knew where to go to. For example, the nuclear pore complex molecular machines go to just where they are needed (the nuclear membrane), and the splicesome and RNA polymerase II complex go to appropriate places in the cell.  To have an analogy for the whole storyline of the construction and target reaching of such molecular machines occurring accidentally, we would have to imagine something like tornadoes passing through junkyards, constructing many cars, and also blowing the cars to just the right places to pick up a million scattered people who needed rides.  How accidentally unachievable would that be? 

molecular machines acting like motors

Reason #6: such molecular machinery behavior occurs massively every dayIf something occurs only very rarely, we might regard is as accident. For example, if you come to your door with a friend, and realize you lost your  key, and your friend suggests he tries his key on your door, and his key opens your door, you may regard this as a lucky coincidence, having seen such a thing only once in your life. But when some type of lucky thing occurs all the time in massive numbers, in some way you cannot account for, chance and coincidence are no longer reasonable explanations. 

How often does there occur in your body the assembly and correct positional targeting of the molecular machines I have listed above? Billions or trillions of times every day. For example, very many  instances of the nuclear pore complex discussed above are needed in the construction of a new cell, and it has been estimated that the human body makes 300 billion new cells every day. What would be the chance of the totality of such daily feats of construction occurring accidentally? Something like the chance of a winter ice storm constructing a thousand-mile-high ice arch stretching all the way from Europe to North America. 

Tuesday, September 5, 2023

What You Read Would Be So Different If Hi-Tech Companies Were Better at Judging Reliability

A small class of literature gatekeepers has enormous control over what information and opinions end up being viewed by the masses. Such a class includes publishers, editors, peer reviewers, site moderators, book store owners, librarians, and people at high-tech companies who control what type of posts and articles end up in search results and news feeds. We often fail to realize how much power such people once had and still have over what type of things we read.  One key point often overlooked is that enormous control is influenced by people who are not preventing publication of anything but merely controlling likelihood of readership. For example, a librarian at a public library exerts enormous control by deciding which books end up on the libraries of bookshelf, and also by putting certain books on a "New Books" shelf or a "Recommended Books" shelf where they will be far more likely to be found. 

One obvious way in which this class of literature gatekeepers exerts control is by approving and rejecting proposed articles, papers and manuscripts. Scientists who act as anonymous peer reviewers help to enforce prevailing groupthink and reigning dogmas by rejecting for publication papers that defy assumptions that supposedly reign in their fields of study.  The people who perform such censorship  often justify it as "quality control."  The class of book editors and newspaper editors once exerted the most enormous control over what the public read, by controlling what ended up on the printed page. With the rise of the Internet, the enormous power of such a class has been greatly lessened because of the ease in which people can publish content online. 

But as this class of literature gatekeepers lost a good deal of their once enormous power, another class of literature gatekeepers had an enormous rise in their power. This was the class of gatekeepers controlling Internet search results and the content of digitized feeds such as Facebook feeds and the Apple News app feed.  Consider the average person today. He will probably not visit a public library very often, and he will not visit a bookstore very often. But almost every day such a person will search for information using a search engine such as Google or Bing. And almost every day such a person will access various types of feeds that provide a stream of stories, articles and posts.  Among the most popular feeds are the Google News feed, the Apple News feed, and the Facebook feed that is typically a mixture of posts and stories from people you don't know, and posts from people you do know, who are some of your Facebook friends. 

The most enormous power is possessed by the people who control the search results of search engines and the news feeds such as those mentioned above. There are strong reasons for suspecting that such persons are not very skillfully using such power. Some reasons are the abundance of low-quality items appearing in search engine results and news feeds such as the Google News feed and the Apple News feed. Below are some of the main examples of these low-quality items:

(1) Anonymous clickbait results appearing in news feeds such as science news feeds.  Clickbait is a gigantic problem that mars the reliability of stories appearing on feeds such as Science news feeds. There currently exists an economic ecosystem that strongly incentivizes the appearance of interesting-sounding but misleading science stories. This ecosystem is described in my post "Why the Academia Cyberspace Profit Complex Keeps Giving Misleading Brain Research Reports."  What I say in that post about brain research holds true for many other types of science. 

After a scientific paper has been written up and published, it is announced with a press release issued by the main academic institution involved in the research. Nowadays the press releases of universities and colleges are notorious for making sensationalized claims that are not warranted by anything discovered in the research being discovered. Often a tentative claim made in a scientific paper (basically a "perhaps" or a "maybe") will be stated as if it is was simply a discovery of a definite fact.  Other times a university press release will make some important-sounding claim that was never made in the scientific paper writing up the research.  

Authorship anonymity is a large factor that facilitates the appearance of misleading university and college press releases.  Nowadays university and college press releases typically appear without any person listed as the author. So when a lie or misleading statement occurs (as it very often does), you can never point the figure and identify one particular person who was lying.  When PR men at universities are thinking to themselves "no one will blame me specifically if the press release has an error," they feel more free to say misleading and untrue things that make unimpressive research sound important.  

There are complex economic reasons why press releases so erroneous keep appearing so often, and why they are passed on in clickbait Internet stories that lead to pages containing ads that generate revenue. To understand those reasons you have to "follow the money" and look at which parties are profiting from such unreliable but interesting-sounding stories. The reasons are explained in this post, and sketched in the diagram below.

who profits from science clickbait

Ads on web pages are a pretty good indicator of whether clickbait is occurring, although not all clickbait involves pages with ads. For example, a university press office may publish a press release web page with no ads, but a clickbait headline not matching anything actually discovered in the described research.  The purpose in this case is not generate ad revenue, but largely to glorify the university or its researchers. 

(2) Anonymous web pages lacking a statement of authorship by one particular person, often containing ads.  When a page lacks a statement of authorship by one particular person, the person or people writing the page may feel free to make dubious or false statements, having no sense that they will be held personally accountable. Many pages containing errors and distortions appear at the top of Internet search results when you search for a particular topic. Such pages often include the anonymously written pages of wikipedia.org, which often have errors, particularly when they deal with controversial topics.  Apparently the algorithms of companies such as Google fail to properly penalize authorship anonymity. 

Often such pages contain ads, and the presence of such ads compromises the credibility of such pages. The ads make us wonder whether the primary purpose of the page was to provide accurate information, or instead to serve as a vehicle for generating revenue by the display of the ads. 

An example of a rather poor page returning near the top of Google search results is the page here. The page appears as the third result when I search for "memory recall" using Google. The page has no listed author.  The page has near its beginning two ads, one very large.  One is a dubious ad hawking CBD oil "for brain and memory." There is little evidence that CBD oil does anything to benefit memory. The page also has near its beginning a column of "product reviews" that are basically ads.  The page includes several important claims that are very dubious and not backed up by any evidence. The page claims, "Increased activity in Globus pallidus, anterior cingulate gyrus, thalamus, and cerebellum is seen during recall."  There is no robust evidence for such a thing. See my page here for a discussion of the relevant evidence, which fails to show any clear evidence of some part of the brain working harder during recall or recollection.  The references the page gives are mostly to very old papers or low-quality papers, such as a study of monkeys using a sample size of only one. 

(3) Pages on very subtle and complex topics requiring very many hours of thought or scholarship, written by people who have never written much or studied extensively the topic they are writing about.  

We see examples of such pages showing up extremely abundantly in the top search results returned by search engines such as Google and Bing. An example is that when I search for "memory recall" on Google, I get as the third item in the search results an article written by Emilie Le Beau Lucchesi. A page describing the author mentions a bachelor's degree in journalism and a PhD in communication, "with an emphasis on media framing, message construction and stigma communication." This does not suggest the author is a scholar of the brain. Lucchesi's article contains important statements that are unfounded. Lucchesi makes the incorrect claim that "Scientists continue to learn how the brain stores and retrieves memories using brain mapping technology." No such thing has happened, and scientists are not at all learning how a brain could either store or retrieve memories. Scientists lack any credible theory of how a brain could do either of these things. Lucchesi writes this:

"Scientists use the term engram to describe the physical process the brain uses for holding a specific memory. In the past, scientists identified memory engrams in the hippocampus, amygdala and cortex."

To the contrary, no robust evidence has ever been found of engrams in the human brain. No claims to have discovered such claims will hold up to diligent scrutiny. Scientists have never found the slightest solid evidence of human learned information by microscopically examining brain tissue living or dead. Lucchesi's article refers at length to some mouse study but does not mention either the name of the study or give a link to it, merely mentioning its lead author. Lucchesi is probably referring to the study here, which is a Questionable Research Practices study involving way-too-small study group sizes such as only 10 mice and only 12 mice, with the study group sizes sometimes dropping down to ridiculously small sizes such as only 4 mice per study group.  The poorly-designed study used no pre-registration, no blinding protocol, and it confesses, "No statistical methods were used to predetermine sample sizes."  A decently designed experimental paper will use statistical methods to calculate an adequate sample sizes (in other words, study group sizes), and then use such sample sizes. 

We can excuse Lucchesi for her misstatements in this article and for discussing this poor mouse study, because brains and neuroscience are not any areas she has written much about, and none of her three books deal with such a topic or any scientific topic. If search engines were better, her erring article on a topic she rarely writes about would not have appeared on the first page of search results when I searched for information on "memory recall." 

What steps could hi-tech companies take to reduce so much junk from appearing at the top of their search results, and to reduce the occurrence of so many misleading stories on their news feeds? Among the steps they could take would be these:

(1) Severely penalize all anonymous authorship, including pages anonymously written and listing some prestigious institution as ita source.  Algorithms and human judges should severely penalize all web pages, press releases, posts and articles written by anonymous authors, regardless of whether such literature lists some prestigious institution as its source. For example, a press release issued by Harvard University should be treated as probable unreliable junk unless it lists one particular person (or one or two persons) as its author.  Since in recent years the nation's most prestigious institutions (such as leading universities and NASA) routinely have issued erroneous and unreliable press releases without listing specific human authors of such press releases, anonymous authorship from some esteemed authorship should not be regarded as any indication of reliability. If the major hi-tech companies were to announce such a policy, the colleges and universities that keep polluting our science news feeds with low credibility anonymously-written press releases would notice that their press releases are not attracting attention, and would switch to press releases with a single named author. Doing that would improve reliability, because a person is more likely to be truthful when writing a page that lists himself as the author.

Currently we seem to have the opposite of such a principle occurring. Search engines routinely display anonymously written pages on wikipedia.org at the top of search results. Such pages are often of very poor quality, particularly when controversial topics are discussed. 

(2) Severely penalize all pages with ads, with the penalty proportional to the number and size of ads on a page.  Ads on web pages generate revenue for the people publishing the site. The presence of a single ad on a web page is a reason for doubting the credibility of the statements on that page. As soon as we see an ad on a web page, we should start asking: was the page written primarily to teach truth, or to generate revenue for the web site? The more ads we see on a page, the more we should suspect that the page was written and published primarily to generate ad revenue. There is a direct relation between pages written primarily to cause ad views and the credibility of the pages. If you are writing a page mainly for the sake of generating ad views, you will tend to write enticing headlines or make sensational claims not justified by any evidence you site, because such headlines and content works better as clickbait. For example, if you write a sensational hype headline of "Breakthrough in the search for extraterrestrials" (one not matching anything discovered), that clickbait headline can appear as a link on other pages, causing many people to click on the link to get to your page.  

(3) Penalize in search results all pages written by authors writing on a complex topic that they have not written much about, and promote in search results pages written by authors writing on a complex topic that they have written much about.

Major search engine companies engage in constant "web crawling" in which their automated "spiders" search every corner of the internet, and analyze the results. It would not be very hard for "rich as Croesus" search engine companies to maintain databases that try to get a rough idea of  authors and their expertise or what they have often written about. For example, if there are twenty pages scattered around the web in which Bob Worthington offers long reviews of military action games, then a search engine company should be able to detect such a pattern, and record Bob Worthington as something of an expert on military action games. Similarly, if such a search engine picks up a claim that Sally Jenkins has a PhD in cell biology, then it should be able to store that fact in some expertise database. 

Probably the search engine companies already have something like such databases. But they don't seem to use them well. What should occur is something like this: when someone writes for the first time on some deep topic that he has never written about, such a page should be penalized in search results. But when someone writes about a complex topic he has written very much about, such a writing should be promoted in search results, getting a higher rating. 

If such a principle were followed, we would not get in our search results a result like the very poor new piece at the frequently-erring but slick site Quanta Magazine, a piece entitled "The Usefulness of a Memory Guides Where the Brain Saves It." The article repeats the myth that patient HM could not form new memories (a myth I debunk here), and repeats the groundless achievement legend that research on him "helped scientists discover that new memories first formed in the hippocampus and then were gradually transferred to the neocortex," a migration myth that multiplies the explanatory problems of explaining how a brain could store memories.  The writer (identified as a "writing intern") also fails to see how bad it is that the "evidence" provided to back up the speculative theory he is promoting is imaginary-data evidence consisting not of real tests with real subjects but merely tests with simulated humans in a computer experiment (which is rather like trying to prove the effectiveness of your new gun model by showing it seems to work well inside the play world of a video game).  

(4) Penalize scientific papers hidden behind paywalls,  have a system for rating the quality and relevance of scientific papers and scientific articles that show up on the first page or two of search results, and use such a system to quality-check results appearing very early in search results. 

Very often on the first page of the search results of major engines, we get links to very poor scientific papers very guilty of Questionable Research Practices. An example is when I search for "memory recall" on Google, I get on the first page of search results a link to the poorly designed study "Memory recall involves a transient break in excitatory-inhibitory balance." The study is guilty of the usual Questionable Research Practices so abundant in today's experimental neuroscience: lack of pre-registration, lack of any blinding protocol, the use of way-too-small study group sizes as small as only 12, and the lack of any sample size calculation to determine whether the samples sizes used were adequate. A study last year found that thousands of participants are needed for accurate correlation studies involving brain scanning, but the paper in question (relying on brain scanning) used only about 19 subjects. 

Another example of a poor paper showing up on the first page of Google search results is the paper "Prefrontal feature representations drive memory recall." The paper is behind a paywall, and does not mention how many subjects it used; we may presume it is some way-too-small sample size (when scientists use decent-sized study group sizes, they almost always mention such sizes in their paper abstracts). The study involved scanning the hippocampus region of brains of mice while they performed a memory activity, a region maybe the size of a grain of sand. Given the very tiny size of a hippocampus in mouse brains, that is such a ridiculous and error-prone method that it is virtually never used. Although I can't judge the paper, based on its abstract I can say there seems to be no sense at all in including a link to it on the main page of Google search results when someone uses the phrase "memory recall."  All scientific papers behind paywalls should be severely penalized in search results, not appearing on the first three pages of search results.  

But, you may object, it would be too hard for a search engine company to evaluate all of the countless thousands or millions of papers published each year, to rate their quality. But such an effort would not be needed.  Such a search engine company could simply rate the papers that show up on the first page or two of search results when someone searches for common search phrases such as "exercise benefits," "memory recall," "COVID prevention," and so forth.  Doing that would require evaluating the quality of only a few thousand papers. 

(5) Promote in search results lengthy "deep dive" articles sounding like someone had researched a topic for a long time, and penalize vacuous-sounding articles reading like college freshman  efforts. 

When I search for "memory recall" using Google, I get a vacuous-sounding article as the second search result. The anonymously-written article sounds like something a college freshman might have written while half-watching a movie on TV. It includes no references to other sources, no mention of any specific observations, and seems to tell us nothing useful. Some of its generalizations are untrue, such as its goofy generalization that "every time a memory is accessed for retrieval, that process modifies the memory itself," and its generalization that "the ability to access a given memory typically declines over time." No, I very clearly remember saluting the passing coffin of President John Kennedy, and there has been no decline or change in that memory over time.  Why is an article like this appearing as the second result in Google's search results when I search for "memory recall"?  Their algorithm for producing search results clearly needs more work.

Friday, September 1, 2023

The Data of the "Starship Smithereens" Paper Shows That Nothing Very Odd Was Found

Harvard astronomer Avi Loeb somehow got the idea that a  2014 meteor (the CNEOS 2014-01-08 meteor) may have been an interstellar spacecraft that blew up high in the sky. Loeb has recently finished his million-dollar oceanic expedition looking for what he hoped would be remnants of a crashed extraterrestrial spaceship, an expedition he organized.  He found no sign of anything looking like a spaceship or any of its parts. Loeb claims to have found tiny round specks only about a millimeter in size. All that he recovered were some tiny metal specks. The metal specks he found are just like metal sea specks found all over the world.  But it seemed like Loeb was trying to convince the press that he discovered smithereens of a starship. 

 The result was stories such as a CBS News story story entitled "Harvard professor Avi Loeb believes he's found fragments of alien technology." Analyzing such stories it seems hard to pin down Loeb as explicitly stating that he believes the sea specks he found are starship smithereens, specks of an extraterrestrial spaceship. But clearly Loeb was doing very much to raise such an idea in the minds of the press, and he was doing nothing to correct story titles like the CBS News story title.

Now finally Loeb's team has released a preprint of a scientific paper on the tiny sea specks that were dredged up.  It seems that Loeb's grand claims have been ramped down very much. The paper makes no explicit claims to have discovered any traces of an interstellar spaceship. It merely claims that traces were discovered that "likely" have some interstellar origin.  The paper mentions a purely natural origin, saying, "We suggest that the 'BeLaU' abundance pattern could have originated from a highly differentiated magma ocean of a planet with an iron core outside the solar system or from more exotic sources." The first part refers to a purely natural origin (such as a meteor from another solar system), and the "more exotic sources" may refer to all kinds of other possibilities, such as an extraterrestrial spaceship.

Is there any justification for these claims that the sea specks recovered by Loeb's expedition likely came from beyond our solar system? There is not.  The paper presents no good evidence that the dredged-up sea specks came from beyond our solar system. The paper is also guilty of a "hide the bad news" presentation, where it's like the authors are trying to make it very hard for readers to discover the relevant facts behind their central claim. 

Referring to the path of a meteor that exploded in the sky, and referring to "spherules" that are tiny round specks dredged up from the ocean, the paper claims this: "Mass spectrometry of 47 spherules near the high-yield regions along IM1’s path reveals a distinct extra-solar abundance pattern for 5 of them, while background spherules have abundances consistent with a solar system origin." This claim about a special status of 5 of the 47 speck-sized spherules is false, and we should immediately be suspicious upon hearing this claim of "a distinct extra-solar abundance pattern." Humans have not well-studied the compositions of objects known to have entered our solar system from beyond the solar system. So there could never be a match in which someone found some particular chemical composition in a rock or sphere or spherule, and said, "Yes, that matches the characteristics of interstellar objects."  In fact, the five supposedly special speck-like spherules Loeb refers to have the same composition as countless other such spherules scattered all over the world. 

In Section 7.6 of the paper, the authors make this incorrect claim: "The spherules with enrichment of beryllium (Be), lanthanum (La) and uranium (U), labeled 'BeLaU', appear to have an exotic composition different from other solar system material." To try and back up its claim about five of its tiny sea specks, the paper refers us to its graphs 12, 13 and 14. Those are extremely confusing graphs that use a logarithmic scale, and fail to tell us what is the elemental composition of the five supposedly special specks.  

It is as if the paper was tying to prevent us from easily discovering the composition of these five supposedly special specks. To clearly display their composition would be very easy to do. The paper could have had five nice clear pie charts, that anyone could have easily read to discover the composition of these  five supposedly special specks. Instead we have graphs that seem like that they were designed to be as confusing as possible. The appendix of the paper reveals the composition of the these  five supposedly special specks, but in a way that is hard to decipher. There is a table that gives the data. The table appears in landscape mode, meaning you have to tilt your head to read it.  Also, violating the rules of clear data presentation, each column of element abundances is stated using a different scale.   

confusing science paper
             Part of Appendix A,  rotated to make reading easier

Laboriously jumping through hoops to read this data, you can get to the bottom of the matter. The six supposedly special specks (labeled with a subclass of BeLaU) appear at the top of the Appendix 1 table, on page 28 of the paper. At first it looks like some of these specks have lots of Beryllium, but that isn't the case. If you look at the Beryllium column header (labeled Be), we see that the Beryllium numbers are stated in fractions of parts-per-million, fractions of .025 parts per million. So when we see 4587 as the  Beryllium abundance of one of the BeLaU specks, the biggest Beryllium abundance of any of the specks Loeb dredged up from the sea, that merely means a Beryllium abundance of 4587 times .000001 times .025, which equals a Beryllium abundance of 0.00011, about 1 part in 10,000. 

Similarly, if you look at the Lanthanum column header, on page 29 of the paper, you see that the Lanthanum numbers are in fractions of parts-per-million, fractions of .235 parts per million. So when we see 1108 as the Lanthanum abundance of one of the BeLaU specks, the biggest  Lanthanum abundance of any of the specks Loeb dredged up from the seathat merely means a  Lanthanum  abundance of  1108 times .000001 times .235 which equals a Lanthanum abundance of 0.00026, only about 2 parts in 10,000.

Similarly, if you look at the Uranium column header on page 30 of the paper, you see that the Uranium numbers are in fractions of parts-per-million, fractions of .0081 parts per million. So when we see 1892 as the Uranium abundance of one of Loeb's specks, the biggest Uranium abundance of any of the specks Loeb dredged up from the seathat merely means a  Uranium  abundance of  1892 times .000001 times .0081 which equals a Uranium abundance of 0.000015, only about 15 parts in a million. 

The table below summarizes the data, showing the strangest things Loeb was able to find in his little specks, after exhaustively looking for any strange thing. The fractions displayed are simple abundance ratios (so, for example, .000001 means 1 part in a million, and .0001 means one part in 10,000). 


Element

Highest amount in any of  Loeb's spherule specks 

Amount in meteorite (reported before 2023)

Amount in rocks (reported before 2023)

Amount in tiny spherules (reported before 2023)

Beryllium

.00011

.000000386

.000003

.000049 to .000200

Never reported?

Lanthanum.

.00026

?

.0000038

.0000095

.000297 to .000920 (in Finland)

Never reported?

Uranium

.000015

.00000017

.00000022

.000090 

 (in Chinese phosphate rocks)

.000090 

 (in rocks from various countries)

Never reported?

Two papers I link to in the last row refer to uranium levels of about 90 milligrams per kilogram in earthly rocks, which is an abundance of about .000090.

The results shown above are the strangest bit of strangeness that the Loeb "starship smithereens" expedition has to report. And it's nothing very strange at all. It is merely that in one of the 70+ tiny spherules that were analyzed, there was maybe a tiny bit more beryllium and  maybe a tiny bit more lanthanum than you might have expected to find, based on previous analytic reports analyzing these elements in meteorites and rocks.  The reported Beryllium level of Loeb's sea speck with the most Beryllium did not even match the highest level of Beryllium reported in igneous rocks, being only half of the Beryllium  level of 200 parts per million reported in the paper hereThe reported Lanthanum level of Loeb's sea speck with the most Lanthanum did not even match the highest level of Lanthanum reported in Finnish rocks, being three times smaller than the Lanthanum level of 920 parts per million reported in the paper here. The uranium level of the spherule with the most uranium is not even remarkable, and many earthly rocks have uranium levels far greater.  Does that mean something very  remarkable was found? No, it doesn't. Traces of rare elements are found in various concentrations that may easily vary by a hundred times from sample to sample. You could explain the whole difference under the simple idea that Loeb was using state-of-the-art equipment that is better at finding trace concentrations than the older equipment used to get the numbers in the right column above. 

A story in the tabloid press is claiming this about Loeb's spherules: "The lanthanum and uranium were 500 times more plentiful than in earthly rocks and beryllium hundreds of times so."  That is not at all correct. My table above shows that the highest level of lanthanum in any of Loeb's spherules was three times smaller than a level reported in Finnish rocks (920 parts per million), and that the highest level of beryllium found in any of Loeb's spherules was only half of a level of beryllium reported in some igneous rocks. There was no real uranium anomaly, since earthly rocks mined for uranium have even higher levels of uranium than in any of the specks (in parts per million). The average amount of lanthanum and uranium and beryllium found in the full set of Loeb's spherules was not more plentiful than in earthly rocks. 

Nothing very unusual has been found from Loeb's million dollar expedition. Loeb's paper has made the groundless claim that five of the tiny specks gathered by the mission "reveal a distinct extra-solar abundance pattern."  The data gathered by Loeb does not even suggest that the specks came from outer space. A 2001 scientific paper ("Magnetic spherules: cosmic dust or markers of a meteoric impact?") reports that tiny magnetic spherules have been found all over the world:

"In the past hundred years, magnetic spherules were found in various geological environments, namely in the Antarctic and Greenland ice and glacial sediments, in deep-sea floor cores, in meteorite fall areas...in volcanic and ..metamorphic rocks. Magnetic spherules found in recent sediments and oceanic floor around the industrial centers may also be the products of air pollution (probably over 99%)." 

When science is done properly, you wait for a decent amount of data justification before you go announcing grand conclusions such as visitations from outside of the solar system. You don't go drawing conclusions based on tiny irregularities in only five speck-sized things. And it's pretty ridiculous to take something that's probably the result of mere pollution and to claim that it came from another solar system. It's rather like someone in Los Angeles saying today's smog came from Alpha Centauri.  

Loeb recently made the groundless claim that some of his tiny sea specks came from the 2014 CNEOS 2014-01-08 meteor, which has been inappropriately given a name of IM1, standing for "interstellar meteor 1." We do not actually know that this meteor came from beyond the solar system.  Referring to the sea specks I discuss above, on his blog Loeb recently made this groundless claim: "Five of these millimeter-size marbles originated as molten droplets from the surface of IM1 when it was exposed to the immense heat from the fireball generated by its friction on air on January 8, 2014."   At www.space.com we read some reasons for rejecting all such claims:

"Matthew Genge, a planetary scientist at Imperial College London who specializes in meteorites, said that connecting the spheres with the 2014 fireball — or any meteorite fragments with any other meteor — is impossible. 'Meteorite ablation debris has been found, but not from an instrumentally observed fireball,' Genge told Space.com via email. 'There never has been a micrometeorite derived from a specific fireball event, and never will be, since it is an impossibility.' Peter Brown, an astronomer at the University of Western Ontario, agreed with Genge. If the meteor did in fact enter Earth's atmosphere at the speeds reported, Brown said, it would have been vaporized into fragments much smaller than the spherules Loeb's expedition discovered.  'There has never been a meteorite recovered from any object that hits the atmosphere moving at more than 28 kilometers a second [62,600 mph],' said Brown, who studies meteors and small solar system bodies such as asteroids. 'Any solids that would remain would be essentially aerosol-size.'  (In a 2022 paper in The Astrophysical Journal, Loeb claimed that IM1 was moving between 52 and 58 km per second, or 116,000 to 130,000 mph.)"

In the same blog post Loeb makes this simply untrue claim: "Five unique spherules... showed a composition pattern of elements from outside the solar system, never seen before." No, there was nothing very unusual about the element composition of the five strangest of Loeb's spherules, and the main irregularities are those I list in the table above, which are unimpressive. 

In a NY Post article we seem to have Loeb trying to create the misleading impression that the government estimated something with 99.999% confidence. We read this: "No less an unimpeachable source than the US Space Command went on to confirm, with '99.999 percent confidence' that the tiny spherical objects were interstellar, he said."  The government has said nothing at all about Loeb's sea-speck spherules.  Loeb is referring here to a letter from a government official who cited a '99.999 percent confidence'  estimate made by Loeb himself, not by the government. That estimate wasn't about any spherical objects Loeb recovered, but about whether the IM1/CNEOS 2014-01-08 meteor was interstellar.  No one at the government made any confidence estimate about either Loeb's IM1/CNEOS 2014-01-08 meteor or Loeb's recovered spherical objects.  The letter from someone in the government is shown in Loeb's post here.  When that letter refers to a 99.999 percent confidence estimate, it is merely referring to Loeb's own estimate, not a government estimate. 

In that NY Post article Loeb says, "The composition of uranium is 1,000 times what you find on earth.”  This is very untrue. The spherule speck that had the most uranium of all the sea specks Loeb recovered and analyzed had a level of uranium of 0.000015, several times less than the level found in earthly rocks mined for uranium. Such rocks have a level of about 90 milligrams of uranium per kilogram, as reported here, which is a level of 0.000090.

A very recent article in Science magazine gives us this quote about the paper I analyze above:

"But others are dismissive of the preprint, which has not been peer reviewed. Although the geochemical analysis of the debris is solid, the conclusions that Loeb and his colleagues hang on them are 'nonsense,' says Martin Schiller, a cosmochemist at the University of Copenhagen. 'I’m surprised anyone would take it seriously.' Larry Nittler, a cosmochemist at Arizona State University (ASU), calls it 'very weak sauce.' "

In the same Science article we read a reason for doubting Loeb's continued claims that the IM1/CNEOS 2014-01-08 meteor was interstellar, based on supposed confirmation from a government satellite: 

"A new study out this month in The Astrophysical Journal examined 17 known fireballs captured by both classified U.S. sensors and independent observations. The study showed that the government sensors often overestimated speed, with the errors getting worse the faster things got. 'A third of the time, the numbers are just way off,' says Steve Desch, an ASU astrophysicist."

During the Middle Ages for centuries there was a great enthusiasm for collecting the relics of saints. For centuries people would dig up bones or teeth (often just little bits of bone), and claim that they had magnificent healing powers on the grounds that they belonged to a canonized Catholic saint. Such relics would be displayed in churches, and people would make pilgrimages to see them. Such relic pitchmen remind me of Loeb's antics in trying to glorify his tiny sea specks that are probably mere specks of pollution. The difference is that the medieval relic stories actually seemed to do some good. Possibly because of a placebo effect, countless people would report cures after touching or seeing the alleged relics of some Catholic saint. But Loeb's sea speck antics seem to do very little good other than help sell more copies of his latest book.  

Here is the latest tabloid headline sounding like something that we might hear from a carnival barker:

"EXCLUSIVE: Truth about my 'alien' encounter... How I found bombshell interstellar objects a mile beneath the sea - and their limitless potential for life on Earth, by scientist AVI LOEB."

If you gave some specks of smudge pried from the bottom of your shoes the kind of "royal treatment" given to Loeb's sea specks -- that of analyzing the abundance of all of their elements -- you would probably be able to find some tiny bit of strangeness somewhere about as impressive as the strangest thing reported in the "starship smithereens" paper I discuss above. There no evidence of "starship smithereens" in Loeb's paper and no evidence of anything from beyond the solar system, but merely evidence of pareidolia in which a scientist claims to see some trace of something he is longing to see. No strong evidence has been provided that Loeb's spherule specks even came from beyond Earth. It reminds me of things I mention in my post "When Scientists Claim to See Things They Never Saw." 

Postscript: A co-author of the paper  (Charles Hoskinson) issued a tweet claiming, "We discovered spherules that appear to be from a different solar system due to their ultra high abundance of Beryllium, Lanthanum, and Uranium (thus BeLaU)."  But as shown above, none of the abundances were greater than 2 parts in 10,000, and all of the levels were similar to previously reported levels of such elements in some earthly rocks.  So what on Earth was Hoskinson doing using the phrase "ultra high abundance"?  Another scientist points out that 67 atomic bomb tests were done by the US between 1946 and 1958 within a few hundred kilometers of where Loeb's spherules were recovered, a fact that could account for minor element irregularities in such sea specks. 

In an October, 2023 paper Avi Loeb repeats his untrue claim that the US government made a  99.999% estimate about the likelihood of an interstellar origin of a meteor. In the paper he incorrectly states, "In 2022 the US Space Command issued a formal letter to NASA certifying a 99.999% likelihood that the object was interstellar in origin."  This statement has a reference to the document here, which merely mentions a 99.999% likelihood  made by Loeb himself, not the US government. The careless repetition of this misstatement by the press is an example of the kind of "hook, line and sinker" science journalism that goes on these days. 

Loeb states this:

" Mass spectrometry of 47 spherules near the high-yield regions along IM1’s path revealed a distinct extra-solar abundance pattern for 5 of them, while background spherules showed abundances consistent with a solar system origin. The unique spherules showed an excess of Be, La and U, by up to three orders of magnitude relative to the solar system standard of CI chondrites."

The claim of a "distinct extra-solar abundance pattern" for five of the tiny specks is groundless.  In the second sentence Loeb gives us two  examples of objectionable speech. First, he makes a comparison between his metal specks and "CI chondrites," which are primarily stony meteorites. Of course, if you compare two types of different things, there will be some discrepancy in the elemental abundances.  Also Loeb engages in the trick of claiming a difference of "up to three orders of magnitude,"  a very imprecise phrase that could refer to any difference between about 10 and about 1000.  The use of imprecise,  vague or misleading language such as this is an indication that we do not have an example of robust science, which is characterized by precise and accurate statements. 

In his latest October, 2023 paper Loeb suggests a non-technological explanation for his little sea specks, coming up with some wild speculation of weird natural events in some other solar system. But since the element content of the sea specks is not very unusual, and since there's no reason to think that specks have anything to do with CNEOS 2014-01-08, and since the case that CNEOS 2014-01-08 was from another solar system is weak,  this wild speculation has little credibility.   

With his typical dogmatism, Ethan Siegel has an article entitled "Harvard astronomer’s 'alien spherules' are industrial pollutants." It would have better to say that the stranger spherules have an element abundance similar to those of industrial pollutants.  Siegel mentions the new paper here, finding that the levels of Beryllium, Lanthanum, and Uranium in Loeb's strangest spherules are similar to those in coal ash pollutants. A later article by Siegel has more debunking of Loeb's spherule claims.