Showing posts with label evolution. Show all posts
Showing posts with label evolution. Show all posts

Thursday, July 9, 2026

A new version of my review article on emergence

On the arXiv, I have posted a new version of my review article, Emergence: from physics to biology, sociology, and computer science.

I have added expanded sections on molecular structure, quantitative measures of causal emergence, and biological evolution.

There are also many minor additions and corrections. I hope the hyperlinked Table of Contents is helpful.

I welcome feedback and suggestions. I am sure there is much more to do.

Tuesday, June 30, 2026

Biological evolution and emergence

 The theory of evolution explains the origin of biological diversity and levels of similarity between species. A characteristic of emergence is that many iterations of a simple law (natural selection of the fittest to reproduce) can produce novel, diverse and rich structures. In biological evolution many generations in a population can produce new traits and species. 

Many of the most debated issues about evolution relate to the different characteristics of emergence and are briefly discussed below.

Scales

Central to emergence are the ideas of “many” and of scales. The former can take two forms: a system composed of many interacting components, or a system that undergoes many iterations according to a rule that is repeated many times. For evolution, both forms of “many” are relevant and have several dimensions. Evolution occurs in a population, i.e., a community of many members of a species living in a specific environment. Each member of the population has a specific genotype (many genes), which largely determines biological characteristics, from proteins to organs, defined as the phenotype. The environment also consists of many interacting species. Natural selection can act at multiple levels: on genes, cells, organisms, species, and groups of species.

Microscopic and macroscopic scales can also manifest in different ways. In terms of length, the micro- and macro- scales can be defined in terms of genotypes and phenotypes, respectively. In terms of time, microevolution and macroevolution roughly correspond to directly observable timescales and geological timescales, respectively. They are associated with the emergence of new traits within a species and new species, respectively.

Novelty

Development of new traits and species occurs over many generations, due to the repetition of the rule of natural selection.

Evolution theory uses concepts such as natural selection, survival of the fittest, niches, and hierarchical trees, that are not present in chemistry and physics. 

Connecting micro- and macro- properties

As for other systems, this is one of the great challenges of emergence. Genotypes and phenotypes are extremely well characterised. Genotype-phenotype maps seek to connect these micro- and macro- levels. A detailed understanding of how microevolution leads to macroevolution is a challenge.

Discontinuities

In microevolution, new traits occur within a species due to (continuous) adaptation to the environment. In contrast, in macroevolution, new organs and species can occur suddenly (at least on geological timescales). An example is the Cambrian explosion of new life forms. Extinctions can also represent discontinuities.

Evolution of a population occurs in response to changes in an environment. New traits, new species, and extinctions can be viewed as qualitative changes due to quantitative changes. For example, small changes in the oxygen concentration in the atmosphere is one (among many) hypotheses for the cause of the Cambrian explosion.

Using techniques from statistical physics, the transition of a species from survival to extinction can be viewed as a non-equilibrium phase transition to an absorbing state. The order parameter is the population and a toy model is directed population.1

Diversity with limitations

All species are based on the same biochemistry of DNA and proteins. Yet from these same building blocks there is an incredible diversity: more than 8 million distinct species, including more than 10,000 species of birds and more than 15,000 species of ants. Darwin said nature produces “endless forms most beautiful.”

But there are limitations. For example, the number of species with more than one head, brain, heart, or liver is limited. There are many more genotypes than phenotypes. 

The dominant view is that evolution is driven by random genetic mutations. Debates have arisen about how much evolution is limited (constrained) by morphology and environment.

Ball stated: (p. 332)

“convergent evolution is often regarded as a sign that certain shapes or structures are ideal adaptations to particular environments for physical reasons: wings consisting of flat, thin membranes are best for flying, torpedo-shaped bodies a streamlined for efficient swimming, and so on… There is a tendency in evolutionary biology to regard natural selection as a process with an infinite palette: anything is possible so long as it doesn't break the laws of physics. But the laws of physics might impose more constraint than that, precisely because biology uses rather than merely suffers them.”

Universality

Not all mutations produce a change in phenotype. There are neutral mutations. There are many more genotypes than phenotypes. In other words, genotype-phenotype maps are many-to-one.

Species that are unrelated or distantly related (in the tree of life) sometimes have traits or behaviours that are similar. Convergent evolution is the hypothesis that natural selection produced the same outcome in a different context. 

Modularity at the mesoscale

The economist Simon pointed out that evolution can occur on much faster time scales than might be expected because of modularity. According to Clune et al.

“A long-standing, open question in biology is how populations are capable of rapidly adapting to novel environments, a trait called evolvability [1]. A major contributor to evolvability is the fact that many biological entities are modular, especially the many biological processes and structures that can be modelled as networks, such as metabolic pathways, gene regulation, protein interactions and animal brains [1–7].”

Ball highlighted how domains in proteins provide functional modules that evolution uses: (pp. 174-5)

“the evolution of metazoan proteins is not so much a slow affair of letting random genetic mutations change one amino acid for another and seeing what effect it produces. Rather, it constitutes a reshuffling of already functional modules to produce multidomain molecules with new potential - a strategy much more likely to yield successful results…  the “unit” of molecular evolution here is not really the base pair of DNA or the amino acid or protein, or the gene itself, by the peers at a scale intermediate between the two: the module of a domain. It seems that this shuffling, rather than the slow mutation of primary base sequences, is what has driven the evolution of animals.”

Johnston et al. considered an algorithmic picture of evolution that 

“suggests that symmetric structures preferentially arise not just due to natural selection but also because they require less specific information to encode and are therefore much more likely to appear as phenotypic variation through random mutations… many genotype–phenotype maps are exponentially biased toward phenotypes with low descriptional complexity. A preference for symmetry is a special case of this bias… Lower descriptional complexity also correlates with higher mutational robustness, which may aid the evolution of complex modular assemblies of multiple components.”

Self-organisation

Complex biological structures, from proteins to organisms, have formed spontaneously due to evolution over millions of years. Their intricacy and functionality have led to claims of purpose and design. However, this is argued to be an “apparent” design, just like an economy whose self-organisation appears “as if” it is guided by an “invisible hand.”

Kauffman claimed that self-organisation is as important as natural selection in driving evolution.

Unpredictability

A contested question about evolution is the role of contingency (historical accidents) and whether the evolution of complex life forms, particularly humans, was an accident of history or inevitable.

Irreducibility

Until recently, evolutionary biology has been dominated by a reductionist gene-centric view, popularised by Dawkins. However, recent discussions about systems biology, evo-devo, and epigenetics have questioned this view. Some characterise these alternative views as a form of structuralism.

Complexity

An algorithmic picture of evolution suggests that simplicity spontaneously emerges as many genotype-phenotype maps may be biased towards phenotypes with low descriptional complexity. 

Toy models

An earlier post discussed the key role that toy models, such as “bean bag” genetics, have played in evolutionary theory.

Cross-fertilisation of fields

Ideas from evolution have stimulated the development of genetic algorithms in computer science.

Drossel has reviewed connections between evolution and statistical physics, including a wide range of toy models. Examples include spin glass models that give rise to rugged landscapes for fitness and can describe hierarchical structures, comparable to Darwin’s tree of life. Goldenfeld and Woese argued that evolution can be viewed as a collective phenomenon far from equilibrium. The toy model central to their discussion is directed percolation.

I welcome comments. My knowledge of biology is limited, and scientifically some the ideas above can be contentious. (Never mind philosophy, politics, or theology!)

Tuesday, February 24, 2026

Information theoretic measures for emergence and causality

The relationship between emergence and causation is contentious, with a long history. Most discussions are qualitative. Presented with a new system, how does one identify the microscopic and macroscopic scales that may be most useful for understanding and describing the system? Can Judea Pearl’s seminal ideas about causality be implemented practically for understanding emergence?

Broadly speaking, a weakness of discussions of emergence and causality is that it is hard to define these concepts in a rigorous and quantitative manner that makes them amenable to empirical testing, with respect to theoretical models and to experimental data. 

Fortunately, in the past decade, there have been some specific proposals to address this issue, mostly using information theory. A helpful recent review is by Yuan et al. 

“Two primary challenges take precedence in understanding emergence from a causal perspective. The first is establishing a quantitative definition of emergence, whereas the second involves identifying emergent behaviors or phenomena through data analysis.

To address the first challenge, two prominent quantitative theories of emergence have emerged in the past decade. The first is Erik Hoel et al.’s theory of causal emergence [19] whereas the second is Fernando E. Rosas et al.’s theory of emergence based on partial information decomposition [24].

Hoel et al.’s theory of causal emergence specifically addresses complex systems that are modeled using Markov chains. It employs the concept of effective information (EI) to quantify the extent of causal influence within Markov chains and enables comparisons of EI values across different scales [19,25]. Causal emergence is defined by the difference in the EI values between the macro-level and micro-level."

One perspective on causal emergence is that it occurs when the dynamics of a system at the macro-level is described more efficiently by macro-variables than by the dynamics of variables from the micro-level.

Klein et al. used Hoel’s information-theoretic measures of causal emergence to analyse protein interaction networks (interactomes) in over 1800 species, containing more than eight million protein–protein interactions, across different scales. They showed the emergence of ‘macroscales’ that are associated with lower noise and uncertainty. The nodes in the macroscale description of the network are more resilient than those in less coarse-grained descriptions. Greater causal emergence (i.e., a stronger macroscale description) was generally seen in multicellular organisms compared to single-cell organisms. The authors quantified causal emergence in terms of mutual information (between large and small scales) and effective information (a measure of the certainty in the connectivity of a network). Philip Ball (2023) (pages 218-220) gives an account of this work in terms of the emergence of multicellularity in biological evolution. He introduced the term causal spreading (pages 225-7), arguing that over the history of evolution the locus of causation has changed.

Yuan et al. continue

"However, in Hoel’s theory of causal emergence, it is essential to establish a coarse-graining strategy beforehand. Alternatively, the strategy can be derived by maximizing the effective information (EI) [19]. However, this task becomes challenging for large-scale systems due to the computational complexity involved. To address these problems, Rosas et al. introduced a new quantitative definition of causal emergence [24] that does not depend on coarse-graining methods, drawing from partial information decomposition (PID)-related theory. PID is an approach developed by Williams et al., which seeks to decompose the mutual information between a target and source variables into non-overlapping information atoms: unique, redundant, and synergistic information [29]…"

The Figure below is taken from Rosas et al. Xt^j (j=1,…,n) are microscopic variables that define a Markov chain. Vt is a macroscopic variable that is completely determined by the microscopic variables.

“Diagram of causally emergent relationships. Causally emergent features have predictive power beyond individual components. Downward causation takes place when that predictive power refers to individual elements; causal decoupling when it refers to itself or other high-order features.”

Rosas et al. applied the method to specific systems, including Conway’s Game of Life, Reynolds’ flocking model, and neural activity as measured by electrocorticography. More recently, it was used to describe emergence in computer science, including the identification of modular structures. Calculations were performed for specific examples, including Ehrenfest’s urn model for diffusion, the Ising model with Glauber dynamics, a Hopfield neural network model for associative memory.

Yuan et al. also state the following:

"The second challenge pertains to the identification of emergence from data. In an effort to address this issue, Rosas et al. derived a numerical method [24]. However, it is important to acknowledge that this method offers only a sufficient condition for emergence and is an approximate approach. Another limitation is that a coarse-grained macro-state variable should be given beforehand to apply this method."

Sas et al. recently stated

“Empirical applications of this framework to study emergence … including the study of gene regulatory networks [22], the dynamics of the human brain [23], the internal dynamics of reservoir computing [24], and the formation of useful internal representations in machine learning [25].”

Yuan et al. also discuss two significant connections between causal emergence and machine learning. First, machine learning can be used to improve calculations of causal emergence. Second, causal emergence measures can be used to better understand how machine learning works and improve it.

The work described above built on earlier work by Crutchfield, who claimed that the identification of emergence and hierarchies could be made operational, stating that “different scales are delineated by a succession of divergences in statistical complexity at lower levels.” More recently, Rupe and Crutchfield have reported progress towards identifying emergent self-organisation in a system.

Although this work on quantitative measures of emergence based on information theory represents significant progress, there are many open problems. Examples include the extension to non-Markovian systems and the development of computationally feasible methods for large systems. The latter is particularly important in physical systems where spontaneous symmetry breaking occurs, as this only happens in the thermodynamic limit of an infinite system.

There is an unrecognised similarity between the work described above and techniques recently developed to characterise phase transitions in statistical mechanics models such as the Ising model and classical dimer models. Coarse-graining (CG) is optimised by maximising the Real-Space Mutual Information (RSMI) between a spatial block and its distant environment. 

In general, maximising mutual information is notoriously hard but can be done using state-of-the-art machine learning algorithms. Gokmen et al. have developed an algorithm that they claim “can, unsupervised, construct order parameters, locate phase transitions, and identify spatial correlations and symmetries for complex and large-dimensional real-space data.” Furthermore, the optimal CG explicitly identifies the scaling operators associated with the critical point. 

The classical dimer model provides a stringent test as “the relevant low-energy degrees of freedom are profoundly different from the microscopic building blocks of the theory and change qualitatively throughout the phase diagram.” In other words, the emergent entities (quasiparticles such as vortices associated with the height field, which is described by a sine-Gordon field theory) are different from the dimers.

It is encouraging to see that two different scientific communities have developed similar ideas to address this challenging problem of making discussions about emergence and causality more concrete and quantitative.

Friday, January 16, 2026

Responding to scientific uncertainty

Science provides an impressive path to certainty in some areas, particularly in physics. However, as scientists seek to describe increasingly complex entities, moving from chemistry to biology, and then to humans and societies, the level of uncertainty increases.

One observes a wide range of responses to scientific knowledge being uncertain. Here are a few.

Denial. Science is about facts and absolute truth. There really isn’t a problem. We should just trust the scientists.

Minimisation. There is some uncertainty, but it isn’t anything to be concerned about. Some scientists will also minimise any uncertainty about their own research. This may occur because of career ambition. Others will minimise public discussion of uncertainty to try and avoid promoting the science scepticism discussed below.

Optimistic perseverance. The uncertainty is openly acknowledged. Some of the uncertainty does not matter for what we need to know. Other uncertainties can be reduced by further scientific work, such by more precise measurements with new instruments or by developing more sophisticated theories.

Total scepticism. There is a suspicion about the validity of most scientific knowledge, particularly that which is perceived to have philosophical, religious, or political implications.

Suspicion about science

In spite of the success of science at describing the material world and leading to powerful and useful technologies, there is much public suspicion of science. On the one hand, this is understandable given that science has led to technologies with undesirable health, environmental and social consequences. Some scientists, governments and companies have lied about these consequences and hidden them from the public. Human subjects have been abused in medical experiments. Drugs that were claimed to be effective and safe turned out to be ineffective or have undesirable side effects. Science has been used for ideological purposes. Sometimes scientists have faked results to advance their own careers. However, these failures should not undermine our trust in reliable scientific knowledge. Distinctions should be made between the bodies of knowledge, the applications of that knowledge, and the actions of institutions. I now discuss several common claims in public discussion that are used to justify scepticism of scientific knowledge.

Science is always changing. 

One day, scientists tell you that chocolate is good for your health, and the next year they say it is bad for you. And that is just the start. Then there are eggs, wine, running marathons, and cheese. They just can’t make up their mind. So why should we trust them? At one time, they believed in phlogiston and the aether. Now they say they don’t exist. Aristotle was replaced by Newton, who was replaced by Einstein. So why believe in human-induced climate change, biological evolution, vaccines, the Big Bang theory, or Einstein’s theories?

It is true that scientific knowledge does develop and change over time. However, today we have incredibly detailed observations and theories in physics, astronomy, chemistry, biology, and geology. Any future changes will be relatively minor because they will have to be consistent with all the knowledge we have now. Furthermore, when theories change, such as when Einstein superseded Newton, they don’t show that the old theory was completely wrong, but rather that it applied in a limited domain. For example, Newton’s theories of motion and gravity are extremely reliable when it comes to objects that are much larger than atoms, less dense than a black hole, and are moving at speeds less than about 10,000 kilometres per second. This is why engineers spend years learning Newton’s theories, not Einstein’s. If you want to build a good bridge or a rocket, Newton is good enough. He is not wrong.

Update. (Jan. 19). I just discovered that the NY Times had a recent op-ed Science Keeps Changing. So Why Should We Trust It?

“Well, that’s just a theory.” 

In popular debate, such a refrain may be applied to the theory of biological evolution, the Big Bang theory in cosmology, or human-induced climate change. The claimant usually wants to dismiss a particular theory as just idle speculation. Here, the term “theory” is used in the same sense as everyday speculations, such as “I have a theory as to why the president resigned,” or “I have a theory about why my computer is running so slowly.” These are just stories that sound somewhat plausible. In contrast, scientific theories in physics, such as quantum theory and Einstein’s theories of relativity, have precisely defined mathematical formulations that have been checked for logical consistency, made specific predictions, and tested to great precision in experiments. They are not “just theories.” For example, for the Big Bang theory about the beginning of the universe and Darwin’s theory of biological evolution and diversity, there are many independent lines of evidence that are consistent with each theory.  

Scientists cannot be trusted. 

They are not committed to the truth, but rather to their own interests and agendas, related to their careers, politics, and religion. They close ranks and support the status quo of current scientific “dogma”, rather than being open to original thinkers who critique it and propose alternative theories. They don’t want to lose their well-paid jobs and lucrative grants. 

On the one hand, scientists can be conservative and resistant to new ideas. On the other hand, there are significant career incentives to overturn existing knowledge and have your radical new theory accepted. That is how some scientists become famous and win Nobel Prizes. The reasons it does not happen very often are not necessarily for social or ideological reasons. Many of the theories we have today can explain an awful lot. It requires a lot of evidence, carefully acquired and checked, to convince people that those theories need to be modified, let alone abandoned. This may take decades. But it does happen. An example is the Big Bang theory of the universe, whose acceptance was initially resisted because it went against the prevailing view that the universe did not have a beginning. In biology, the discovery in 1970 of the enzyme reverse transcriptase went against a popular version of the “Central dogma” of molecular biology that DNA was always converted to RNA and not the reverse. That discovery led to a Nobel Prize.

I don’t trust scientists. I will do my own research. There is lots of good material from unbiased sources on the internet.

The internet provides a range of information and perspectives on practically any issue imaginable, including science. The material is particularly vast and controversial on biological evolution, the beginning of the universe, fundamental physics, the age of the earth, climate change, and medicine. Since the covid-19 pandemic, scepticism of the effectiveness and safety of vaccines has increased. 

Ivermectin is a drug that was developed as a treatment for parasite worms. Its incredible success was recognised by the award of the 2015 Nobel Prize in Physiology or Medicine to William Campbell and Satoshi Omura, who discovered the drug. During the pandemic, high-profile politicians and social media influencers promoted ivermectin as a treatment for covid-19, even after systematic medical studies showed it was ineffective. Recently, it has gained a reputation as a “miracle” drug that can even cure cancer, but this is being suppressed by the medical establishment. All clinical trials have shown the drug is ineffective for human ailments, beyond deworming. Nevertheless, there are groups on social media with hundreds of thousands of members that discuss the conspiracy, how to get the drug, and the experiences of participants using it to treat a wide range of ailments. Danny Lemoi, a founder of one of the largest groups, died in 2023 after taking massive daily doses of the drug for several years to treat a heart condition. Afterwards, one member of the group wrote “No one can convince me that he died because of ivermectin. He ultimately died because of our failed western medicine which only cares about profits and not the cure.”

Fans of ivermectin claim that they are escaping the biases and vested interests of the medical establishment and Big Pharma as they pursue the truth. However, they are not escaping bias and vested interests. Successful social influencers build their reputations and million-dollar incomes from promoting scepticism. If there is no conspiracy, just scientific uncertainty and occasional incompetence and malpractice, their following collapses. Populist politicians build their careers on criticism of and stoking resentment towards elites, such as the medical establishment. The authority of the medical establishment is replaced with the authority of the popular opinion of a group of people whose views are shaped by social media algorithms, intuition, and anecdotal experience.

My purpose in giving the example of Ivermectin is not to start a detailed critique of science scepticism. Rather, it is to illustrate the role that the interplay of trust, authority, and tradition plays in how we determine what is true and what to act on. There are two competing traditions here: the populism of alternative medicine and the elitism of professional medicine. Each has its own sources of authority. In the end, it boils down to who we trust. We do not have the time, energy, resources or inclination to check the veracity of every single piece of information we have access to. We take shortcuts. This is what tradition does for us, for better and worse. Thus, we cannot escape tradition. We are all swimming in traditions, many of which are in conflict with one another. The question is whether we are aware of it and what we do with that awareness.

Saturday, December 23, 2023

Niels Bohr on emergence

Until this week, I did not know that Bohr ever thought about emergence.

Ernst Mayr was one of the leading evolutionary biologists in the twentieth century and was influential in the development of the modern philosophy of biology. He particularly emphasised the importance of emergence and the limitations of reductionism. In the preface to his 1997 book, This is Biology: the Science of the Living World, Mayr recounts the development of his thinking about emergence.

At first I thought that this phenomenon of emergence, as it is now called, was restricted to the living world, and indeed, in a lecture I gave in the early 1950s in Copenhagen, I made the claim that emergence was the one of the diagnostic features of the of the organic world. The whole concept of emergence at the time was considered to be rather metaphysical. When Niels Bohr who was who was in the audience, stood up during the discussion, I was fully prepared for an annihilating refutation. However, much to my surprise, he did not at all object to the concept of emergence, but only to my notion that it provided a demarcation between the physical and the biological sciences. Citing the case of water whose "aquosity" could not be predicted from the characteristics of its two components, hydrogen and oxygen, Bohr stated that emergence is rampant in the inaminate world.  (page xii).

Later in the book Mayr pillars Bohr for his support of vitalism, including claims that vitalism has a "quantum" foundation.

Thursday, November 2, 2023

Diversity is a common characteristic of emergent properties

Consider a system composed of many interacting parts. I take the defining characteristic of an emergent property is novelty. That is, the whole has a property not possessed by the parts alone. I argue that there are five other characteristics of emergent properties. These characteristics are common but they are neither necessary nor sufficient for novelty.

1. Discontinuities

2. Unpredictability

3. Universality

4. Irreducibility

5. Modification of parts and their relations

I now add another characteristic.

6. Diversity

Although a system may be composed of only a small number of different components and interactions, the large number of possible emergent states that the system can take is amazing. Every snowflake is different. Water is found in 18 distinct solid states. All proteins are composed of linear chains of 20 different amino acids. Yet in the human body there are more than 100,000 different proteins and all perform specific biochemical functions. We encounter an incredible diversity of human personalities, cultures, and languages. 

A related idea is that "simple models can describe complex behaviour". Here "complex" is often taken to mean diverse. Examples, how simple Ising models with a few competing interactions can describe a devil's staircase of states or the multitude of atomic orderings found in binary alloys.

Perhaps the most stunning case of diversity is life on earth. Billions of different plant and animal species are all an expression of different linear combinations of the four base pairs of DNA: A, G, T, and C.

One might argue that this diversity is just a result of combinatorics. For example, if one considers a chain of just ten amino acids there are 10^13 different possible linear sequences. But this does not mean that all these sequences will produce a functional protein, i.e., one that will fold rapidly (one the timescale of milliseconds) into a stable tertiary structure, and one that can perform a useful biochemical function. 

Friday, January 27, 2023

Science and the universe are awesome

Since we are surrounded by scientific knowledge. We are so used to it that we can take science for granted and not reflect on how amazing science truly is. And how amazing the universe is that science reveals. Things that we know, learn, and do today in science would have been inconceivable decades ago, let alone centuries ago.

What specific things do you think are particularly awesome? This question was stimulated by Frank Wilczek's recent book, Fundamentals: Ten Keys to Reality. In writing the book, he says "what began as an exposition grew into a contemplation."

 My answer to the question has some significant overlap with Wilczek's ten. 

Below I list some of the things that I find awesome. I consider two classes: what science can do and what we learn about the universe from science.

Science works! It is amazing what science can do.

We can understand the material world.

Einstein said, "The most incomprehensible thing about the world is that it is comprehensible." In a previous post, I explored some different dimensions of the fact that the universe is comprehensible. The mystery includes human capabilities, both intellectual and physical, and the malleability of the material world.

We can make precise measurements.

Scientists have created incredibly powerful and specialised instruments for making very precise measurements such as spectrometers, telescopes and microscopes. Scientists can measure the tension in a single strand of DNA, the magnetic moment of an electron to a precision of one part in one billion billion, the spectrum of light emitted by a galaxy that is ten billion light years away, ...

We can predict the outcome of new experiments.

Scientists construct theories in their minds, on pieces of paper, in mathematical equations, and in computers. One way to evaluate the possible validity of a theory is to propose new experiments and predict the outcome. Famous examples include the existence of the chemical element aluminium, the existence of the planet Neptune, radio waves, a specific excited quantum state of the atomic nucleus of carbon atoms, the pollinator moth for Darwin's orchid, the deflection of the path of light from a distant star by our sun, gravitational waves, the Cosmic Microwave Background, quarks, the Higgs boson, the Berezinskii-Kosterlitz-Thouless phase transition, the hexatic phase, edge states in integer spin antiferromagnetic chains, topological insulators, ... Predictions are particularly impressive when they are unexpected and controversial.

We can use mathematics. 

Eugene Wigner received the Nobel Prize in Physics in 1963. In 1960 he published an essay "The Unreasonable Effectiveness of Mathematics in the Natural Sciences that concludes

The miracle of the appropriateness of the language of mathematics for the formulation of the laws of physics is a wonderful gift which we neither understand nor deserve. 

We can manipulate and control nature.

Scientists and engineers can move single atoms, design drugs, make computers, build atom bombs, heart pacemakers, and mobile phones, manipulate genes, ......

We know so much but we know so little. 

On the one hand, the achievements of science are amazing. Yet, in spite of this, there are still significant mysteries and challenges. Examples include the nature of dark matter or human consciousness, a quantum theory of gravity, fine-tuning of fundamental constants, the quantum-classical boundary, protein folding, the nature of glasses, and how to calculate the properties of complex systems.

It is awesome what science reveals to us about the universe.

The immense scales of the observable universe

Our sun is just one star among the more than two hundred billion that make up our galaxy, the Milky Way. And that is just one of one trillion galaxies in the whole universe. It takes light from the most distant galaxies tens of billions of years to travel to us.

Length, time, and energy scales over many many orders of magnitude

These go far beyond our everyday experience and what we can see with the naked eye (from a millimetre to a kilometre). On the large scale, the visible universe involves distances of billions of light years (10^25 metres). On the small scale, there is the sub-structure of nucleons, which is smaller than femtometres (10^-15 m).  This wide range of length scales is nicely illustrated in the wonderful movie Powers of Ten and its update, The Cosmic Eye. There are corresponding time, energy, and temperature scales varying over many many orders of magnitude. For example, as one goes from ultracold atomic gases to quark-gluon plasmas, the  relevant energy and temperature scales vary over more than 20 orders of magnitude! At every scale, there are distinct phenomena and structures. 

Universal laws that are simple to state

The universe exhibits a diversity of rich and complex behaviour. Yet it can understand much of it in terms of simple universal laws that are easy to state, e.g., Newton's laws of motion, the laws of thermodynamics, Maxwell's equations of electromagnetism, Schrodinger's equation of quantum mechanics, the genetic code, ....  And, these are just a few of these laws. One does not need a multitude of laws to describe a multitude of instances of a multitude of phenomena.

Just a few building blocks

There are just a few fundamental particles in the standard model (leptons, neutrinos, and gauge bosons). Everything is made of them. They are the building blocks of atoms. They each have just a few physical properties: charge, spin, mass, and colour. Every single particle of a particular type in the universe has exactly the same properties. Exactly. As far as we know, they have been exactly the same throughout time, going back to the beginning of the universe, and whether they are in your body, or in a star in a distant galaxy.

Atoms are the building blocks of chemical compounds. Every single atom of a particular chemical element (and nuclear isotope) is absolutely identical. This allows astronomers to determine the chemical composition of distant stars, galaxies, and dust clouds.

Humans, plants, and animals all have the same molecular building blocks and there are just a few of them. Any DNA molecule is composed of just four different base pairs (denoted A, G, T, C) and proteins are composed of just twenty different amino acids.

There are two amazing things here. First, there are just so few building blocks. Second, every one of these building blocks is absolutely identical.

Emergence: simple rules produce complex behaviour

Humans, cells, and crystals can be viewed as systems composed of many interacting components. The components and their interactions can often be understood and described in simple terms. Nevertheless, from these interactions complex structures and properties can emerge.

Nature appears to be fine-tuned for life

This covers not just the values of fundamental physical constants that lead to the notion of fine-tuning and the anthropic principle. Water has unique physical and chemical properties that allow it to play a crucial role in life, such as the surface of lakes freezing before the bottom and aiding protein folding.

The intricate and subtle "machinery" of biomolecules

Proteins have very unique structures that are intimately connected to their specific functions, whether as catalysts or light sensors.


What do you think are the most amazing things about science and what we learn from it?


Friday, September 2, 2022

The value of "simple" models for complex systems

Significant understanding of emergent phenomena in quantum materials has come from the study of model Hamiltonians such as those associated with the names Hubbard, Anderson, Kondo, Heisenberg, Kitaev, Haldane, BCS,...

I had not appreciated until recently that an early key to the Modern Synthesis of evolutionary biology (that brought together Darwinian natural selection with Mendelian genetics) was the development of simple mathematical models. The discussion below is taken from

Towards a unified science of cultural evolution 
Alex Mesoudi, Andrew Whiten and Kevin N. Laland 
Significant advances were made in the study of biological [micro]evolution before its molecular basis was understood, in no small part through the use of simplified mathematical models, pioneered by Fisher (1930), Wright (1931), and J.B.S. Haldane (1932)... 
Mathematical models such as [those for cultural evolution and gene-culture coevolution] are often treated with suspicion and even hostility by some social scientists, who consider them to be oversimplifications of reality... The alternatives..., however, are usually either analysis at a single (purely genetic or purely cultural) level or vague verbal accounts of “complex interactions,” neither of which we believe to be productive. Gene-culture analyses have repeatedly revealed circumstances under which the interactions between genetic and cultural processes lead populations to different equilibria than those predicted by single level models or anticipated in verbal accounts... as illustrated by the aforementioned examples of dairy farming and handedness.  
Interestingly, fifty years ago the same reservations about simplifying assumptions were voiced about the use of population genetic models in biology by the prominent evolutionary biologist Ernst Mayr (1963). He argued that using such models was akin to treating genetics as pulling coloured beans from a bag (coining the phrase “beanbag genetics”), ignoring complex physiological and developmental processes that lead to interactions between genes. 
 

In his classic article “A Defense ofBeanbag Genetics,” J. B. S. Haldane (1964) countered that the simplification of reality embodied in these models is the very reason for their usefulness. Such simplification can significantly aid our understanding of processes that are too complex to be considered through verbal arguments alone, because mathematical models force their authors to specify explicitly and exactly all of their assumptions, to focus on major factors, and to generate logically sound conclusions. Indeed, such conclusions are often counterintuitive to human minds relying solely on informal verbal reasoning. 

Haldane (1964) provided several examples in which empirical facts follow the predictions of population genetic models in spite of their simplifying assumptions, and noted that models can often highlight the kind of data that need to be collected to evaluate a particular theory. Ultimately, Haldane won the argument, and population genetic modelling is now an established and invaluable tool in evolutionary biology (Crow 2001). We can only echo Haldane’s defence and argue that the same arguments apply to the use of similar mathematical models in the social sciences.

A more recent version of J.B. S. Haldane's argument is Not Just a Theory—The Utility of Mathematical Models in Evolutionary Biology Maria R. Servedio,Yaniv Brandvain, Sumit Dhole, Courtney L. Fitzpatrick, Emma E. Goldberg, Caitlin A. Stern, Jeremy Van Cleve, D. Justin Yeh 


All models are wrong but some are useful. I first learnt this aphorism from Scott Page, in his wonderful course Model Thinking at Coursera.  This short talk discusses how models help us think more clearly. Simple quantitative models, such as agent-based models, in the social sciences, have the value that their assumptions can be clearly stated, and then the consequences of these assumptions can be investigated in a rigorous manner.

There is also a nice discussion of the importance of model building for science in John Holland's beautiful book, Emergence. 

Monday, August 22, 2022

Hysteresis, hype, niches, nudges and social change

The world is a mess. Most people want a better world. Sometimes nothing changes. Sometimes things change incredibly rapidly. Sometimes changes are positive. Other times the change is negative. Often this change is unanticipated, even by experts who have been studying the relevant topic for decades. Wicked problems are things that seem to be incredibly resilient to change. Examples of rapid changes that were (largely) positive and unanticipated were the peaceful collapse of the former Soviet empire, smoking in public becoming taboo, and increased public concern about climate change. Examples of negative changes include the rise of Trumpism, misinformation on social media, and the global financial crisis of 2008.

Many people in government, public policy, NGOs, and social activists want to implement policies and take actions that will produce outcomes that (they believe) are positive. Here I discuss some basic but very important insights from "social physics", such as discussed in my previous two posts.

Suppose the system of interest can be modelled by some type of Ising model where the pseudospin corresponds to two choices (good and bad) for each agent in the system. The policy maker wants to change something such as increase the incentive for agents to make the "good" choice. There are two qualitatively different possible behaviours and they are shown in the Figure below (taken from Bouchaud). 

The vertical axis is the "magnetisation", i.e, the fraction of agents who make the good choice. The horizontal axis is the "external field", i.e, the level of incentive provided for agents to make the good choice. 


Case I. Smooth curve (blue). This occurs when the interaction between agents is weaker than some threshold strength. Suppose that a small but not insignificant minority of agents are already making the good choice and then incentive is increased slightly. If one is near the steep part of the blue curve then this "nudge" can produce a desired outcome for the society.

Case II. Discontinuous curve (red). This occurs when the interaction between agents is greater than some threshold strength. People's choices are influenced more by their friends than by what the government or an NGO is telling them to do. Then one has to provided very large incentives to get a change in agent choice, far beyond the incentive required for a single isolated agent. The system is stuck in a state that is not good for the society as a whole. It is a metastable state, as shown in the figure below.

On the other hand, if the "polarisation field" is sitting near a critical value (5 in the figure, a tipping point), then a "nudge" can lead to a dramatic change for good. 

I think there are important implications for social activists of all stripes. Realistic expectations are key.

1. Don't expect even the best-designed and well-intentioned policy or action to necessarily have the impact you hope for.

2. Be sceptical about hype and ideology. In the public space there are a lot of claims, whether from political parties, pundits, or NGOs, that if we just do X (change this law, donate money, do what my book says, ...) then the good Y will inevitably follow.

The problem with unrealistic expectations is that they lead to disappointment, disillusionment, and burnout. People give up. Then the next fad or "silver bullet" comes along...

Inspired by a rugged landscape perspective, a better and more sustainable approach is that of learning and adaptation. One identifies what one thinks the best "nudge" is, tries something, evaluates the effect, adapts, and tries out some new ideas. One does not claim or expect the first few iterations to produce a significant desired effect. Here, somewhat "random" sampling of the landscape may help. Here a diversity of perspectives and methods can play a positive role. A more concrete version of this argument is in a paper concerned with public health initiatives. Rugged landscapes: complexity and implementation science, by Joseph T. Ornstein, Ross A. Hammond, Margaret Padek, Stephanie Mazzucca & Ross C. Brownson 

Postscript. After posting this I remember reading a recent article in The Economist pointing out how nudges often do not work.

Evidence for behavioural interventions looks increasingly shaky 
The academic literature is plagued by publication bias 

It references three recent Letters in PNAS, including this one, that come to the opposite conclusion to an earlier PNAS paper.
Stephanie Mertens, Mario Herberz, Ulf J. J. Hahnel, and Tobias Brosch

Friday, July 29, 2022

Famous last words

If you ever write a popular book about science I suggest you spend a lot of time honing your very last paragraph. If it is eloquent, grand, and hyperbolic it may be so widely quoted that many people will think that this is actually what the book is about or has proven. Here are a few examples that I often see.
Where then shall we find the source of truth and the moral inspiration for a really scientific socialist humanism? Only, we suggest, in the sources of science itself,..... it is the conclusion to which the search for authenticity necessarily leads. The ancient covenant is in pieces; man at last knows that he is alone in the unfeeling immmensity of the universe, out of which he emerged only by chance. Neither his destiny nor his duty have been written down. The kingdom above or the darkness below: it is for him to choose.''
Jacques MonodChance and Necessity: An Essay on the Natural Philosophy of Modem Biology, trans. Austryn Wainhouse (New York: Knopf, 1971), p. 167
But if there is no solace in the fruits of our research, there is at least some consolation in the research itself. Men and women are not content to comfort themselves with tales of gods and giants, or to confine their thoughts to the daily affairs of life; they also build telescopes and satellites and accelerators, and sit at their desks for endless hours working out the meaning of the data they gather. The effort to understand the universe is one of the very few things which lifts human life a little above the level of farce and gives it some of the grace of tragedy.
Steven Weinberg, The First Three Minutes (Basic Books, 1977), pages 154-155.
If we do discover a complete theory, it should in time be understandable in broad principle by everyone, .... Then we shall all ...[discuss] why it is that we and the universe exist. If we find the answer to that, it would be the ultimate triumph of human reason - for then we would truly know the mind of God.  
Stephen Hawking, A Brief History of Time
There is grandeur in this view of life, with its several powers, having been originally breathed by the Creator into a few forms or into one; and that, whilst this planet has gone circling on according to the fixed law of gravity, from so simple a beginning endless forms most beautiful and most wonderful have been, and are being evolved.
Charles Darwin, The Origin of Species

Can you think of any other examples of famous last paragraphs?

Friday, April 8, 2022

Why is there so much symmetry in biological systems?

 One of the biggest questions in biology is, What is the relationship between genotypes and phenotypes? In different words, how does a specific gene (DNA sequence) encode information that allows a very specific biological structure with a unique function to emerge?

Like big questions in many fields, this is a question about emergence.

In biology, this mapping from genotype to phenotype occurs at many levels from protein structure to human personality. An example is how the RNA encodes the structure of a SARS-CoV2 virion.

A fascinating thing about biological structures is that many have a certain amount of symmetry. The human body has reflection symmetry and many virions have icosahedral symmetry. What is the origin of this tendency to symmetry? Could evolution produce it?

Scientists will sometimes make statements such as the following about evolution.

Symmetric structures preferentially arise not just due to natural selection but also because they require less specific information to encode and are therefore much more likely to appear as phenotypic variation through random mutations.

How do we know this is true? Can such a statement be falsified? Or at least, can we produce concrete models or biological systems that are consistent with this statement?

There is a fascinating paper in PNAS that addresses the questions above.

Symmetry and simplicity spontaneously emerge from the algorithmic nature of evolution 
Iain G. Johnston, Kamaludin Dingle, Sam F. Greenbury, Chico Q. Camargo, Jonathan P. K. Doye, Sebastian E. Ahnert, and Ard A. Louis 

Here are a few highlights from the article. First, how one gets specific about information content and algorithms.
Genetic mutations are random in the sense that they occur independently of the phenotypic variation they produce. This does not, however, mean that the probability P(p) that a Genotype-Phenotype [GP] map produces a phenotype p upon random sampling of genotypes will be anything like a uniformly random distribution. 
Instead, ... arguments based on the coding theorem of algorithmic information theory (AIT) (7) predict that the P(p) of many GP maps should be highly biased toward phenotypes with low Kolmogorov complexity K(p) (8). 
High symmetry can, in turn, be linked to low K(p) (6911). An intuitive explanation for this algorithmic bias toward symmetry proceeds in two steps: 
1) Symmetric phenotypes typically need less information to encode algorithmically, due to repetition of subunits. This higher compressibility reduces constraints on genotypes, implying that more genotypes will map to simpler, more symmetric phenotypes than to more complex asymmetric ones (23). 
2) Upon random mutations these symmetric phenotypes are much more likely to arise as potential variation (1213), so that a strong bias toward symmetry may emerge even without natural selection for symmetry.
The authors consider several concrete models and biological systems that illustrate this bias toward symmetry. The first involves the structure of protein complexes, as given in the Protein Data Base (PDB).


A) Protein complexes self-assemble from individual units. 

(B) Frequency of 6-mer protein complex topologies found in the PDB versus the number of interface types, a measure of complexity 
K˜(p). 
Symmetry groups are in standard Schoenflies notation: C6D3C3C2, and C1. There is a strong preference for low-complexity/high-symmetry structures. 

(C) Histograms of scaled frequencies of symmetries for 6-mer topologies found in the PDB (dark red) versus the frequencies by symmetry of the morphospace of all possible 6-mers illustrate that symmetric structures are hugely overrepresented in the PDB database. 

Note the logarithmic scales for the probabilities (frequencies), meaning that the probabilities span four orders of magnitude. The authors claim that "many genotype–phenotype maps are exponentially biased toward phenotypes with low descriptional complexity. "
This intuition that simpler outputs are more likely to appear upon random inputs into a computer programming language can be precisely quantified in the field of AIT (7), where the Kolmogorov complexity K(p) of a string p is formally defined as a shortest program that generates p on a suitably chosen universal Turing machine (UTM). 

From AIT the authors produce a bound (equation 1, and below), that exhibits the exponential decay of probability with complexity, similar to that seen in their graphs, such as the one shown below, for a model gene regulatory network that is modeled by 60 ordinary differential equations (ODEs). The red dashed line is the bound below.

𝑃(𝑝)2𝑎𝐾˜(𝑝)𝑏,  [1


Scaled frequency vs. complexity for the budding yeast ODE cell cycle model (30). Phenotypes are grouped by complexity of the time output of the key CLB2/SIC1 complex concentration. Higher frequency means a larger fraction of parameters generate this time curve. The red circle denotes the wild-type phenotype, which is one of the simplest and most likely phenotypes to appear. The dashed line shows a possible upper bound from Eq. 1. There is a clear bias toward low-complexity outputs.

One minor comment is that I was surprised that the authors did not reference the classic 1956 paper by Crick and Watson. They introduced the concept of "genetic economy". Prior to any knowledge of the actual structure of virions, they predicted that virions would have icosahedral symmetry because that reduced the cost of the genome coding for the structure of the virion.

Hence, it would be interesting to explore the relationship between the PNAS paper and this one.
There is a nice New York Times article about the PNAS paper. I thank Sophie van Houtryve for bringing that to my attention leading me to the PNAS paper.

What does this movie tell us about the modern university?

Last night, my wife and I watched the movie, Wit. You can watch the full movie here  (free with ads). I should warn that some of the conten...