Showing posts with label entropy. Show all posts
Showing posts with label entropy. Show all posts

Tuesday, February 24, 2026

Information theoretic measures for emergence and causality

The relationship between emergence and causation is contentious, with a long history. Most discussions are qualitative. Presented with a new system, how does one identify the microscopic and macroscopic scales that may be most useful for understanding and describing the system? Can Judea Pearl’s seminal ideas about causality be implemented practically for understanding emergence?

Broadly speaking, a weakness of discussions of emergence and causality is that it is hard to define these concepts in a rigorous and quantitative manner that makes them amenable to empirical testing, with respect to theoretical models and to experimental data. 

Fortunately, in the past decade, there have been some specific proposals to address this issue, mostly using information theory. A helpful recent review is by Yuan et al. 

“Two primary challenges take precedence in understanding emergence from a causal perspective. The first is establishing a quantitative definition of emergence, whereas the second involves identifying emergent behaviors or phenomena through data analysis.

To address the first challenge, two prominent quantitative theories of emergence have emerged in the past decade. The first is Erik Hoel et al.’s theory of causal emergence [19] whereas the second is Fernando E. Rosas et al.’s theory of emergence based on partial information decomposition [24].

Hoel et al.’s theory of causal emergence specifically addresses complex systems that are modeled using Markov chains. It employs the concept of effective information (EI) to quantify the extent of causal influence within Markov chains and enables comparisons of EI values across different scales [19,25]. Causal emergence is defined by the difference in the EI values between the macro-level and micro-level."

One perspective on causal emergence is that it occurs when the dynamics of a system at the macro-level is described more efficiently by macro-variables than by the dynamics of variables from the micro-level.

Klein et al. used Hoel’s information-theoretic measures of causal emergence to analyse protein interaction networks (interactomes) in over 1800 species, containing more than eight million protein–protein interactions, across different scales. They showed the emergence of ‘macroscales’ that are associated with lower noise and uncertainty. The nodes in the macroscale description of the network are more resilient than those in less coarse-grained descriptions. Greater causal emergence (i.e., a stronger macroscale description) was generally seen in multicellular organisms compared to single-cell organisms. The authors quantified causal emergence in terms of mutual information (between large and small scales) and effective information (a measure of the certainty in the connectivity of a network). Philip Ball (2023) (pages 218-220) gives an account of this work in terms of the emergence of multicellularity in biological evolution. He introduced the term causal spreading (pages 225-7), arguing that over the history of evolution the locus of causation has changed.

Yuan et al. continue

"However, in Hoel’s theory of causal emergence, it is essential to establish a coarse-graining strategy beforehand. Alternatively, the strategy can be derived by maximizing the effective information (EI) [19]. However, this task becomes challenging for large-scale systems due to the computational complexity involved. To address these problems, Rosas et al. introduced a new quantitative definition of causal emergence [24] that does not depend on coarse-graining methods, drawing from partial information decomposition (PID)-related theory. PID is an approach developed by Williams et al., which seeks to decompose the mutual information between a target and source variables into non-overlapping information atoms: unique, redundant, and synergistic information [29]…"

The Figure below is taken from Rosas et al. Xt^j (j=1,…,n) are microscopic variables that define a Markov chain. Vt is a macroscopic variable that is completely determined by the microscopic variables.

“Diagram of causally emergent relationships. Causally emergent features have predictive power beyond individual components. Downward causation takes place when that predictive power refers to individual elements; causal decoupling when it refers to itself or other high-order features.”

Rosas et al. applied the method to specific systems, including Conway’s Game of Life, Reynolds’ flocking model, and neural activity as measured by electrocorticography. More recently, it was used to describe emergence in computer science, including the identification of modular structures. Calculations were performed for specific examples, including Ehrenfest’s urn model for diffusion, the Ising model with Glauber dynamics, a Hopfield neural network model for associative memory.

Yuan et al. also state the following:

"The second challenge pertains to the identification of emergence from data. In an effort to address this issue, Rosas et al. derived a numerical method [24]. However, it is important to acknowledge that this method offers only a sufficient condition for emergence and is an approximate approach. Another limitation is that a coarse-grained macro-state variable should be given beforehand to apply this method."

Sas et al. recently stated

“Empirical applications of this framework to study emergence … including the study of gene regulatory networks [22], the dynamics of the human brain [23], the internal dynamics of reservoir computing [24], and the formation of useful internal representations in machine learning [25].”

Yuan et al. also discuss two significant connections between causal emergence and machine learning. First, machine learning can be used to improve calculations of causal emergence. Second, causal emergence measures can be used to better understand how machine learning works and improve it.

The work described above built on earlier work by Crutchfield, who claimed that the identification of emergence and hierarchies could be made operational, stating that “different scales are delineated by a succession of divergences in statistical complexity at lower levels.” More recently, Rupe and Crutchfield have reported progress towards identifying emergent self-organisation in a system.

Although this work on quantitative measures of emergence based on information theory represents significant progress, there are many open problems. Examples include the extension to non-Markovian systems and the development of computationally feasible methods for large systems. The latter is particularly important in physical systems where spontaneous symmetry breaking occurs, as this only happens in the thermodynamic limit of an infinite system.

There is an unrecognised similarity between the work described above and techniques recently developed to characterise phase transitions in statistical mechanics models such as the Ising model and classical dimer models. Coarse-graining (CG) is optimised by maximising the Real-Space Mutual Information (RSMI) between a spatial block and its distant environment. 

In general, maximising mutual information is notoriously hard but can be done using state-of-the-art machine learning algorithms. Gokmen et al. have developed an algorithm that they claim “can, unsupervised, construct order parameters, locate phase transitions, and identify spatial correlations and symmetries for complex and large-dimensional real-space data.” Furthermore, the optimal CG explicitly identifies the scaling operators associated with the critical point. 

The classical dimer model provides a stringent test as “the relevant low-energy degrees of freedom are profoundly different from the microscopic building blocks of the theory and change qualitatively throughout the phase diagram.” In other words, the emergent entities (quasiparticles such as vortices associated with the height field, which is described by a sine-Gordon field theory) are different from the dimers.

It is encouraging to see that two different scientific communities have developed similar ideas to address this challenging problem of making discussions about emergence and causality more concrete and quantitative.

Monday, January 26, 2026

What is absolute temperature?

The concept and reality of absolute temperature is amazing. It tells us something fundamental about the universe, including physical limits as to what is possible. The existence of absolute temperature is intimately connected with the existence of entropy as a thermodynamic state function. It also hints at the underlying quantum nature of reality.

Aside: Unfortunately, the Wikipedia page on this topic is mediocre and garbled. For example, it continues the myth that temperature is related to kinetic energy.

The zeroth law of thermodynamics allows the definition of empirical temperature. It is an equilibrium state variable that indicates whether a thermodynamic system will remain in the same state upon being brought into thermal contact with another system. Thermometers are systems with a single state variable.

Absolute temperature is a specific temperature scale that is central to thermodynamics and statistical mechanics. 

There are several equivalent definitions of absolute temperature. They start at different points. Except for the first one, the others show that the existence of absolute temperature is intimately connected to the second law and to entropy being an extensive quantity.

This is nicely discussed by Zemansky in chapter 8 of his text Heat and Thermodynamics, Fifth Edition (1968). [This was the text for my second year undergrad thermo course at ANU in 1980. At the time, I did not fully appreciate how profound some of it is. I just enjoyed all the multivariable calculus.] 

1. Ideal gas thermometers.

Consider a fixed mass of ideal gas whose volume is fixed. An ideal gas is defined as any gas at a temperature and pressure much larger than the critical temperature and pressure for the gas-liquid transition. Suppose the system is cooled and heated, and the pressure is measured as a function of the temperature measured by a separate thermometer calibrated by the Celsius scale. The pressure versus temperature curve is a straight line. If this line is extrapolated to zero pressure, this occurs at -273.15 degrees Celsius. The straight line has different slopes for different gases, but they all intercept the x-axis at the same point. Alternatively, one can take the pressure as fixed and measure the volume of the gas versus temperature. Extrapolation to zero volume also occurs at -273.15 degrees. 

This suggests that something special is happening at -273.15 degrees Celsius. One can define a special temperature scale where this temperature is zero. Historically, this was the beginning of the concept of absolute temperature.

However, we should be cautious about this approach. This is just an extrapolation and does not allow for the fact that ideal gases are rather special or that some very different physics might kick in below the critical temperature of helium.

2. The efficiency of Carnot cycles. 

This follows Zemansky (page 208). Consider a Carnot cycle abcda, where b to c and d to a are isothermal processes, between the same two reversible adiabatic surfaces, and involve heat transfers Q and Q_3, respectively. The absolute temperature scale T is defined by 

T/T_3 = Q/Q_3

with T_3 = 273.16, when the process d to a occurs at the triple point of water.

3. Integrating factor for heat

Heat is not a state property. It depends on processes. The first law says Delta Q = Delta U + P Delta V. If we consider a quasi-static process and integrate the heat transfer along the path taken (in state space), the result may depend on the path taken. On the other hand, if one integrates dQ/T, one finds that the result is independent of the path. This can then be used to define a new state variable, the entropy. 

The brief discussion above misses some subtle and profound features that only became clear in the 1960s following the work of Pippard, Turner, Landsberg, and Sears, which was inspired by an axiomatic approach to thermodynamics developed by Caratheodory.

Zemansky states

It is an extraordinary circumstance that not only does an integrating factor exist for the dQ of any system, but this integrating factor is a function of temperature only and is the same function for all systems! This universal character enables us to define an absolute temperature.

4. Applying the second law to a composite system

This treatment follows Schroeder, Thermal Physics (Section 3.1)

Schroeder defines entropy in terms of a multiplicity of states. However, I prefer to define entropy as the state function which tells us whether or not two states are accessible from one another by an adiabatic process. There are multiple possible versions of this empirical entropy state function, but let's choose one that is extensive, i.e., scales with the mass and volume of the system.

Consider an adiabatically isolated system containing an internal partition through which the conduction of heat can occur. Denote the two parts of the system by A and B. The entropy of each part can be written as a function U of its internal energy. 

The total entropy of the system can be written 

S = S_A (U_A) + S_B (U_B)

If the system is in thermal equilibrium, by the second law, the entropy of the whole system must be a minimum as a function of U_A and U_B.

Now, dU_A = - dU_B as the composite system is adiabatically isolated. Hence, we have.


The left-hand (right-hand) side of the equation only depends on the properties of system A (B). Thus, it is an intensive state variable which determines whether the system will be in equilibrium with another system. Hence, by the zeroth law, it defines a temperature scale.

T is the absolute temperature.

Monday, January 5, 2026

Maxwell's demon and the history of the second law of thermodynamics

I recently reread Warmth Disperses and Time Passes: The History of Heat by Hans Christian von Baeyer

As a popular book, it provides a beautiful and enthralling account of the discovery of the first and second laws of thermodynamics. The book is a great companion to teaching and learning thermodynamics and statistical mechanics. The narrative is unified by the puzzle of Maxwell's demon.

Aside: The book was first published in 1998 with the title Maxwell's Demon. My guess is that the publisher changed the title because most people have probably not heard of the demon, unlike Schrodinger's cat.

Baeyer captures both the wonder of the subject and the fascinating story of how the science of thermodynamics developed. He describes quirky personalities and illustrates how science proceeds with a mixture of brilliant insights, clever experiments, false leads, and forgotten discoveries. It is easy and compelling reading.

I appreciated that there is a lack of hype, in contrast to too many popular science books.

The book is enhanced by showing that the story is not over. Many reports of the demise of the demon have been premature. The penultimate chapter discusses Zurek's definition of entropy in terms of algorithmic randomness. The last chapter considers molecular motors, such as kinesin, which can be viewed as ratchets driven by thermal noise.

Physical insights

The first and second laws tell us something about the fundamental nature of the universe. Although they are macroscopic and may have some (debatable) microscopic justification,  they can be viewed as fundamental.

Central to the development of the first law was the notion of the mechanical equivalent of heat.

There are three rather different ways to formulate the second law: a Carnot cycle represents an engine of optimal efficiency, heat never passes from a cold to a hot body, and the arrow of time. It is profound that these formulations are equivalent and not something that was anticipated. We should marvel at this.

Entropy can be viewed as the absence of information. Consequently, the second law can be viewed as statistical.

Things I want to understand

A good book stimulates us to want to engage more with its subject. Some things I want to understand are the entropy of the initial state of the universe, Boltzmann's H theorem, Feynman's ratchet, Shannon's information theory, molecular motors, Zurek's definition of entropy, and Gerald Holton's book, Thematic origins of scientific thought.

A recent tutorial is A Friendly Guide to Exorcising Maxwell’s Demon, by A. de Oliveira Junior, Jonatan Bohr Brask, and Rafael Chaves

Beautiful things missed

As a popular book, I think the length and scope of topics are right. Nevertheless, in a longer book, here are some things I would enjoy reading about: the zeroth and third laws, the contributions of Gibbs, the ergodic hypothesis, Brownian motion and evidence for atoms, the role of thermodynamics (and statistical mechanics) in the development of quantum theory (blackbody radiation, Einstein solid, identical particle statistics, and the Sackur-Tetrode equation) and perhaps phase transitions.

Two quibbles

von Baeyer has a somewhat reductionist perspective that the true nature of thermodynamics was revealed by the microscopic descriptions of Maxwell and Boltzmann.

I will write separate posts on why I am not comfortable with the following two statements.

Temperature IS the average kinetic energy of molecules.

Entropy was mysterious until Boltzmann's definition S=k ln W. 

Friday, August 22, 2025

The two-state model for spin crossover in organometallics

Previously, I discussed how spin-crossover is a misnomer for organometallic compounds and proposed that an effective Hamiltonian to describe the rich states and phase transitions is an Ising model in "magnetic field".

I introduce the two-state model that defines the model without the Ising interactions. To save me time on formatting in HTML, here is a pdf file that describes the model and what comparisons with experimental data (such as that below) tells us.

Future posts will consider how elastic interactions produce the Ising interaction and how frustrated interactions can produce multi-step transitions.

Friday, June 27, 2025

Thermodynamics and emergence

Novelty. 

Temperature and entropy are emergent properties. Classically, they are defined by the zeroth and second laws of thermodynamics, respectively. The individual particles that make up a system in thermodynamic equilibrium do not have these properties. Kadanoff provided an example illustrating the qualitative difference between macro- and micro-perspectives. He pointed out how deterministic behaviour can emerge at the macroscale from stochastic behaviour at the microscale. The many individual molecules in a dilute gas can be viewed as undergoing stochastic motion. However, collectively they are described by an equation of state such as the ideal gas law.

 Primas gave a technical argument, involving C* algebras, that temperature is emergent: it belongs to an algebra of contextual observables but not to the algebra of intrinsic observables.44 Following this perspective, Bishop argued that temperature and the chemical potential are (contextually) emergent.

Intra-stratum closure. 

The laws of thermodynamics, the equations of thermodynamics (such as TdS = dU + pdV), and state functions such as S(U,V), provide a complete description of processes involving equilibrium states. A knowledge of microscopic details, such as the atomic constituents or forces of interaction, is not necessary for the description.

Irreducibility. 

A common view is that thermodynamics can be derived from statistical mechanics. However, this is contentious. David Deutsch claimed that the second law of thermodynamics is an “emergent law”: it cannot be derived from microscopic laws, like the principle of testability.

Lieb and Yngvason stated that the derivation from statistical mechanics of the law of entropy increase “is a goal that has so far eluded the deepest thinkers.”  In contrast, Weinberg claimed that Maxwell, Boltzmann, and Gibbs “showed that the principles of thermodynamics could in fact be deduced mathematically, by an analysis of the probabilities of different configurations… Nevertheless, even though thermodynamics has been explained in terms of particles and forces, it continues to deal with emergent concepts like temperature and entropy that lose all meaning on the level of individual particles.” (Dreams of A Final Theory, pages 40-41)

I agree that thermodynamic properties (e.g., equations of state, the temperature dependence of heat capacity, and phase transitions) can be deduced from statistical mechanics. However, thermodynamic principles, such as the second law, are not thermodynamic properties. Furthermore, these thermodynamic principles are required to justify the equations of statistical mechanics, such as the partition function, that are used to calculate thermodynamic properties. 

Macro hints of microscopics.

The Sackur-Tetrode equation for the entropy of an ideal gas hinted at the quantisation of phase space. The Gibbs paradox hinted that fundamental particles are indistinguishable. The third law of thermodynamics hints at quantum degeneracy.

Monday, April 29, 2024

Emergence of the arrow of time

Time has a direction. Microscopic equations of motion in classical and quantum mechanics have time-reversible symmetry. But this symmetry is broken for many macroscopic phenomena. This observation is encoded in the second law of thermodynamics. We experience the flow of time and distinguish past, present, and future. The arrow of time is manifest in phenomena that occur at scales covering many orders of magnitude. Here are some of these different arrows of time, listed in order of increasing time scales. These are discussed by Tony Leggett in chapter 5 of The Problems of Physics.

Elementary particle physics. CP violation is observed in certain phenomena associated with the weak nuclear interaction, such as the decay of neutral kaons observed in 1964. The CPT symmetry theorem shows that any local quantum field theory that is invariant under the “proper” Lorentz transformations must also be invariant under combined CPT transformations. This means that CP violation means that time-reversal symmetry is broken. In 1989, the direction violation of T symmetry was observed.

Electromagnetism. When an electric charge is accelerated an electromagnetic wave propagates out from the charge towards infinity. Energy is transferred from the charge to its environment. We do not observe a wave that propagates from infinity into the accelerating charge, i.e., energy being transferred from the environment to the charge. Yet this possibility is allowed by the equations of motion for electromagnetism. There is an absence of the “advanced” solution to the equations of motion. 

Thermodynamics. Irreversibility happens in isolated systems. Heat never travels from a cold body to a hotter one. Fluids spontaneously mix. There is a time ordering of the thermodynamic states of isolated macroscopic systems. The thermodynamic entropy encodes this ordering.

Psychological experience. We remember the past and think we can affect the future. We don’t think we can affect the past or know the future.

Biological evolution. Over time species adapt to their environment and become more complex and more diverse.

Cosmology. There was a beginning to the universe. The universe is expanding not contracting. Density perturbations grow independent of cosmic time (Hawking and Laflamme).

It is debatable to what extent these arrows of time are related to one another. 

The problem of how statistical mechanics connects time-reversible microscopic dynamics with macroscopic irreversibility is subtle and contentious. Joel Lebowitz claimed this problem was solved by Boltzmann, provided the distinction between typical and average behaviour are accepted, along with the Past Hypothesis. This states that the universe was initially in a state of extremely low entropy. David Wallace discussed the need to accept the idea of probabilities in law of physics and that the competing interpretations of probability as frequency or ignorance matter. In contrast, David Deutsch claims that the second law of thermodynamics is an “emergent law”: it cannot be derived from microscopic laws, like the principle of testability.

I find the Past Hypothesis fascinating because it connects the arrow of time seen in the laboratory and everyday life (time scales of microseconds to years) to cosmology, covering timescales of the lifetime of the universe (10^10 years) and the “initial” state of the universe, perhaps at the end of the inflationary epoch (10^-33 seconds). This also raises questions about how to formulate the Second Law and the concept of entropy in the presence of gravity and on cosmological length and time scales. 

Monday, July 31, 2023

What is a complex system?

What do we mean when we say a particular system is "complex"? We main have some intuition that it means there are many degrees of freedom and/or that it is hard to understand. "Complexity" is sometimes used as a buzzword, just like "emergence." There are many research institutes that claim to be studying "complex systems" and there is something called "complexity theory". Complexity seems to mean different things to different people.

I am particularly interested in understanding the relationship between emergence and complexity. To do this we first need to be more precise about what we mean by both terms. A concrete question is the following. Consider a system that exhibits emergent properties. Often that will be associated with a hierarchy of scales. For example: atoms, molecules, proteins, DNA, genes, cells, organs, people. The corresponding hierarchy of research fields is physics, chemistry, biochemistry, genetics, cell biology, physiology, psychology. Within physics a hierarchy is quarks and leptons, nuclei and electrons, atoms, molecules, liquid, and fluid. 

In More is Different, Anderson states that as one goes up the hierarchy the system scale and complexity increases. This makes sense when complexity is defined in terms of the number of degrees of freedom in the system (e.g., the size of the Hilbert space needed to describe the complete state of the system). On the other hand, the system state and its dynamics become simpler as one goes up the hierarchy.  The state of the liquid can be described completely in terms of the density, temperature, and the equation of state. The dynamics of the fluid can be described by the Navier-Stokes equation. Although that is hard to solve in the regime of turbulence, the system is still arguably a lot simpler than quantum chromodynamics (QCD)! Thus, we need to be clearer about what we mean by complexity.

To address these issues I found the following article very helpful and stimulating.

What is a complex system? by James Ladyman, James Lambert, and Karoline Wiesner 

It was published in 2013 in a philosophy journal, has been cited more than 800 times, and is co-authored by two philosophers of science and a physicist.

[I just discovered that Ladyman and Wiesner published a book with the same title in 2020. It is an expansion of the 2013 article.].

In 1999 the journal Science had a special issue that focussed on complex systems, with an Introduction entitled, Beyond Reductionism. Eight survey articles covered complexity in physics, chemistry, biology, earth science, and economics.

Ladyman et al., begin by pointing out how each of the authors of these articles chooses different properties to define what complexity is associated. These characteristics include non-linearity, feedback, spontaneous order, robustness and lack of central control, emergence, hierarchical organisation, and numerosity.

The problem is that these characteristics are not equivalent. If we do choose a specific definition for a complex system, the difficult problem then remains of determining whether each of the characteristics above is necessary, sufficient, both, or neither for the system to be complex (as defined). This is similar to what happens with attempts to define emergence.

Information content is sometimes used to quantify complexity. Shannon entropy and Kolmogorov complexity (Sections 3.1, 3.2) are discussed. The latter is also known as algorithmic complexity. This is the length of the shortest computer program (algorithm) that can be written to produce the entity as output. A problem with both these measures are they are non-computable.

Deterministic complexity is different from statistical complexity (Section 3.3). A deterministic measure treats a completely random sequence of 0s and 1s as having maximal complexity. A statistical measure treats a completely random sequence as having minimal complexity. Both Shannon and algorithmic complexity are deterministic.

Section 4 makes some important and helpful distinctions about different measures of complexity.

3 targets of measures: methods used, data obtained, system itself

3 types of measures: difficulty of description, difficulty of creation, or degree of organisation

They then review three distinct measures that have been proposed logical depth (Charles Bennett), thermodynamic depth (Seth Lloyd and Heinz Pagels), and effective complexity (Murray Gell-Mann).

Logical depth and effective complexity are complementary quantities. The Mandelbrot set is example of a system (set of data) that exhibits a complex structure that has a high information content. It is difficult to describe. It has a large logical depth.

Created by Wolfgang Beyer with the program Ultra Fractal 3. 

On the other hand, the effective complexity of the set is quite small since it can be generated using the simple equation

z_n+1 = c + z_n^2

c is a complex number and the Mandelbrot set is the values of c for which the iterative map is bounded.

Ladyman et al, prefer the definition of a complex system below, but do acknowledge its limitations.

(Physical account) A complex system is an ensemble of many elements which are interacting in a disordered way, resulting in robust organisation and memory. 

(Data-driven account) A system is complex if it can generate data series with high  statistical complexity. 

What is statistical complexity? It relates to degrees of pattern and some they refer to as causal state reconstruction. It is applied to data sets, not systems or methods. Central to their definition is the idea of the epsilon-machine, something introduced in a long and very mathematical article from 2001, Computational Mechanics: Pattern and Prediction, Structure and Simplicity, by Shalizi and Crutchfield.

The article concludes with a philosophical question. Do patterns really exist? This relates to debates about scientific realism versus instrumentalism. The authors advocate something known as "rainforest realism", that has been advanced by Daniel Dennett, Don Ross, and David Wallace. A pattern is real if one can construct an epsilon-machine that can simulate the phenomena and predict its behaviour.

I don't have a full appreciation or understanding of where the article ends up. Nevertheless, the journey there is helpful as it clarifies some of the subtleties and complexities (!) of trying to be more precise about what we mean by a "complex system".

Monday, June 26, 2023

What is really fundamental in science?

What do we mean when we say something in science is fundamental? When is an entity or a theory more fundamental or less fundamental than something else? For example, are quarks and leptons more fundamental than atoms? Is statistical mechanics more fundamental than thermodynamics? Is physics more fundamental than chemistry or biology? In a fractional quantum Hall state, are electrons or the fractionally charged quasiparticles more fundamental?

Answers depend on who you ask. Physicists such as Phil Anderson, Steven Weinberg, Bob Laughlin, Richard Feynman, Frank Wilczek, and Albert Einstein have different views.

In 2017-8, the Foundational Questions Institute (FQXi) held an essay contest to address the question, “What is Fundamental?” Of the 200 entries, 15 prize-winning essays have been published in a single volume. The editors give a nice overview in the Introduction.

This post is mostly about the essay, Fundamental? of the first prize winner, Emily Adlam, a philosopher of physics. She contrasts two provocative statements.

Fundamental means we have won. The job is done and we can all go home.

Fundamental means we have lost. Fundamental is an admission of defeat.

This raises the question of whether being fundamental is objective or subjective.

Examples are given from scientific history to argue that what is considered to be fundamental has changed with time. The reductionism has led to the drive to explain everything in terms of smaller and smaller entities, that are deemed 'more fundamental". But we find that smaller does not always mean simpler.

Perhaps we should ask what needs explaining and what constitutes a scientific explanation. For example, Adlam asks whether explaining the fact that the initial state of the universe had a low entropy [the "past hypothesis"] is really possible or should be an important goal.

She draws on the issue of the distinction between objective and subjective probabilities. Probabilities in statistical mechanics are subjective: they are a statement about our own ignorance about the details of the motion of individual atoms and not any underlying randomness in nature. In contrast, probabilities in quantum theory reflect objective chance.

as realists about science we must surely maintain that there is a need for science to explain the existence of the sorts of regularities that allow us to make reliable predictions... but there is no similarly pressing need to explain why these regularities take some particular form rather than another. Yet our paradigmatic mechanical explanations do not seem to be capable of explaining the regularity without also explaining the form, and so increasingly in modern physics we find ourselves unable to explain either. 

It is in this context that we naturally turn to objective chance. The claim that quantum particles just have some sort of fundamental inbuilt tendency to turn out to be spin up on some proportion of measurements and spin down on some proportion of measurements does indeed look like an attempt to explain a regularity (the fact that measurements on quantum particles exhibit predictable statistics) without explaining the specific form (the particular sequence of results obtained in any given set of experiments). But given the problematic status of objective chance, this sort of nonexplanation is not really much better than simply refraining from explanation at all. 

Why is it that objective chances seem to be the only thing we have in our arsenal when it comes to explaining regularities without explaining their specific form? It seems likely that part of the problem is the reductionism that still dominates the thinking of most of those who consider themselves realists about science

In summary, (according to the Editors) Adlam argues that "science should be able to explain the existence of the sorts of regularities that allow us to make reliable predictions. But this does not necessarily mean that it must also explain why these regularities take some particular form." 

we are in dire need of another paradigm shift. And this time, instead of simply changing our attitudes about what sorts of things require explanation, we may have to change our attitudes about what counts as an explanation in the first place. 

Here, she is arguing that what is fundamental is subjective, being a matter of values and taste.

In our standard scientific thinking the fundamental is elided with ultimate truth: getting to grips with the fundamental is the promised land, the endgame of science. 

She then raises questions about the vision and hopes of scientific reductionists. 

In this spirit, the original hope of the reductionists was that things would get simpler as we got further down, and eventually we would be left with an ontology so simple that it would seem reasonable to regard this ontology as truly fundamental and to demand no further explanation. 

But the reductionist vision seems increasingly to have failed. 

When we theorise beyond the standard model [BSM] we usually find it necessary to expand the ontology still more: witness the extra dimensions required to make string theory mathematically consistent.

It is not just strings. Peter Woit has emphasised how BSM theories, such as supersymmetry, introduce many more particles and parameters.

... the messiness deep down is a sign that the universe works not ‘bottom-up’ but rather ‘top-down,’ ... in many cases, things get simpler as we go further up.

Our best current theories are renormalisable, meaning that many different possible variants on the underlying microscopic physics all give rise to the same macroscopic physical theory, known as an infrared fixed point. This is usually glossed as providing an explanation of why it is that we can do sensible macroscopic physics even without having detailed knowledge of the underlying microscopic theories. 

For example, elasticity theory, thermodynamics and fluid dynamics all work without knowing anything about atoms, statistical mechanics, and quantum theory.

But one might argue that this is getting things the wrong way round: the laws of nature don’t start with little pieces and build the universe from the bottom up, rather they apply simple macroscopic constraints to the universe as a whole and work out what needs to happen on a more fine-grained level in order to satisfy these constraints.

This is rather reminiscent of Laughlin's views about what is fundamental.

Finally, I mention two other essays that I look forward to reading as I think they make particularly pertinent points.

Marc Séguin (Chap. 6) distinguishes "between epistemological fundamentality (the fundamentality of our scientific theories) and ontological fundamentality (the fundamentality of the world itself, irrespective of our description of it)."

"In Chap. 12, Gregory Derry argues that a fundamental explanatory structure should have four key attributes: irreducibility, generality, commensurability, and fertility."

[Quotes are from the Introduction by the Editors].

Some would argue that the Standard Model is fundamental, at least on some level. But it involves 19 parameters that have to be fixed from experiment. Related questions about the Fundamental Constants, have been explored in a 2007 paper by Frank Wilczek.

Again, I thank Peter Evans for bringing this volume to my attention.

Saturday, June 17, 2023

Why do deep learning algorithms work so well?

I am interested in analogues between cognitive science and artificial intelligence. Emergent phenomena occur in both, there have been some fruitful cross-fertilisation of ideas, and the extent of the analogues is relevant to debates on fundamental questions concerning human consciousness.

Given my general ignorance and confusion on some of the basics of neural networks, AI, and deep learning, I am looking for useful and understandable resources.

Related questions are explored in a nice informative article from 2017 in Quanta magazine, New Theory Cracks Open the Black Box of Deep Learning by Natalie Wolchover.

Like a brain, a deep neural network has layers of neurons — artificial ones that are figments of computer memory. When a neuron fires, it sends signals to connected neurons in the layer above. During deep learning, connections in the network are strengthened or weakened as needed to make the system better at sending signals from input data — the pixels of a photo of a dog, for instance — up through the layers to neurons associated with the right high-level concepts, such as “dog.” 

After a deep neural network has “learned” from thousands of sample dog photos, it can identify dogs in new photos as accurately as people can. The magic leap from special cases to general concepts during learning gives deep neural networks their power, just as it underlies human reasoning, creativity and the other faculties collectively termed “intelligence.” 

Experts wonder what it is about deep learning that enables generalization — and to what extent brains apprehend reality in the same way.

The article describes work by Naftali Tishby and collaborators that provides some insight into why deep learning methods work so well. This was first described in purely theoretical terms in a 2000 preprint

The information bottleneck method, Naftali Tishby, Fernando C. Pereira, William Bialek 

The idea is that a network rids noisy input data of extraneous details as if by squeezing the information through a bottleneck, retaining only the features most relevant to general concepts.

Tishby was stimulated in new directions in

2014 after reading a surprising paper by the physicists David Schwab and Pankaj Mehta

 An exact mapping between the Variational Renormalization Group and Deep Learning 

[They] discovered that a deep-learning algorithm invented by Geoffrey Hinton called the “deep belief net” works, in a particular case, exactly like renormalization [group methods in statistical physics... When they]. applied the deep belief net to a model of a magnet at its “critical point,” where the system is fractal, or self-similar at every scale, they found that the network automatically used the renormalization-like procedure to discover the model’s state. 

Although this connection was a valuable new insight, the specific case of a scale-free system, is not relevant to many deep learning situations.

Tishby and Ravid Shwartz-Ziv discovered that 

Over the course of training, common patterns in the training data become reflected in the strengths of the connections, and the network becomes expert at correctly labeling the data, such as by recognizing a dog, a word, or a 1.

...layer by layer, the networks converged to the information bottleneck theoretical bound: a theoretical limit derived in Tishby, Pereira and Bialek’s original paper that represents the absolute best the system can do at extracting relevant information. At the bound, the network has compressed the input as much as possible without sacrificing the ability to accurately predict its label...

...deep learning proceeds in two phases: a short “fitting” phase, during which the network learns to label its training data, and a much longer “compression” phase, during which it becomes good at generalization, as measured by its performance at labeling new test data.

What these new discoveries teach us about the relationship between learning in humans and in machines is contentious and explored briefly in the article. Although neural nets were inspired by the structure of the human brain the connection with the neural nets used today is tenuous.

The mystery of how brains sift signals from our senses and elevate them to the level of our conscious awareness drove much of the early interest in deep neural networks among AI pioneers, who hoped to reverse-engineer the brain’s learning rules. AI practitioners have since largely abandoned that path in the mad dash for technological progress, instead slapping on bells and whistles that boost performance with little regard for biological plausibility.

Monday, December 2, 2019

Ising model basics

The Ising model is a paradigm in both statistical mechanics and condensed matter physics. Today for most theorists it is so familiar that some of its historical and conceptual significance is lost.
Previously, I posted about what students can learn from computer simulations of the Ising model.

If you had to talk about the Ising model to an experimental chemist what would you say?
[Last week I had to do this].

The Ising model is the simplest effective model Hamiltonian that can describe a thermodynamic system that undergoes a first-order phase transition and has a phase diagram containing a critical point.

On each site i of a lattice one defines a spin sigma_i= +1 or -1, representing spin up or spin down.

The Hamiltonian H is

J_ij describes the interaction between spins on sites i and j. In the simplest version the interactions are only between nearest neighbours, and have the same value J.
h is the external magnetic field.

If J is positive, the ground state at h=0 is a ferromagnet.
If J is negative, the ground state at h=0 is an anti-ferromagnet for a bipartite lattice.

[Caution: just like for the Heisenberg model, some authors define the Hamiltonian with the opposite sign of J].

For h=0 there is a critical point at a finite temperature Tc, for lattices of dimension two and higher.

The spins sigma_i= +/- 1 defined at each lattice site i, were originally to represent the atomic magnetic moments in a ferromagnetic material. However, the sigma's can represent any two states of the site i. For example, the ``spin'' or pseudo-spin can represent the presence or absence of an atom or molecule in a ``lattice gas'', atom A or atom B in a binary alloy (mixture), or the low-spin and high-spin states in a spin-crossover material.

The mean-field theory of the Ising model is mathematically equivalent to the thermodynamic theory of binary mixtures with an entropy of an ideal mixture.
There is a nice discussion of such mixtures in Section 5.4 [and the associated problems] of Introduction to Thermal Physics by Schroeder.
[Here are the slides for a lecture I have given based on that text].
Chapter 15 of the text by Dill and Bromberg is also helpful as it has more detail.
Neither text makes an explicit connection to the Ising model. Following this paper on alloys, one has

This is shown in Section 8.1.2 of James Sethna's text, Statistical MechanicsEntropy, Order Parameters and Complexity.

When interactions beyond nearest-neighbours are included in the Ising model or when the lattice is frustrated (e.g. fcc or triangular) a richer phase diagram is possible. Examples include the ANNNI model and some models for spin-state ice considered by Jace Cruddas and Ben Powell.

Friday, January 18, 2019

First-order transitions and critical points in spin-crossover compounds

An interesting feature of spin-crossover compounds is that the transition from low-spin to high-spin with increasing temperature is usually a first-order phase transition. This is associated with hysteresis and the temperature range of the hysteresis varies significantly between compounds.
If there was no interaction between the transition metal ions the transition would be a smooth crossover. This is nicely illustrated in a figure taken from the paper below.

Abrupt versus Gradual Spin-Crossover in FeII(phen)2(NCS)2 and FeIII(dedtc)3 Compared by X-ray Absorption and Emission Spectroscopy and Quantum-Chemical Calculations 
Stefan Mebs, Beatrice Braun, Ramona Kositzki, Christian Limberg, and Michael Haumann


For the first compound, the transition is abrupt [much earlier work found a narrow hysteresis region of about 0.15 K]. For the second compound, the transition is a crossover.

The authors fit their data to an empirical equation that has a parameter n, describing the "interactions". You have to read the Supplementary Material to find the details. This equation cannot describe hysteresis.

 However, there is an elegant analytical theory going back to a paper by Wajnflasz and Pick from 1971. This is nicely summarised in the first section of a paper by Kamel Boukheddaden, Isidor Shteto, Benoit Hôo, and François Varret.
The system can be described by the Ising model

where the Ising spin denotes the high- and low-spin states. Delta is the energy difference between them and ln g the entropy difference.
The mean-field Hamiltonian for q nearest neighbours is

There are two independent dimensionless variables, d and r. Solving for the fraction of high-spin states (HS) versus temperature gives the graphs below for different values of d.
The vertical arrows show the hysteresis region for a specific value of d=2. 
As d increases the hysteresis region gets smaller. Above the critical value of d=r/2, the crossover temperature T0=Delta/ln g is larger than the mean-field critical temperature Tc= qJ, and the transition is no longer first-order but a crossover.
Using DFT-based quantum chemistry, the authors calculate the change in vibrational frequencies and the associated entropy change for the SCO transition in a single molecule. The values for compounds 1 and 2 are 0.68  and 0.21 meV/K, respectively. The spin entropy changes are 0.21 and 0.22 meV/K respectively. The total entropy changes are thus 0.89 and 0.43 meV/K respectively. The values of Delta are 175 and 125 meV, respectively. The corresponding crossover temperatures are 210 and 360 K, compared to the experimental values of 176 and 285 K.

If we assume that J is roughly the same for both compounds, then the fact that the entropy change is half as big for compound 2, means r is twice as big. This naturally explains why the second compound has a smooth crossover, compared to the first, which is very close to the critical point.

Wednesday, May 30, 2018

Broken symmetry, order, and entropy

One of the greatest joys of teaching is having students ask questions that you do not know the answer to. In the last week of the course PHYS2020 Thermodynamics and Condensed Matter Physics for second year undergrads at UQ, I give two lectures about critical points, universality, critical exponents, broken symmetry, order parameters, and Landau theory.

Many students find this quite challenging. However, I think it is important that students be exposed to two of the most important ideas of theoretical physics from the twentieth century: broken symmetry and universality. Furthermore, there is no technical reason why second year undergrads cannot learn this material. Since the text, Thermal Physics by Schroeder, does not cover this material we have finally settled on a chapter from a book by Hoch.

After my last lecture, a student asked an excellent question along the lines of
"Why is it that broken symmetry occurs at lower temperatures?
How is this related to entropy and order?"

This led me to wondering whether there were any rigorous results that answer the question. I could not find anything in a quick search.
Do you know of anything?

I was wondering whether something like the following conjecture was true:
Conjecture. Consider a physically reasonable Hamiltonian H for an infinite system. Suppose H is invariant under some symmetry group G. Let rho(T) be the equilibrium density matrix at temperature T. Then for sufficiently large T, rho(T) is also invariant under G.
Maybe this is equivalent to
Lemma. At sufficiently high temperatures, the von Neumann entropy S (rho) = - Tr( rho ln (rho)) is maximal if rho is invariant under G. 
This looks to me like the kind of thing that people like Elliot Lieb, David Ruelle, Y. Sinai, ... might have tackled at some point.

I welcome ideas and suggestions.

Friday, April 27, 2018

Relating frustrated spin models and flat bands in tight-binding models

What kind of theory paper to I enjoy?
Here are some personal tastes
- "simple" enough I can understand it
- physical insight
- some analytical results
- some pretty pictures that illuminate

This week I read the following paper which I consider nicely meets these criteria.

Band touching from real-space topology in frustrated hopping models
Doron L. Bergman, Congjun Wu, and Leon Balents

The quantum spin antiferromagnetic Heisenberg model on the kagome lattice attracts a lot of attention because it may have a spin liquid ground state, for spin-1/2 and spin 1. This is arguably driven by the large spin frustration. A reflection of this frustration is that the classical model has a non-zero entropy at zero temperature due to a manifold of degenerate states. For this reason, the kagome lattice is sometimes said to be "maximally frustrated". This is in contrast to the triangular lattice for which their is a unique classical ground state and the spin-1/2 model exhibits long-range order.

The kagome lattice is also of interest because of the band structure for the tight-binding model has a flat band, i.e. it is dispersionless. This means that in the presence of interactions the electrons in this band may be strongly correlated and susceptible to instability to new states of matter.

The question arises as to whether there is any connection between these two properties of models on a particular "frustrated" lattice: flat bands and a manifold of degenerate classical ground states.

The purpose of this paper is to show that for a whole class of lattices, in two and three dimensions, that there is an close relationship between these properties.
It turns out that a key feature is that the flat bands touch a dispersive band at one point in k-space.

My interest was stimulated by the work of some of my UQ colleagues on a class of organometallic compounds that exhibit a kagomene lattice (that interpolates between kagome and honeycomb (graphene). The associated band structure (taken from this paper) is shown below.

The abstract states:
We demonstrate that this band touching is related to states which exhibit nontrivial topology in real-space. Specifically, these states have support [i.e. non-zero values] on one-dimensional loops which wind around the entire system 􏰀with periodic boundary conditions􏰁. A counting argument is given that determines, in each case, whether there is band touching or none, in precise correspondence to the result of straightforward diagonalization. When they are present, the topological structure protects the band touchings in the sense that they can only be removed by perturbations, which also split the degeneracy of the flat band.
I know illustrate this with the kagome lattice.

It has a three site basis (mu=1,2,3) and so there are three bands. If q is the Bloch wave vector, the Bloch states for the flat band can be written

One of these plaquette states is shown on the left below. 
A key point is that there is constructive interference between these plaquette states. Thus, one can take superpositions of them. On the right is the superposition of three neighbouring plaquette states.

A whole line of plaquette states can lead to visualising something with nontrivial topology.

The authors then show how similar physics occurs in other two- and three-dimensional lattice models. The one below is the dice lattice.
Finally, they show that the corresponding Hubbard model leads to a Heisenberg model in the classical limit does have macroscopic degeneracy.

I thank Ben Powell for bringing the paper to my attention.

Wednesday, March 14, 2018

"Bad fluids" near the superfluid transition

There is an interesting preprint
Viscosity Bound Violation in Viscoelastic Fermi Liquids 
 Matthew P. Gochan, Hua Li, Kevin S. Bedell

They consider the unitary Fermi gas within the framework of Fermi liquid theory. This system undergoes a superfluid transition at a temperature of about 0.17 times T_F (the Fermi temperature). They calculate the shear viscosity as a function of temperature. (I think) the complete temperature dependence is obtained by interpolating between the low-temperature and high-temperature limits.

The motivation for the study is the conjectured universal bound for the ratio of the shear viscosity to the entropy density, based on the AdS-CFT conjecture, beloved by string theorists.

The authors find that the conjectured bound is violated because the viscosity can become arbitrarily small near the superfluid transition due to large scattering from superfluid fluctuations. This is because the mean free path becomes arbitrarily small, i.e. the system is similar to a bad metal.
Unfortunately, the preprint does not reference some earlier relevant work on the shear viscosity of the unitary Fermi gas or on the bad metal near a Mott transition.

I thank Alejandro Mezio for bringing the preprint to my attention.

Monday, March 12, 2018

A new class of "spin ice" materials

Two of my UQ colleagues have just finished a nice paper:
Spin-state ice in geometrically frustrated spin-crossover materials 
Jace Cruddas, B. J. Powell

The paper brings together two fascinating topics I have written about before, spin crossover materials and spin ice. One thing that it is a little worrying and disappointing about spin ice materials is that there seem to be only two (?) of them!
This paper argues that some spin crossover materials may be a new class of materials that realise ice physics (residual entropy, emergent gauge fields, monopoles, ...) Here, the Ising spin variable is the two possible spin states (High Spin and Low Spin). These materials have the potential advantage that they may be tuneable due to the creativity of synthetic chemists.
The mechanism of the interaction between spins is rather unique and interesting. It is not an exchange interaction but rather and effective interaction mediated by the spin-lattice interaction, which in these compounds is arguably large.
It is also interesting that the sign of the frustrating interactions (which are key to the stability of the ice) is determined by the anharmonic potential associated with intermolecular interactions in the spin crossover compound.

Friday, February 23, 2018

Spin ice in a nutshell

What is spin ice? What its definitive and experimental signatures?

A good place to start is the lucid discussion by Roderich Moessner and Art Ramirez in a 2006 article on Geometrical Frustration. They emphasise two organising principles: local constraints on neigbouring spins and the emergence of new entities such as gauge fields.

First, let's discuss the "ice" bit since this involves some beautiful chemistry, physics, statistical mechanics, and history. In the solid phase of water at atmospheric pressure (ice Ih) the water molecules form a hexagonal lattice, with the oxygen atoms located a the vertices of the lattice. The molecules interact with one another via hydrogen bonds.


Now the key point is that there are many different ways of orienting the water molecules (arranging the protons). The only constraint is that one has to have two protons covalently bonded to the oxygen and two protons on next-nearest neighbour water molecules hydrogen bonded to the oxygen. This is known as the ice rule. Suppose we assign an Ising spin variable (+1,-1)=(in, out)  = (covalent, Hbond) to each "bond" on the lattice. Then the ice rule is that on each tetrahedron the sum of the four "spins" must be zero.

How much degeneracy is there?
There are 2^4= 16 possible spin states on a tetrahedron. But, only six (a fraction of 3/8) satisfy the ice rule. To see this, put +1 on site one, then one must put +1 on one of the other three sites, and -1 on the other two. This gives 6 = 2 x 3 options.
If one neglects the interaction between vertices, the thermodynamic entropy per tetrahedron (water molecule) is

S = k ln (3/2)

Historical asides.
This "residual" entropy in ice was observed experimentally by William Giauque in the chemistry department at Berkeley in the 1930s.
Linus Pauling explained this in 1935, even arguing it as evidence for a specific crystal structure of ice.
Pauling's picture led to the ice-type models that are very important  (from a mathematical and conceptual point of view) in classical statistical mechanics as they are exactly soluble in two dimensions.
In 1956 Phil Anderson (who else!) noted that Pauling's problem was equivalent to that of Ising spins on a pyrochlore lattice.
It was not until four decades later than an experimental realisation was observed in a magnetic material. The experimental data is shown below.


But there is much more to spin ice. The local constraints lead naturally to an emergent gauge field (a pseudo-magnetic field), analogues of "magnetic monopoles", and unusual spin correlations (algebraic correlations without criticality). I now discuss the latter as they can be viewed as a "smoking gun" of spin ice.

The "magnetic field" B satisfies the constraint Div B =0. As a result the spin correlations have a dipolar form, i.e. they have a distance and directional dependence similar to the magnetic field associated with a magnetic dipole. This means the spin correlations fall off algebraically. This is in contrast to conventional magnets where spin correlations decay exponentially, except at a critical point. Furthermore, if one plots or measures the static spin structure factor S(q) one finds "pinch points" occur in high symmetry planes. The figure below shows an experimental measurement for Holonium Titanate, taken from here.

Thursday, June 8, 2017

A lucid lecture on the last 50 years of superconductivity

At the weekly condensed matter theory cake meeting today we watched a video of a KITP blackboard talk given by Piers Coleman in 2015.
Superconducting Surprises: five decades of discovery, in both temperature and time!

It is a very nice exposition of the history and some of the key physics.

A couple of minor comments.

Organic superconductors were discovered in 1980 not 1973.

Piers claims that the difference between the thermodynamic entropy of the superconducting and metallic states (determined from integrating the temperature dependent specific heat) is related to the quantum entanglement entropy of the superconducting ground state.
The relationship between entanglement entropy (defined on a pure quantum state (at zero temperature) which is divided in two) and thermal entropies (defined for a bulk system in a mixed state at finite temperature) is an incredibly subtle and complex issue that I don't think is resolved. See for example the discussion in this paper.

What does this movie tell us about the modern university?

Last night, my wife and I watched the movie, Wit. You can watch the full movie here  (free with ads). I should warn that some of the conten...