Showing posts with label randomness. Show all posts
Showing posts with label randomness. Show all posts

Friday, July 25, 2025

Reviewing emergent computational abilities in Large Language Models

Two years ago, I wrote a post about a paper by Wei et al, Emergent Abilities of Large Language Models

Then last year, I posted about a paper Are Emergent Abilities of Large Language Models a Mirage? that criticised the first paper.

There is more to the story. The first paper has now been cited over 3,600 times. There is a helpful review of the state of the field.

Emergent Abilities in Large Language Models: A Survey

Leonardo Berti, Flavio Giorgi, Gjergji Kasneci

It begins with a discussion of what emergence is, quoting from Phil Anderson's More is Different article [which emphasised how new properties may appear when a system becomes large] and John Hopfield's Neural networks and physical systems with emergent collective computational abilities, which was the basis of his recent Nobel Prize. Hopfield stated

"Computational properties of use to biological organisms or the construction of computers can emerge as collective properties of systems having a large number of simple equivalent components (or neurons)."

Berti et al. observe, "Fast forward to the LLM era, notice how Hopfield's observations encompass all the computational tasks that LLMs can perform."

They discuss emergent abilities as in-context learning, defined as the "capability to generalise from a few examples to new tasks and concepts on which they have not been directly trained."

Here, I put this review in the broader context of the role of emergence in other areas of science.

Scales. 

Simple scales that describe how large an LLM is include the amount of computation, the number of model parameters, and the size of the training dataset. More complicated measures of scale include the number of layers in a deep neural network and the complexity of the training tasks.

Berti et al. note that the emergence of new computational abilities does not just follow from increases in the simple scales but can be tied to the training process. I note that this subtlety is consistent with experience in biology. Simple scales would be the length of an amino acid chain in a protein or base pairs in a DNA molecule, the number of proteins in a cell or the number of cells in an organism. More subtle scales include the number of protein interactions in a proteome or gene networks in a cell. Deducing what the relevant scales are is non-trivial. Furthermore, as emphasised by Denis Noble and Robert Bishop, context matters, e.g., a protein may only have a specific function if it is located in a specific cell.

Novelty. 

When they become sufficiently "large", LLMs have computational abilities that they were not explicitly designed for and that "small" versions do not have. 

The emergent abilities range "from advanced reasoning and in-context learning to coding and problem-solving."

The original paper by Wei et al. listed 137 emergent abilities in an Appendix!

Berti et al. give another example.

"Chen et al. [15] introduced a novel framework called AgentVerse, designed to enable and study collaboration among multiple AI agents. Through these interactions, the framework reveals emergent behaviors such as spontaneous cooperation, competition, negotiation, and the development of innovative strategies that were not explicitly programmed."

An alternative to defining novelty in terms of a comparison of the whole to the parts is to compare properties of the whole to those of a random configuration of the system. The performance of some LLMs is near-random (e.g., random guessing) until a critical threshold is reached (e.g., in size) when the emergent ability appears.

Discontinuities.

Are there quantitative objective measures that can be used to identify the emergence of a new computational ability? Researchers are struggling to find agreed-upon metrics that show clear discontinuities. That was the essential point of Are Emergent Abilities of Large Language Models a Mirage? 

In condensed matter physics, the emergence of a new state of matter is (usually) associated with symmetry breaking and an order parameter. Figuring out what the relevant broken symmetry and the order parameter often requires brilliant insight and may even lead to a Nobel Prize (Neel, Josephson, Ginzburg, Leggett,...) A similar argument can be made with respect to the development of the Standard Model of elementary particles and gauge fields. Furthermore, the discontinuities only exist in the thermodynamic limit (i.e., in the limit of an infinite system), and there are many subtleties associated with how the data from finite-size computer simulations should be plotted to show that the system really does exhibit a phase transition.

Unpredictability.

The observation of new computational abilities in LLMs was unanticipated and surprised many people, including the designers of the specific LLMs involved. This is similar to what happens in condensed matter physics, where new states of matter have mostly been discovered by serendipity.

Some authors seem surprised that it is difficult to predict emergent abilities. "While early scaling laws provided some insight, they often fail to anticipate discontinuous leaps in performance."

Given the largely "black box" nature of LLMs, I don't find it the unpredictability surprising. It is hard for condensed matter systems, and they are much better characterised and understood.

Modular structures at the mesoscale.

Modularity is a common characteristic of emergence. In a wide range of systems, from physics to biology to economics, a key step in the development of the theory of a specific emergent phenomenon has been the identification of a mesoscale (intermediate between the micro- and macro-scales) at which modular structures emerge. These modules interact weakly with one another, and the whole system can be understood in these terms. Identification of these structures and the effective theories describing them has usually required brilliant insight. An example is the concepts of quasiparticles in quantum many-body physics, pioneered by Landau.

Berti et al. do not mention the importance of this issue. However, they do mention that "functional modules emerge naturally during training" [Ref. 7,43,81,84] and that "specialised circuits activate at certain scaling thresholds [24]".

Modularity may be related to an earlier post, Why do deep learning algorithms work so well? In the training process, a neural network rids noisy input data of extraneous details...There is a connection between the deep learning algorithm, known as the "deep belief net" of Geoffrey Hinton, and renormalisation group methods (which can be key to identifying modularity and effective interactions).

Is emergence good or bad?

Undesirable and dangerous capabilities can emerge. Those observed include deception, manipulation, exploitation, and sycophancy.

These concerns parallel discussions in economics. Libertarians, the Austrian school, and Federich Hayek tend to see the emergence as only producing socially desirable outcomes, such as the efficiency of free markets [the invisible hand of Adam Smith]. However, emergence also produces bubbles and crashes and recessions.

Resistance to control

A holy grail is the design, manipulation, and control of emergent properties. This ambitious goal is promoted in materials science, medicine, engineering, economics, public policy, business management, and social activism. However, it largely remains elusive, arguably due to the complexity and unpredictability of the systems of interest. Emergent properties of LLMs may turn out to offer similar hopes, frustrations, and disappointments. We should try, but have realistic expectations.

Toy models.

This is not discussed in the review. As I have argued before, a key to understanding a specific emergent phenomenon is the development of toy models that illustrate the phenomenon and the possible essential ingredients for it to occur. The following paper may be a step in that direction.

An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem

Yoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, Chris Mingard, Ard A. Louis

In a similar vein, another possibly relevant paper is the review

Statistical Mechanics of Deep Learning

Yasaman Bahri, Jonathan Kadmon, Jeffrey Pennington1, Sam S. Schoenholz, Jascha Sohl-Dickstein and Surya Ganguli

They considered a toy model for the error landscape for a neural network, and show that the error function for a deep neural net of depth D corresponds to the energy function for a D-spin spherical spin glass. [Section 3.2 in their paper].

Friday, January 3, 2025

Self-organised criticality and emergence in economics

A nice preprint illustrates how emergence is central to some of the biggest questions in economics and finance. Emergent phenomena occur as many economic agents interact resulting in a system with properties that the individual agents do not have.

The Self-Organized Criticality Paradigm in Economics & Finance

Jean-Philippe Bouchaud

The paper illustrates several key characteristics of emergence (novel properties, universality, unpredictability, ...) and the value of toy models in elucidating it. Furthermore, it illustrates the elusive nature of the "holy grail" of controlling emergent properties. 

The basic idea of self-organised criticality

"The seminal idea of Per Bak is to think of model parameters themselves as dynamical variables, in such a way that the system spontaneously evolves towards the critical point, or at least visits its neighbourhood frequently enough"

A key property of systems exhibiting criticality is power laws in the probability distribution of a property. This means that there are "fat tails" in the probability distribution and extreme events are much more likely than in a system with a Gaussian probability distribution.

Big questions

The two questions below are similar in that they concern the puzzle of how markets produce fluctuations that are much larger than expected when one tries to explain their behaviour in terms of the choices of individual agents.

A big question in economics

"A longstanding puzzle in business cycle analysis is that large fluctuations in aggregate economic activity sometimes arise from what appear to be relatively small impulses. For example, large swings in investment spending and output have been attributed to changes in monetary policy that had very modest effects on long-term real interest rates."

This is the "small shocks, large business cycle puzzle", a term coined by Ben Bernanke, Mark Gertler and Simon Gilchrist in a 1996 paper. It begins with the paragraph above. [Bernanke shared the 2022 Nobel Prize in Economics for his work on business cycles].

A big question in finance

The excess volatility puzzle in financial markets was identified by Robert Shiller: The volatility "is at least five times larger than it "should" be in the absence of feedback". In the views of some, this puzzle highlights the failings of the efficient market hypothesis and the rationality of investors, two foundations of neoclassical economics. [Shiller shared the 2013 Nobel Prize in Economics for this work]. 

"Asset prices frequently undergo large jumps for no particular reason, when financial economics asserts that only unexpected news can move prices. Volatility is an intermittent, scale-invariant process that resembles the velocity field in turbulent flows..." (page 2)

Emergent properties

Close to a critical point, the system is characterised by fat-tailed fluctuations and long memory correlations.

Avalanches. They allow very small perturbations to generate large disruptions.

Dragon Kings

Minsky moment

The holy grail: control of emergent properties

It would be nice to understand superconductivity well enough  to design a room-temperature superconductor. But, this pales in significance compared to the "holy grail" of being about to manage economic markets to prevent bubbles, crashes, and recessions.

Bouchaud argues that  the quest for efficiency and the necessity of resilience may be mutually incompatible. This is because markets may tend towards self-organised criticality which is characterised by fragility and unpredictability (Black swans).

The paper has the following conclusion

"the main policy consequence of fragility in socio-economic systems is that any welfare function that system operators, policy makers of regulators seek to optimize should contain a measure of the robustness of the solution to small perturbations, or to the uncertainty about parameters value.

Adding such a resilience penalty will for sure increase costs and degrade strict economic performance, but will keep the solution at a safe distance away from the cliff edge. As argued by Taleb [159], and also using a different language in Ref. [160], good policies should ideally lead to “anti-fragile” systems, i.e., systems that spontaneously improve when buffeted by large shocks."

Toy models

Toy models are key to understanding emergent phenomena. They ignore almost all details to the point that critics claim that the models are oversimplified. The modest goal of their proponents is simply to identify what ingredients may be essential for a phenomenon to occur. Bouchaud reviews several such models. All provide significant insight.

A trivial example (Section 2.1)

He considers an Ornstein-Uhlenbeck process for a system relaxing to equilibrium. As the damping rate tends to zero [κ⋆ → 0], the relaxation time and the variance of fluctuations diverge at the same rate. In other words, "in the limit of marginal stability κ⋆ →0, the system both amplifies exogenous shocks [i.e., those originating outside the system] and becomes auto-correlated over very long time scales."

The critical branching transition (Section 2.2)

The model describes diverse systems: "sand pile avalanches, brain activity, epidemic propagation, default/bankruptcy waves, word of mouth,..."

The model involves the parameter R0 which became famous during the COVID-19 pandemic. R0 is the average number of uninfected people who become infected due to contact with an infected individual. For sand piles R0 is the average number of grains that start rolling in response to a single rolling grain.

when R0 = 1 the distribution of avalanche sizes is a scale-free, power-law distribution 1/S^3/2, with infinite mean.

"most avalanches are of small size, although some can be very large. In other words, the system looks stable, but occasionally goes haywire with no apparent cause."

A generalised Lotka-Volterra model (Sections 3.3 and 4.2) 

This provides an analogue between economic production networks and ecology. Last year I reviewed recent work on this model, concerning how to understand the interplay of evolution and ecology.

A key result is how in the large N limit (i.e., a large number of interacting species/agents) qualitatively different behaviour occurs. Ecosystems and economies can collapse. 

 "any small change in the fitness of one species can have dramatic consequences on the whole system – in the present case, mass extinctions...

"most complex optimisation systems are, in a sense, fragile, as the solution to the optimisation problem is highly sensitive to the precise value of the parameters of the specific instance one wants to solve, like the Aij entries in the Lotka-Volterra model. Small changes of these parameters can completely upend the structure of the optimal state, and trigger large-scale rearrangements,..." 

Balancing stick problem (Section 3.4)

 The better one is able to stabilize the system, the more difficult it becomes to predict its future evolution! 

Propagation of production delays along the supply chain (Section 4.1)


An agent-based firm network model (Section 4.3)

This has the phase diagram shown below. The horizontal axis is the strength of forces counteracting supply/demand and profit imbalances. The vertical axis is the perishability of goods.

There are four distinct phases.

Leftmost region (a, violet): the economy collapses; 

Middle region (b, blue): the economy reaches equilibrium relatively quickly;

Right region (c, yellow): the economy is in perpetual disequilibrium, with purely endogenous fluctuations. 

The green vertical sliver (d) corresponds to a deflationary equilibrium

Phase diagrams illustrate how quantitative changes can produce qualitative differences.

Universality

The toy models considered describe emergent phenomena in diverse systems, including in fields other than economics and finance. 

Here are a few other recent papers by Bouchaud that are relevant to this discussion.

Navigating through Economic Complexity: Phase Diagrams & Parameter Sloppiness

From statistical physics to social sciences: the pitfalls of multi-disciplinarity

This includes the opening address from a workshop on "More is Different" at the College de France in 2022.

Monday, July 22, 2024

Clarity about the relationship of emergence, complexity, predictability, and universality

Emergence means different things to different people. Except, that practically everyone likes it! Or at least, likes using the word. Terms associated with emergence include novelty, unpredictability, universality, stratification, and self-organisation. We need to be clearer about what we mean by each of these terms and how they are related or unrelated. Significant progress is reported in a recent preprint.

Software in the natural world: A computational approach to hierarchical emergence

Fernando E. Rosas, Bernhard C. Geiger, Andrea I Luppi, Anil K. Seth, Daniel Polani, Michael Gastpar, Pedro A.M. Mediano

This preprint is the subject of a nice article in Quanta Magazine.

The New Math of How Large-Scale Order Emerges by Philip Ball

Ball defines emergence in terms of unpredictability. He states: 

"Loosely, the behavior of a complex system might be considered emergent if it can’t be predicted from the properties of the parts alone."

He describes the work of Rosas et al. as follows, 

"A complex system exhibits emergence, according to the new framework, by organizing itself into a hierarchy of levels that each operate independently of the details of the lower levels."

This is defining emergence in terms of universality. Rosas et al. use an analogy with software, which runs independently of the details of the hardware of the computer and does not depend on microscopic details such as electron dynamics.

There are three types of closure associated with emergence: informational, causal, and computational.

Informational closure means that to predict the dynamics of the system at the macroscale one does not need any additional  information from the microscale.

Equilibrium thermodynamics is a nice example. 

Causal closure means that the system can be controlled at the macroscale without any knowledge of lower-level information.

"Interventions we make at the macro level, such as changing the software code by typing on the keyboard, are not made more reliable by trying to alter individual electron trajectories."

"...we can use macroscopic variables like pressure and viscosity to talk about (and control) fluid flow, and knowing the positions and trajectories of individual molecules doesn’t add useful information for those purposes. And we can describe the market economy by considering companies as single entities, ignoring any details about the individuals that constitute them."

Computational closure is a more technical concept. 

"a conceptual device called the ε-(epsilon) machine. This device can exist in some finite set of states and can predict its own future state on the basis of its current one. It’s a bit like an elevator, said Rosas; an input to the machine, like pressing a button, will cause the machine to transition to a different state (floor) in a deterministic way that depends on its past history — namely, its current floor, whether it’s going up or down and which other buttons were pressed already. Of course an elevator has myriad component parts, but you don’t need to think about them. Likewise, an ε-machine is an optimal way to represent how unspecified interactions between component parts “compute” — or, one might say, cause — the machine’s future state."

Aside: epsilon-machines featured significantly in my previous post about What is a complex system? 

"Computational mechanics allows the web of interactions between a complex system’s components to be reduced to the simplest description, called its causal state."

"...for an emergent system that is computationally closed, the machines at each level can be constructed by coarse-graining the components on just the level below: They are, in the researchers’ terminology, “strongly lumpable.”"

In some sense, this may be related to the notion of quasiparticles and effective interactions in many-body physics. 

Aside: In 1962, Herbert Simon identified hierarchies as an essential feature of complex systems, both natural and artificial. A key property of a level in the hierarchy is that it is nearly decomposable into smaller units, i.e., it can be viewed as a collection of weakly interacting units. The time required for the evolution of the whole system is significantly decreased due to the hierarchical character. The construction of an artificial complex system, such as a clock, is faster and more reliable if different units are first assembled separately and then the units are brought together into the whole. Simon argues that the reduction in time scales due to modularity is why biological evolution can occur on realistic time scales.  The 1962 article is reprinted in The Sciences of the Artificial.

The paper by Rosas et al. is one of the most important ones I have encountered in the past few years. I am slowly digesting it.

The beauty of the paper that it is mathematically rigorous. All the concepts are precisely defined and the central results are actually theorems. This replaces the vagueness of most discussions of emergence, including by myself.

The paper has helpful figures and considers concrete examples including Ehrenfest's Urn, an Ising model with Glauber dynamics, and a Hopfield neural network model.

I thank Gerard Milburn for bringing the Quanta article to my attention.

Friday, January 5, 2024

Certain benefits of Bayes

Best wishes for the New Year! One thing I hope to achieve this year is an actual understanding of things "Bayesian".

I am particularly interested because it gives a way to be more quantitative and precise about some of the intuitions that I use in science. For example, I tend to be skeptical of new experimental results (often hyped) that claim to go against well-established theories, regardless of how good the "statistics" of the touted result.

In this vein, Phil Anderson argued that Bayesian methods should have been used to rule out the significance of "discoveries" such as the 10 keV neutrino and the fifth force. In 1992 he wrote a Physics Today column on the subject.

An interesting metric for mathematical formula is the ratio of profound and wide implications to the simplicity of the formula and its derivation. I suspect that Bayes' formula for conditional probabilities would win first place!

P(A|B) denotes the probability of A given B. 

The proof takes about two lines. If you multiply both sides of the equation about by P(B) the identity holds because both sides of the equation are just different ways of writing P(A and B).

My first attempt to understand the applications and implications of Bayes was reading the relevant sections in Phil Nelson's beautiful book, Physical Models of Living Systems. There is a helpful section entitled, "Bayes formula provides a consistent approach to upgrading our degree of belief in light of new data."

More recently, I found this wonderful and short video very helpful, as it clearly defines terms, uses graphical representations, and gives some concrete examples.

 

A Bayesian perspective highlights the importance of reporting negative results and is the basis of a seminal paper

Why Most Published Research Findings Are False by John P. A. Ioannidis

A measure of the profundity of Bayes is that the Stanford Encyclopedia of Philosophy has two articles on the topic

Bayes Theorem

Bayesian Epistemology



Monday, June 26, 2023

What is really fundamental in science?

What do we mean when we say something in science is fundamental? When is an entity or a theory more fundamental or less fundamental than something else? For example, are quarks and leptons more fundamental than atoms? Is statistical mechanics more fundamental than thermodynamics? Is physics more fundamental than chemistry or biology? In a fractional quantum Hall state, are electrons or the fractionally charged quasiparticles more fundamental?

Answers depend on who you ask. Physicists such as Phil Anderson, Steven Weinberg, Bob Laughlin, Richard Feynman, Frank Wilczek, and Albert Einstein have different views.

In 2017-8, the Foundational Questions Institute (FQXi) held an essay contest to address the question, “What is Fundamental?” Of the 200 entries, 15 prize-winning essays have been published in a single volume. The editors give a nice overview in the Introduction.

This post is mostly about the essay, Fundamental? of the first prize winner, Emily Adlam, a philosopher of physics. She contrasts two provocative statements.

Fundamental means we have won. The job is done and we can all go home.

Fundamental means we have lost. Fundamental is an admission of defeat.

This raises the question of whether being fundamental is objective or subjective.

Examples are given from scientific history to argue that what is considered to be fundamental has changed with time. The reductionism has led to the drive to explain everything in terms of smaller and smaller entities, that are deemed 'more fundamental". But we find that smaller does not always mean simpler.

Perhaps we should ask what needs explaining and what constitutes a scientific explanation. For example, Adlam asks whether explaining the fact that the initial state of the universe had a low entropy [the "past hypothesis"] is really possible or should be an important goal.

She draws on the issue of the distinction between objective and subjective probabilities. Probabilities in statistical mechanics are subjective: they are a statement about our own ignorance about the details of the motion of individual atoms and not any underlying randomness in nature. In contrast, probabilities in quantum theory reflect objective chance.

as realists about science we must surely maintain that there is a need for science to explain the existence of the sorts of regularities that allow us to make reliable predictions... but there is no similarly pressing need to explain why these regularities take some particular form rather than another. Yet our paradigmatic mechanical explanations do not seem to be capable of explaining the regularity without also explaining the form, and so increasingly in modern physics we find ourselves unable to explain either. 

It is in this context that we naturally turn to objective chance. The claim that quantum particles just have some sort of fundamental inbuilt tendency to turn out to be spin up on some proportion of measurements and spin down on some proportion of measurements does indeed look like an attempt to explain a regularity (the fact that measurements on quantum particles exhibit predictable statistics) without explaining the specific form (the particular sequence of results obtained in any given set of experiments). But given the problematic status of objective chance, this sort of nonexplanation is not really much better than simply refraining from explanation at all. 

Why is it that objective chances seem to be the only thing we have in our arsenal when it comes to explaining regularities without explaining their specific form? It seems likely that part of the problem is the reductionism that still dominates the thinking of most of those who consider themselves realists about science

In summary, (according to the Editors) Adlam argues that "science should be able to explain the existence of the sorts of regularities that allow us to make reliable predictions. But this does not necessarily mean that it must also explain why these regularities take some particular form." 

we are in dire need of another paradigm shift. And this time, instead of simply changing our attitudes about what sorts of things require explanation, we may have to change our attitudes about what counts as an explanation in the first place. 

Here, she is arguing that what is fundamental is subjective, being a matter of values and taste.

In our standard scientific thinking the fundamental is elided with ultimate truth: getting to grips with the fundamental is the promised land, the endgame of science. 

She then raises questions about the vision and hopes of scientific reductionists. 

In this spirit, the original hope of the reductionists was that things would get simpler as we got further down, and eventually we would be left with an ontology so simple that it would seem reasonable to regard this ontology as truly fundamental and to demand no further explanation. 

But the reductionist vision seems increasingly to have failed. 

When we theorise beyond the standard model [BSM] we usually find it necessary to expand the ontology still more: witness the extra dimensions required to make string theory mathematically consistent.

It is not just strings. Peter Woit has emphasised how BSM theories, such as supersymmetry, introduce many more particles and parameters.

... the messiness deep down is a sign that the universe works not ‘bottom-up’ but rather ‘top-down,’ ... in many cases, things get simpler as we go further up.

Our best current theories are renormalisable, meaning that many different possible variants on the underlying microscopic physics all give rise to the same macroscopic physical theory, known as an infrared fixed point. This is usually glossed as providing an explanation of why it is that we can do sensible macroscopic physics even without having detailed knowledge of the underlying microscopic theories. 

For example, elasticity theory, thermodynamics and fluid dynamics all work without knowing anything about atoms, statistical mechanics, and quantum theory.

But one might argue that this is getting things the wrong way round: the laws of nature don’t start with little pieces and build the universe from the bottom up, rather they apply simple macroscopic constraints to the universe as a whole and work out what needs to happen on a more fine-grained level in order to satisfy these constraints.

This is rather reminiscent of Laughlin's views about what is fundamental.

Finally, I mention two other essays that I look forward to reading as I think they make particularly pertinent points.

Marc Séguin (Chap. 6) distinguishes "between epistemological fundamentality (the fundamentality of our scientific theories) and ontological fundamentality (the fundamentality of the world itself, irrespective of our description of it)."

"In Chap. 12, Gregory Derry argues that a fundamental explanatory structure should have four key attributes: irreducibility, generality, commensurability, and fertility."

[Quotes are from the Introduction by the Editors].

Some would argue that the Standard Model is fundamental, at least on some level. But it involves 19 parameters that have to be fixed from experiment. Related questions about the Fundamental Constants, have been explored in a 2007 paper by Frank Wilczek.

Again, I thank Peter Evans for bringing this volume to my attention.

Monday, August 22, 2022

Hysteresis, hype, niches, nudges and social change

The world is a mess. Most people want a better world. Sometimes nothing changes. Sometimes things change incredibly rapidly. Sometimes changes are positive. Other times the change is negative. Often this change is unanticipated, even by experts who have been studying the relevant topic for decades. Wicked problems are things that seem to be incredibly resilient to change. Examples of rapid changes that were (largely) positive and unanticipated were the peaceful collapse of the former Soviet empire, smoking in public becoming taboo, and increased public concern about climate change. Examples of negative changes include the rise of Trumpism, misinformation on social media, and the global financial crisis of 2008.

Many people in government, public policy, NGOs, and social activists want to implement policies and take actions that will produce outcomes that (they believe) are positive. Here I discuss some basic but very important insights from "social physics", such as discussed in my previous two posts.

Suppose the system of interest can be modelled by some type of Ising model where the pseudospin corresponds to two choices (good and bad) for each agent in the system. The policy maker wants to change something such as increase the incentive for agents to make the "good" choice. There are two qualitatively different possible behaviours and they are shown in the Figure below (taken from Bouchaud). 

The vertical axis is the "magnetisation", i.e, the fraction of agents who make the good choice. The horizontal axis is the "external field", i.e, the level of incentive provided for agents to make the good choice. 


Case I. Smooth curve (blue). This occurs when the interaction between agents is weaker than some threshold strength. Suppose that a small but not insignificant minority of agents are already making the good choice and then incentive is increased slightly. If one is near the steep part of the blue curve then this "nudge" can produce a desired outcome for the society.

Case II. Discontinuous curve (red). This occurs when the interaction between agents is greater than some threshold strength. People's choices are influenced more by their friends than by what the government or an NGO is telling them to do. Then one has to provided very large incentives to get a change in agent choice, far beyond the incentive required for a single isolated agent. The system is stuck in a state that is not good for the society as a whole. It is a metastable state, as shown in the figure below.

On the other hand, if the "polarisation field" is sitting near a critical value (5 in the figure, a tipping point), then a "nudge" can lead to a dramatic change for good. 

I think there are important implications for social activists of all stripes. Realistic expectations are key.

1. Don't expect even the best-designed and well-intentioned policy or action to necessarily have the impact you hope for.

2. Be sceptical about hype and ideology. In the public space there are a lot of claims, whether from political parties, pundits, or NGOs, that if we just do X (change this law, donate money, do what my book says, ...) then the good Y will inevitably follow.

The problem with unrealistic expectations is that they lead to disappointment, disillusionment, and burnout. People give up. Then the next fad or "silver bullet" comes along...

Inspired by a rugged landscape perspective, a better and more sustainable approach is that of learning and adaptation. One identifies what one thinks the best "nudge" is, tries something, evaluates the effect, adapts, and tries out some new ideas. One does not claim or expect the first few iterations to produce a significant desired effect. Here, somewhat "random" sampling of the landscape may help. Here a diversity of perspectives and methods can play a positive role. A more concrete version of this argument is in a paper concerned with public health initiatives. Rugged landscapes: complexity and implementation science, by Joseph T. Ornstein, Ross A. Hammond, Margaret Padek, Stephanie Mazzucca & Ross C. Brownson 

Postscript. After posting this I remember reading a recent article in The Economist pointing out how nudges often do not work.

Evidence for behavioural interventions looks increasingly shaky 
The academic literature is plagued by publication bias 

It references three recent Letters in PNAS, including this one, that come to the opposite conclusion to an earlier PNAS paper.
Stephanie Mertens, Mario Herberz, Ulf J. J. Hahnel, and Tobias Brosch

Friday, August 12, 2022

Sociological insights from statistical physics

Condensed matter physics and sociology are both about emergence. Phenomena in sociology that are intellectually fascinating and important for public policy often involve qualitative change, tipping points, and collective effects. One example is how social networks influence individual choices, such as whether or not to get vaccinated. In my previous post, I briefly introduced some Ising-type models that allow the investigation of fundamental questions in sociology. The main idea is to include heterogeneities and interactions in models of decision. 

What follows is drawn from Sections 2 and 3 of the following paper from the Journal of Statistical Physics. 

Crises and Collective Socio-Economic Phenomena: Simple Models and Challenges by Jean-Philippe Bouchaud

Bouchaud first considers a homogeneous population which reaches an equilibrium state. This is then described by an Ising model with an interaction (between agents) J, in an external field, F that describes the incentive for the agents to make one of the choices. The state of the model (in the mean-field approximation) is then found by solving the Curie-Weiss equation. In the sociological context, this was first derived by Weidlich and in the economic context re-derived by Brock and Durlauf.  (Aside: The latter paper is in one of the "top-five" economic journals, was published five years after submission, and has been cited more than 2000 times.)

As first noted by Weidlich, a spontaneous “polarization” of the population occurs in the low noise regime β>β c , i.e. [the average equilibrium value of S_z] ϕ ∗≠1/2 even in the absence of any individually preferred choice (i.e. F=0). When F≠0, one of the two equilibria is exponentially more probable than the other, and in principle the population should be locked into the most likely one: ϕ ∗>1/2 whenever F>0 and ϕ ∗<1/2 whenever F<0.

Unfortunately, the equilibrium analysis is not sufficient to draw such an optimistic conclusion. A more detailed analysis of the dynamics is needed, which reveals that the time needed to reach equilibrium is exponentially large in the number of agents, and as noted by Keynes, "in the long run, we are all dead." This situation is well-known to physicists, but is perhaps not so well appreciated in other circles—for example, it is not discussed by Brock and Durlauf.

Bouchaud then discusses the meta-stability associated with the two possible polarisations, as occurs in a first-order phase transition. From a non-equilibrium dynamical analysis, based on a Langevin equation, 

one finds that the time τ needed for the system, starting around ϕ=0, to reach ϕ ∗≈1 is given by: 𝜏 ∝ exp[𝐴𝑁(1−𝐹/𝐽)], where A is a numerical factor. This means that whenever 0<F<J, the system should really be in the socially good minimum ϕ ∗≈1, but the time to reach it is exponentially large in the population size.  The important point about this formula is the presence of the factor N(1−F/J) in the exponential.

In other words, it has no chance of ever getting there on its own for large populations. Only when F reaches J, i.e. when the adoption cost C becomes zero will the population be convinced to shift to the socially optimal equilibrium...

This is very different from the standard model of innovation diffusion, based on a simple differential equation proposed by Bass in 1969 [cited more than 10,000 times].

In physics, the existence of mutually inaccessible minima with different potentials is a pathology of mean-field models that disappears when the interaction is short-ranged. In this case, the transition proceeds through “nucleation”, i.e. droplets of the good minimum appear in space and then grow by flipping spins at the boundaries. 

This suggests an interesting policy solution when social pressure resists the adoption of a beneficial practice or product: subsidize the cost locally, or make the change compulsory there, so that adoption takes place in localized spots from which it will invade the whole population. The very same social pressure that was preventing the change will make it happen as soon as it is initiated somewhere.

This analysis provides concepts to understand wicked problems. Societies get "trapped" in situations that are not for the common good and outside interventions, such as providing incentives for individuals to make better choices, have little impact.

In the next post, I hope to discuss the role of heterogeneity (i.e. the role of a random field in the Ising model). A seminal paper published in the American Journal of Sociology in 1978 is Threshold models of collective behavior  by Mark Granovetter. It has been cited more than 6000 times. The central idea is how changes in heterogeneity can induce a transition between two different collective states.

Aside: The famous Keynes quote was in his 1923 publication, The Tract on Monetary Reform. The fuller quote is “But this long run is a misleading guide to current affairs. In the long run we are all dead. Economists set themselves too easy, too useless a task, if in tempestuous seasons they can only tell us, that when the storm is long past, the ocean is flat again.”

Wednesday, August 3, 2022

Models for collective social phenomena

World news is full of dramatic and unexpected events in politics and economics, from stock market crashes to the rapid rise of extreme political parties. Trust in an institution can evaporate overnight.

The world is plagued by "wicked problems" (corruption, belief in conspiracy theories, poverty, ...) that resist a solution even when considerable resources (money, personnel, expertise, government policy, incentives, social activism) are devoted to addressing the problem. 

Here I introduce some ideas and models that are helpful for efforts to understand these emergent phenomena. Besides rapid change and discontinuities, other relevant properties include herding, trending, tipping points, and resilient equilibria. Some cultural traits or habits are incredibly persistent, even when they are damaging to a community. 

I now consider some key elements for minimal models of these phenomena: discrete choices, utility, incentives, noise, social interactions, and heterogeneity.

Discrete choices

The system consists of N agents {i} who make individual choices. Examples of binary choices are whether or not to buy a particular product, vote for a political candidate, believe a conspiracy theory, accept bribes, get vaccinated, or join a riot. For binary choices, the state of each agent is modelled by an "Ising spin", S_i = +1 or -1. 

Utility

This is the function each agent wants to maximise; what they think they will gain or lose by their decision. This could be happiness, health, ease of life, money, or pleasure.  The utility U_i will depend on the incentives provided to make a particular choice, the personal inclination of the agent, and possibly the state of other agents.

Personal inclination

Let f_i be a number representing the tendency for agent i to choose S_1=+1. 

Incentives

All individuals make their decision based on the incentives offered. Knowledge of incentives is informed by public information.  This incentive F(t) may change with time. For example, the price of a product may decrease due to an advance in technology or a government may run an advertising program for a public health initiative.

Noise

No agent has access to perfect information in order to make their decision. This uncertainty can be modelled by a parameter beta, which increases with decreasing noise. According to the log-it rule the probability that of a particular decision is

1/beta is the analogue of temperature in statistical mechanics and this probability function is the Fermi-Dirac probability distribution! 

Social interactions

No human is an island. Social pressure and imitation play a role in making choices. Even the most "independent-minded" individual makes decisions that are influenced somewhat by the decisions of others they interact with. These "neighbours" may be friends, newspaper columnists, relatives, advertisers, or participants in an internet forum. The utility for an individual may depend on the choices of others. The interaction parameter J_ij is the strength of the influence of agent j on agent i.

Heterogeneity

Everyone is different. People have different sensitivities to different incentives. This diversity reflects different personalities, values, and life circumstances. This heterogeneity can be modelled by assigning a probability distribution rho(f_i).

Putting all the ideas above together the utility function for agent i is the following.


This means that the minimal model to investigate is a Random Field Ising model. It exhibits rich phenomena, many of which are similar to the social phenomena that were mentioned at the beginning of the post. Later posts will explore this.

The discussion above is drawn from a nice paper published in the Journal of Statistical Physics in 2013.

Crises and Collective Socio-Economic Phenomena: Simple Models and Challenges by Jean-Philippe Bouchaud.

What does this movie tell us about the modern university?

Last night, my wife and I watched the movie, Wit. You can watch the full movie here  (free with ads). I should warn that some of the conten...