Week 2: Boltzmann, Free Energy, and Entropy
[jupyter][google colab][reveal][edit]
Abstract:
Introduce the Gibbs–Boltzmann distribution by name and formula, then account for free energy \(F = U - TS\). Name the thermodynamic bath, Schottky’s anomaly, and finite-time dissipation. The MaxEnt derivation of the Boltzmann weights waits for week 5. Worksheet 1 is due; Quiz 1 opens the session.
Worksheet 1 is due at the start of this session. Quiz 1 occupies the first ten minutes: probability, elementary entropy, and Week 1 seeds. Then 110 minutes on the Gibbs–Boltzmann form, free energy, bath, Schottky, and finite time. Do not derive the occupation from MaxEnt today — that is week 5.
This Session
Time plan (120 minutes)
| Minutes | Block |
|---|---|
| 0–10 | Quiz 1 (Moodle; probability / entropy foundations / Week 1 seeds) |
| 10–20 | Collect Worksheet 1; recap seed \(p_i \propto e^{-\beta E_i}\) |
| 20–55 | Gibbs–Boltzmann form and names; \(\log Z\) as CGF; Hopfield / BM colour |
| 55–65 | Break |
| 65–95 | Free-energy decomposition \(F=U-TS\) |
| 95–120 | Bath; Schottky; finite-time naming |
Quiz 1
Invigilated Moodle, own device, no notes, no network, no LLMs. Ten questions on last week’s probability and entropy review and the Boltzmann / entropy seeds pressed in Worksheet 1.
From the Seed to the Formula
Last week stated \(p_i = e^{-\beta E_i}/Z\). Today name that occupation, read \(Z\) as a generating function, and account for free energy. The MaxEnt derivation waits for week 5. Motivations (perpetual motion, bandwidth) remain the frame.
Perpetual Motion and Superintelligence
Imagine in 1925 a world where the automobile is already transforming society, but big promises are being made for things to come. The stock market is soaring, the 1918 pandemic is forgotten. And every major automobile manufacturer is investing heavily on the promise they will each be the first to produce a car that needs no fuel. A perpetual motion machine.
Well, of course that didn’t happen. But I sometimes wonder if what we’re seeing today 100 years later is the modern equivalent of that. In 2026 billions are being invested in promises of superintelligence and artificial general intelligence that will transform everything.
We know why perpetual motion is impossible: the second law of thermodynamics tells us that entropy always increases. So we can’t have motion without entropy production. No matter how clever the design, you cannot extract energy from nothing, and you cannot create a closed system that does useful work indefinitely without an external energy source.
How might we make an equivalent statement for the bizarre claims around superintelligence? Some inspiration comes from Maxwell’s demon, an “intelligent” entity which operates against the laws of thermodynamics. The inspiration comes because the demon suggests that for the second law to hold there must be a relationship between the demon’s decisions and thermodynamic entropy.
One of the resolutions comes from Landauer’s principle, the notion that erasure of information requires heat dissipation. This suggests there are fundamental information-theoretic constraints on intelligent systems, just as there are thermodynamic constraints on engines.
I’ve no doubt that AI technologies will transform our world just as much as the automobile has. But I also have no doubt that the promise of unconstrained superintelligence is just as silly as the promise of perpetual motion.
Carnot and Clausius
The course follows a historical thread as well as a mathematical one. Sadi Carnot (1796–1832) asked, in 1824, what limits the efficiency of a heat engine. Rudolf Clausius (1822–1888) built on Carnot and Kelvin to state the second law of thermodynamics in several equivalent forms, and in 1865 he coined the name entropy for the state function that tracks irreversibility. Boltzmann and Gibbs, later in the same century, gave the microscopic count behind Clausius’s macroscopic \(S\). Shannon and Jaynes, in the twentieth century, reuse the same functional form with different operational readings. Keep that chain in view: engines first, then entropy as a named quantity, then statistics, then information.
Clausius did not give the Boltzmann distribution. He gave the macroscopic balance that any prescription must respect. When we write \(p_i \propto e^{-\beta E_i}/Z\) today and derive it from MaxEnt in week 5, read it as the statistical answer to a constraint Clausius already framed: fixed mean energy, maximum entropy, no perpetual motion.
Information, entropy and intelligence course notebook setup
We install some bespoke code for creating and saving plots as well as loading data sets.
import importlib.utilcmd = install_command('pods')%system {cmd}cmd = install_command('mlai')%system {cmd}The Gibbs–Boltzmann Distribution
Week 1 seeded \(p_i\propto e^{-\beta E_i}\). Today we name the object and fix the formula. The MaxEnt derivation that earns those weights waits for week 5: deriving the occupation from Lagrange multipliers before Shannon \(H\) and Jaynes would put the cart before the horse.
Write the equilibrium occupation of a discrete system with energies \(\{E_i\}\) as \[ p_i = \frac{e^{-\beta E_i}}{Z(\beta)}, \qquad Z(\beta)=\sum_j e^{-\beta E_j}. \] Physicists call \(p_i\) the Boltzmann distribution (or Boltzmann weights) and, for a system exchanging energy with a bath at fixed \(T\), the Gibbs distribution or canonical ensemble. The three names point at the same formula. The normalisation \(Z(\beta)\) is the partition function. Its logarithm \(\log Z(\beta)\) is the cumulant generating function for the energy under this exponential family: derivatives of \(\log Z\) recover the mean energy, the variance (heat capacity, up to factors of \(\beta\)), and higher cumulants. That generating-function reading is why weeks 3 and 5 treat \(Z\) as more than a normalisation constant.
Figure: Occupation of a two-state system as coldness increases. At low \(\beta\) both states are populated; at high \(\beta\) the ground state dominates.
The Gibbs–Boltzmann occupation is not only a statement about gases and magnets. Hopfield networks (Hopfield, 1982) assign an energy to every binary configuration of a recurrent net and treat recall as a descent toward low-energy states — equilibrium statistics are again Gibbs. Ackley, Hinton and Sejnowski (Ackley et al., 1985) made the weights of that energy learnable: a Boltzmann machine is an undirected model whose distribution over configurations is exactly \(p(s)\propto e^{-E(s)/T}\). The 2024 Nobel Prize in Physics, awarded to John Hopfield and Geoffrey Hinton, recognised that physical-systems reading of computation and learning. Name the lineage here so the formula does not feel confined to nineteenth-century heat baths; do not divert the lecture into training algorithms. Week 4 will draw the physical one-bit substrate — a thermal particle in an equal-depth double well — when Landauer prices erasure; week 5 meets the smallest Boltzmann machine as a two-spin MaxEnt model with a correlation constraint.
Coldness and Temperature
The Boltzmann occupation can be written in two ways, \[ p_i = \frac{e^{-E_i/k_B T}}{Z} = \frac{e^{-\beta E_i}}{Z}, \qquad \beta = \frac{1}{k_B T}. \] The mathematics is the same, but the emphasis is different.
\(T\) is the variable of the bath. It is what a thermometer reports, and it is the intensive parameter in the Helmholtz accounting \(F = U - TS\). In that representation energy comes first and entropy is the correction: \(TS\) is the cut the second law takes from \(U\).
\(\beta\) is coldness. It is the variable conjugate to energy. In that representation entropy comes first: you maximise \(S\) (or write a generating function in \(\beta\)) and energy is the constraint. Temperature is then a derived reading, \(T = 1/k_B\beta = (\partial S/\partial U)^{-1}\).
We will use both. Weeks 1–2 keep \(T\) in view so the bath is familiar, and they already compute in \(\beta\) because that is the natural argument of \(Z\). Week 5 makes the \(\beta\)-first order honest: the Lagrange multiplier on a mean-energy constraint is coldness, and it is the natural parameter of the exponential family. Weeks 6–7 then treat \(\beta\) as a coordinate on that family. A student who still hears \(\beta\) as ``one over the thermometer’’ will miss why the geometry is written in those coordinates.
Free Energy Decomposition
Helmholtz free energy \(F=U-TS=-kT\log Z\) accounts for what remains after the entropy takes its cut: \(U\) is total energy, \(TS\) is unavailable, \(F\) is what the bath still allows you to do.
Figure: \(U\), \(S\), and \(F\) for the Worksheet 1 three-state system. Verify \(F=-\ln Z/\beta\) numerically.
Welling, Lu and Holdijk
Only this year, Welling, Lu and Holdijk’s Generative AI and Stochastic Thermodynamics (GAIST, Welling et al. (2026)) was published. Although their focus is on free energy we will now see that free energy and entropy are two sides of the same thermodynamic coin. GAIST explores the shared mathematics of generative models and nonequilibrium thermodynamics are the same mathematics. Although the mathematics is broadly the same ur philosophy is slightly different. GAIST has the slogan free energy is all you need, if this course had a slogan it would more likely be “entropy is all you need.”
GAIST on Boltzmann and Free Energy
The opening essay, ``Why Stochastic Thermodynamics of Machine Learning?’’ (Welling et al., 2026), states the free-energy decomposition we use today: entropy measures the information we are missing; we subtract it from the energy and call what remains free energy, the energy still free to perform work. Maxwell’s demon appears there as a knowledge engine — not yet as Landauer’s cost. That cost is week 3.
Chapter 3 is the physics we need for LO1. Section 3.2 introduces the Hamiltonian, the microcanonical and canonical ensembles, and the passage from the partition function to \(U\), \(S\), and \(F\). Section 3.3 writes the first law and splits an energy change into heat and work: \[ \delta U = \int \delta\rho\,\mathrm{H} + \int \rho\,\delta\mathrm{H}. \] The first term is heat (\(\rho\) changes at fixed Hamiltonian); the second is work (the Hamiltonian changes at fixed \(\rho\)). That is the same accounting as \(F = U - TS\): \(TS\) is the cut taken by ignorance, \(F\) is what remains.
Read Chapter 3, not the whole book. Callen remains the undergraduate thermodynamics reference. GAIST does not discuss perpetual motion or the human bandwidth constraint; those motivations are ours.
Thermodynamic Bath and Schottky’s Anomaly
A thermodynamic bath fixes \(T\) and exchanges energy with a small system, justifying the canonical ensemble. Heat capacity \(C=\partial U/\partial T\) peaks when both states are equally populated — Schottky’s anomaly.
Figure: Mean energy and heat capacity of a two-state system. The Schottky peak appears when \(p_0\approx p_1\approx\frac12\).
Finite Time Costs More Than \(\Delta F\)
Quasi-static processes achieve reversible work \(W=\Delta F\). Finite-time driving dissipates additional energy beyond the equilibrium bound.
Figure: Cartoon of quasi-static versus finite-time driving between two equilibrium states.
Define This Week
Interpret later: how entropy is understood today; equilibrium versus non-equilibrium; the purely entropic Schottky reading. Define-stage answer to “How was entropy discovered?”: Carnot on engines; Clausius names entropy and states the second law (1865); Boltzmann and Gibbs give the statistical count; Shannon and Jaynes reuse \(H\) with different operational readings.
After This Lecture
Next week: Shannon entropy and the partition function as a generating function. Optional LLM probe: is free energy a constraint or a recipe?
Further Reading
Boltzmann machines (optional colour) of Ackley et al. (1985)
Hopfield networks (optional colour) of Hopfield (1982)
Why Stochastic Thermodynamics of Machine Learning? of Welling et al. (2026)
Chapter 3 of Welling et al. (2026)
Chapters 1–4 of Callen (1985)
Chapters 5–6 of Callen (1985)
Thanks!
For more information on these subjects and more you might want to check the following resources.
- company: Trent AI
- book: The Atomic Human
- twitter: @lawrennd
- podcast: The Talking Machines
- newspaper: Guardian Profile Page
- blog: http://inverseprobability.com