Week 4: Maxwell’s Demon and Landauer’s Principle
[jupyter][google colab][reveal][edit]
Abstract:
Maxwell’s demon and Landauer’s principle. Erasing one bit costs \(k_B T\ln 2\). Feedback work is bounded by mutual information; at molecular scale ATP synthase approximates an information engine.
No class test today. Worksheet 2 is released; due 10 November (start of lecture 5).
Define This Week
After This Lecture
Worksheet 2: Maxwell / Landauer, MaxEnt, exponential family. Due 3 November. Quiz 2 is 10 November and will use a new example.
This Session
Time plan (120 minutes)
| Minutes | Block |
|---|---|
| 0–10 | Recap Shannon / partition; release Worksheet 2 |
| 10–50 | Maxwell’s demon; membrane detailed balance and \(\Delta S\) |
| 50–60 | Break |
| 60–95 | Szilard; Feynman; Parrondo; information engines; car vs ATP synthase |
| 95–112 | Boltzmann memory (equal wells); Landauer; erasure cost |
| 112–120 | Intelligence, first cut; human bandwidth; Worksheet 2 |
Information, entropy and intelligence course notebook setup
We install some bespoke code for creating and saving plots as well as loading data sets.
import importlib.utilcmd = install_command('pods')%system {cmd}cmd = install_command('mlai')%system {cmd}Maxwell’s Demon
Maxwell wrote that his demon would act “in contradiction to the second law of thermodynamics” — the law Clausius had formulated for heat engines and irreversible processes. Szilard and Landauer do not replace Clausius; they extend the same no-go to stored outcomes and erased bits. The demon’s sorting policy remains a prescription; it cannot repeal the law Clausius stated.
Maxwell’s Demon
Maxwell’s demon is a thought experiment described by James Clerk Maxwell in his book, Theory of Heat (Maxwell, 1871) on page 308.
But if we conceive a being whose faculties are so sharpened that he can follow every molecule in its course, such a being, whose attributes are still as essentially finite as our own, would be able to do what is at present impossible to us. For we have seen that the molecules in a vessel full of air at uniform temperature are moving with velocities by no means uniform, though the mean velocity of any great number of them, arbitrarily selected, is almost exactly uniform. Now let us suppose that such a vessel is divided into two portions, A and B, by a division in which there is a small hole, and that a being, who can see the individual molecules, opens and closes this hole, so as to allow only the swifter molecules to pass from A to B, and the only the slower ones to pass from B to A. He will thus, without expenditure of work, raise the temperature of B and lower that of A, in contradiction to the second law of thermodynamics.
James Clerk Maxwell in Theory of Heat (Maxwell, 1871) page 308
He goes onto say:
This is only one of the instances in which conclusions which we have draw from our experience of bodies consisting of an immense number of molecules may be found not to be applicable to the more delicate observations and experiments which we may suppose made by one who can perceive and handle the individual molecules which we deal with only in large masses
Figure: Maxwell’s demon was designed to highlight the statistical nature of the second law of thermodynamics.
Figure: Maxwell’s Demon. The demon decides balls are either cold (blue) or hot (red) according to their velocity. Balls are allowed to pass the green membrane from right to left only if they are cold, and from left to right only if they are hot. The displayed entropy is the Shannon entropy of the velocity histogram (a coarse-grained proxy, not full thermodynamic entropy).
Maxwell’s demon allows us to connect thermodynamics with information theory (see e.g. Hosoya et al. (2015);Hosoya et al. (2011);Bub (2001);Brillouin (1951);Szilard (1929)). Landauer (1961) described a fundamental connection between information erasure and energy consumption .
Alemi and Fischer (2019)
Detailed Balance Across the Membrane
In the simulation the membrane does not move. A ball on the left is allowed through only when its speed exceeds \(v_\star\); a ball on the right is allowed through only when its speed is at most \(v_\star\). Between crossings, elastic collisions redistribute energy, so a ball that is cold now can become hot later and then cross. The question is not whether sorting begins — it does — but where the net particle and energy flows settle.
Take equal chamber areas and a 2D ideal gas (the billiard world). On each side the velocity density is Maxwellian, \[ f(\mathbf{v}\mid\beta)=\frac{\beta}{2\pi}\exp\!\left(-\tfrac{1}{2}\beta v^2\right),\qquad \beta=\frac{1}{k_B T}, \] with mean kinetic energy \(k_B T\) per ball (two quadratic degrees of freedom). The one-way effusion flux through a unit aperture is \(\Phi_{\mathrm{tot}}=n/\sqrt{2\pi\beta}\). Restricting to the hot or cold set gives \[\begin{align} \Phi^{\mathrm{hot}}(n,\beta) &= \frac{n}{\pi}\left[v_\star e^{-\beta v_\star^2/2}+\int_{v_\star}^\infty e^{-\beta v^2/2}\,\mathrm{d}v\right],\\ \Phi^{\mathrm{cold}}(n,\beta) &= \Phi_{\mathrm{tot}}(n,\beta)-\Phi^{\mathrm{hot}}(n,\beta). \end{align}\] Particle-number detailed balance across the membrane is \[ \Phi^{\mathrm{hot}}(n_L,\beta_L)=\Phi^{\mathrm{cold}}(n_R,\beta_R). \] Total ball number \(N=N_L+N_R\) and total energy \(E=N_L k_B T_L+N_R k_B T_R\) are conserved. For the simulation parameters (\(v_\star=3\), mean kinetic energy \(12.5\) in code units, equal populations \(N_L=N_R\)) this fixes the temperatures uniquely.
Figure: Left: temperatures on each side of the selective membrane from particle-flux detailed balance at equal populations (simulation energy budget). Right: entropy of the two-Maxwellian state minus the single-temperature equilibrium. The red point is the simulation threshold \(v_\star=3\).
With \(k_B=1\), equal populations, \(v_\star=3\) and total energy matching the simulation’s mean kinetic energy \(12.5\), particle-flux detailed balance gives \(T_L\approx 1.94\) and \(T_R\approx 23.1\). The ideal-gas Shannon entropy of a 2D Maxwellian chamber (position Stirling term plus velocity differential entropy) \[ S(N,\beta)=N\left[2+\log\frac{A}{N}-\log\frac{\beta}{2\pi}\right] \] falls by about \(24\) nats relative to the single-temperature state at the same \(N\) and \(E\). That drop is the thermodynamic content of the sorting you see on the canvas: hot accumulates on the right, cold on the left, and the joint distribution over \((\mathrm{side},\mathbf{v})\) is more concentrated than the unconstrained Maxwellian.
A Maxwellian is symmetric in a way the membrane is not. Conditioning on \(v>v_\star\) selects a high-energy set; conditioning on \(v\le v_\star\) selects a low-energy set. When the two number fluxes match, the energy fluxes generally do not. The true non-equilibrium steady state therefore cannot be a pair of perfect Maxwellians: the left hot tail is depleted by leakage, the right cold body is depleted by leakage the other way, and collisions continually refill those sets. The local-temperature calculation remains the right first cut — it predicts a cold left, a hot right, and a lower gas entropy — but the precise \((T_L,T_R)\) pair is an approximation to a non-Maxwellian steady state.
The selective membrane is Maxwell’s demon written as a dynamical rule. The detailed-balance calculation shows that the rule really does drive the gas toward a lower-entropy steady state. Restoring Clausius requires putting the membrane — or the memory that controls an equivalent trap door — on the thermodynamic ledger. Szilard’s engine isolates one bit of that ledger; Landauer’s principle prices erasing it. The simulation’s ``velocity-bin entropy’’ is a coarse proxy for the same story: as sorting proceeds, the displayed histogram entropy typically drops relative to the unsorted gas.
A Different Perspective
Figure: John Ellis presenting on Maxwell’s demon at the Cambridge philosophical society’s Maxwell day.
John Ellis, Maxwell Day 2024 (Cambridge). Dissipation-first: measurement scatters; the trap door is a switch in a metastable well — its position is already a physical memory degree of freedom, not a separate notebook. Information-first: the Feynman/Szilard/Landauer/Parrondo chain in this lecture. The generality question — mechanism-independent law versus mechanism-specific rebuttals — is Worksheet 2 Part B (i)(ii). Bad textbook Landauer says erasure must cost entropy because otherwise the second law fails; Bennett (1982) derives the demon resolution from logical irreversibility of erasure (phase-space compression) — the non-circular version students should cite against Ellis.
Szilard’s Engine
Figure: Szilard’s Engine
John Norton refers to Szilard’s engine as the “worst thought experiment” Norton (2018) in science.
Szilard (1929): partition, measure which side, expand from \(V_0/2\) to \(V_0\) against a thermal bath. For one molecule the work is \(k_BT\ln 2\). If the outcome is equiprobable, \(H(M)=\ln 2\) nats and \(W_{\mathrm{ext}} = k_B T H(M)\). This is the first link in the chain from Shannon bits (week 2) to Landauer’s joules (next section).
In Feynman Lectures on Computation (Hey ed., Ch. 5) Feynman uses the Szilard box with a piston on the occupied side: information about position is operational fuel. A chain of such strokes is only sustainable if you keep supplying low-entropy records or pay to erase them. In The Feynman Lectures on Physics (Feynman et al., 1963), Ch. 46, he shows that a one-bath ratchet fails because thermal kicks on the pawl destroy rectification — the same fluctuation physics Smoluchowski and Parrondo et al. cite for autonomous demons. The historical chain is Maxwell (sort by velocity), Szilard (sort by position, one bit), Landauer (pay on erase).
Parrondo, Horowitz and Sagawa (Parrondo et al., 2015) frame the thermodynamics of information as non-equilibrium thermodynamics for states updated by measurement. Shannon entropy of the microstate, multiplied by \(k_B\), is the operative entropy for isothermal processes far from equilibrium. A measurement that correlates system \(X\) with outcome \(M\) raises non-equilibrium free energy by \(k_B T I(X;M)\), so feedback can extract work bounded by that mutual information. In a cyclic Szilard engine with error-free measurement, \(I(X;M)=H(M)=\ln 2\) and \(W_{\mathrm{ext}}=k_BT\ln 2\) saturates the bound. The memory where \(M\) is stored must be physical (Landauer: metastable wells, broken ergodicity) — the equal-depth double well drawn in the next block is that claim made concrete, and one cell of an information reservoir. Over measure–feedback–reset, the mutual-information work is paid either during measurement or during erasure — the cost cannot disappear from the ledger. Stochastic thermodynamics and fluctuation theorems now reproduce Szilard and Landauer in the lab; we cite the review rather than reproducing the full formalism here.
Information Engines
This block condenses _information-game/includes/intelligence-thermodynamics-connection.md from the information-engines talk. The full talk also develops Markov blankets, generalised Jarzynski with feedback, and Ashby requisite variety; we name those in week 8. Here the point is operational: intelligence as feedback control is an information engine only if the memory and measurement bandwidth can support the mutual information the second law demands.
At human scale the energy in random thermal motion swamps any ledger we could maintain in bits. A feedback controller that tried to match \(I(X;M)\) to the work flow of a car engine would need memory and measurement bandwidth beyond physical possibility. Clausius and Carnot already state the macroscopic no-go; information thermodynamics explains why importing Szilard into a gearbox does not help.
Figure: Order-of-magnitude contrast: thermal degrees of freedom in a 70 kW engine versus total world data stock (very rough).
Figure: ATP Synthase in action.
ATP synthase synthesises ATP from ADP and phosphate using the proton-motive force. The gradient is both energy reservoir and signal about cellular state. Each proton transit is a discrete event at room temperature; the \(\gamma\) subunit rotates in steps as protons bind and release — a molecular ratchet of the kind Feynman analysed, but coupled to a chemical fuel (gradient) not a single bath. Rough accounting: \(\sim 10^4\) ATP per synaptic event, \(\sim 4\times 10^4\) protons; \(\sim 10^{14}\) synapses with sparse firing gives \(\sim 10^{18}\) protons/s brain-wide, each with several thermal degrees of freedom — petabit-per-second physical throughput distributed across vast numbers of mitochondria and synthase copies, not a single memory register. That is how life improves free-energy conversion where a car engine cannot: nanoscale composition of many information engines, not one demon on a macroscopic shaft. See the information-engines talk and Parrondo et al. (2015) for laboratory Szilárd engines; Roh et al. for cryo-EM structure of rotary proton pumps.
Landauer’s Principle
Boltzmann Memory: Equal Double Well
Parrondo, Horowitz and Sagawa insist that memory is physical: outcomes live in metastable states with broken ergodicity (Parrondo et al., 2015). Draw that claim. A thermal particle in a symmetric double well — equal depth, barrier high compared with \(k_BT\) — is one bit of Gibbs memory. The bit is which well, not an energy bias between the wells.
Equal depth is the point. If the wells differed in depth, part of the story would be energetic preference for one label. With \(E_L=E_R\) the equilibrium Gibbs occupation is uniform, \(p=\tfrac12\), and the thermodynamic cost of resetting an unknown bit is purely informational. Retention is kinetic: rare barrier crossings. That is Landauer’s and Bennett’s metastable-memory geometry in one picture.
Figure: Boltzmann memory: equal-depth double well. Store and read use two metastable basins; erase merges them. Equal depth keeps \(E_0=E_1\), so Landauer’s cost is informational rather than an energy bias between labels.
Landauer’s bound applies to the logically irreversible map that sends both labels to a standard state. If the controller already knows the bit and it sits in the target well, there is nothing to compress. If the bit is unknown and equiprobable, phase-space volume halves and the bath must take at least \(k_BT\ln 2\). Equal depth makes that statement sharp: there is no free-energy difference between \(0\) and \(1\) to confuse with the informational cost (Bennett, 1982; Landauer, 1961).
In the thermodynamics of information, an information reservoir is a subsystem that can change the entropy balance while exchanging negligible energy — ideally a sequence of energetically degenerate bits (Barato and Seifert, 2014; Parrondo et al., 2015). The equal-depth double well is one physical cell of that idealisation: flipping or randomising the bit changes Shannon entropy without an energy bias between values. A tape of such wells is the reservoir. Resetting the tape is Landauer accounting on the reservoir’s entropy. A Boltzmann machine is not this one-bit picture; it is a joint Gibbs distribution \(p(s)\propto e^{-E(s)/T}\) over many binary units with couplings. Week 2 named that lineage; week 5 builds the two-spin case. Here the point is the drawable Landauer bit that makes ``memory is physical’’ concrete.
Is Landauer’s Limit Related to Shannon’s Gaussian Channel Capacity?
Digital memory can be viewed as a communication channel through time - storing a bit is equivalent to transmitting information to a future moment. This perspective immediately suggests that we look for a connection between Landauer’s erasure principle and Shannon’s channel capacity. The connection might arise because both these systems are about maintaining reliable information against thermal noise.
The Landauer limit (Landauer, 1961) is the minimum amount of heat energy that is dissapated when a bit of information is erased. Conceptually it’s the potential energy associated with holding a bit to an identifiable single value that is differentiable from the background thermal noise (representated by temperature).
The Gaussian channel capacity (Shannon, 1948) represents how identifiable a signal \(S\), is relative to the background noise, \(N\). Here we trigger a small exploration of potential relationship between these two values.
When we store a bit in memory, we maintain a signal that can be reliably distinguished from thermal noise, just as in a communication channel. This suggests that Landauer’s limit for erasure of one bit of information, \(E_{min} = k_BT\), and Shannon’s Gaussian channel capacity, \[ C = \frac{1}{2}\log_2\left(1 + \frac{S}{N}\right), \] might be different views of the same limit.
Landauer’s limit states that erasing one bit of information requires a minimum energy of \(E_{\text{min}} = k_BT\). For a communication channel operating over time \(1/B\), the signal power \(S = EB\) and noise power \(N = k_BTB\). This gives us: \[ C = \frac{1}{2}\log_2\left(1 + \frac{S}{N}\right) = \frac{1}{2}\log_2\left(1 + \frac{E}{k_BT}\right) \] where the bandwidth B cancels out in the ratio.
When we operate at Landauer’s limit, setting \(E = k_BT\), we get a signal-to-noise ratio of exactly 1: \[ \frac{S}{N} = \frac{E}{k_BT} = 1 \] This yields a channel capacity of exactly half a bit per second, \[ C = \frac{1}{2}\log_2(2) = \frac{1}{2} \text{ bit/s} \]
The factor of 1/2 appears in Shannon’s formula because of Nyquist’s theorem - we need two samples per cycle at bandwidth B to represent a signal. The bandwidth \(B\) appears in both signal and noise power but cancels in their ratio, showing how Landauer’s energy-per-bit limit connects to Shannon’s bits-per-second capacity.
This connection suggests that Landauer’s limit may correspond to the energy needed to establish a signal-to-noise ratio sufficient to transmit one bit of information per second. The temperature \(T\) may set both the minimum energy scale for information erasure and the noise floor for information transmission.
Implications for Information Engines
This connection suggests that the fundamental limits on information processing may arise from the need to maintain signals above the thermal noise floor. Whether we’re erasing information (Landauer) or transmitting it (Shannon), we need to overcome the same fundamental noise threshold set by temperature.
This perspective suggests that both memory operations (erasure) and communication operations (transmission) are limited by the same physical principles. The temperature \(T\) emerges as a fundamental parameter that sets the scale for both energy requirements and information capacity.
The connection between Landauer’s limit and Shannon’s channel capacity is intriguing but still remains speculative. For Landauer’s original work see Landauer (1961), Bennett’s review and developments see Bennet (1982), and for a more recent overview and connection to developments in non-equilibrium thermodynamics Parrondo et al. (2015).
Landauer (1961): erasing one bit in a bath at temperature \(T\) dissipates at least \(k_B T\ln 2\). Erasure of stored outcomes restores the second law. The demon’s measurement policy is the prescription; it does not repeal the bound. Ellis agrees the demon fails but locates the cost at measurement and gating; Bennett (1982) completes the information ledger: erasure is logically irreversible, so phase-space compression costs at least \(k_B T\ln 2\) per bit — derived, not assumed to save Clausius. Today’s statement is the quasi-static bound (equality in the reversible limit). Finite-time erasure pays a dissipative penalty above \(k_B T\ln 2\); the exact distributional tool is Crooks’ fluctuation theorem, which we meet in week 6 as the bridge to thermodynamic length.
Figure: Landauer’s minimum heat dissipation per erased bit as a function of bath temperature.
kB = 1.380649e-23
def landauer_cost(T_kelvin, n_bits=1):
return n_bits * kB * T_kelvin * np.log(2)# landauer_cost(300) -> ~2.87e-21 joules per bitGAIST on Maxwell and Landauer
Section 3.3 of (Welling et al., 2026), pages 49–52, is the book’s account of today’s pair. Heat is energy stored in fluctuations we cannot couple to; work is energy stored in a few accessible macroscopic degrees of freedom. The split is subjective: it depends on the devices we can build. Maxwell’s demon is the being who can couple to the molecules, and so appears to extract work from a single bath.
The resolution is Landauer and Bennett (Bennett, 1982; Landauer, 1961): the demon must acquire, store, and erase information, and that process increases entropy by at least as much as the system lost. The modern review is Parrondo et al. (2015). GAIST does not develop Szilard’s engine, and does not yet state the \(k_B T\ln 2\) bound as a calculation. Read Landauer’s two pages for that. The finite-time correction to Landauer — you cannot saturate \(k_B T\ln 2\) in finite time — is Section 22.3, and waits until week 8.
Information and Intelligence: First Cut
Landauer: you cannot erase a bit for less than \(k_B T\ln 2\). Embodiment: you cannot communicate at machine bandwidth. The Atomic Human (Lawrence, 2024) takes the second as the defining constraint on human intelligence — we are locked in relative to the machine, and we overcome it by modelling other minds, not by opening a wider channel. Bauby is the extreme of that fence: Shannon lets us count how locked in he is. Probability’s job, on the human side, is to say how that narrow budget is spent. The full intelligence question is week 8.
Further Reading
Szilard engine and demon resolution (logical irreversibility of erasure) of Bennett (1982)
introduction and Szilárd engine section of Parrondo et al. (2015)
physical memory; metastable states of Parrondo et al. (2015)
erasure cost of Landauer (1961)
information reservoirs (optional) of Barato and Seifert (2014)
the whole paper of Landauer (1961)
Section 3.3, pages 49–52 of Welling et al. (2026)
Thanks!
For more information on these subjects and more you might want to check the following resources.
- company: Trent AI
- book: The Atomic Human
- twitter: @lawrennd
- podcast: The Talking Machines
- newspaper: Guardian Profile Page
- blog: http://inverseprobability.com