FW26, William Gates Building
Shannon abstracted a code as probability over symbols. The analogous move: an agent transports probability mass from \(p\) to \(q\).
Optimal intelligence, in this module’s voice, is a protocol under an entropic budget.
Thermodynamics limits mechanical engines
Information theory limits information engines
Same kind of fundamental constraint
What is the thermodynamic cost?
Erasing 1 bit requires: \(Q \geq k_B T \log 2\)
At room temperature: \(\sim 3 \times 10^{-21}\) Joules/bit
Converging perspectives on intelligence:
Unified core: Intelligence as optimal information processing
Implications:
Residual outcome uncertainty is disturbance entropy minus information the regulator has about the disturbance.
One shot, not an MDP: choose a policy \(\pi(a\mid s)\) to minimise outcome entropy.
Randomising between outcome-distinct actions cannot minimise \(H(Z)\).
Among \(H(Z)\)-minimisers there exists a deterministic policy: \(H(A\mid S)=0\).
The theorem is weaker than the popular slogan.
IB gloss, not the theorem: keep only distinctions that matter for action.
Apply three no-gos and name which geometry you are using as the prescription.
{no_gos = [‘Landauer,’ ‘Crooks L^2/tau,’ ‘I+H=C,’ ‘human bandwidth’] prescriptions = [‘Boltzmann/MaxEnt p,’ ‘Crooks geodesic,’ ‘Wasserstein plan,’ ‘Schrodinger bridge’]
Chapters 1–2; §4.2 (Sinkhorn, optional) of Peyré and Cuturi (2019)
Chapter 14 and Sections 22.2–22.3 of Welling et al. (2026)
the whole paper of Crooks (2007)
§§3–4 (optional) of Cuturi (2013)
the Good Regulator Theorem of Conant and Ashby (1970)
requisite variety; Shannon form of Ashby (1956)
Chapter 1 of Lawrence (2024)