Probability Transport and Limits on Intelligence

Neil D. Lawrence

FW26, William Gates Building

This Session

  • Probability transport; three geometries
  • Entropic Good Regulator: \(H(A\mid S)=0\), with caveats
  • Superintelligence as perpetual motion; close Week 1

Information, entropy and intelligence course notebook setup

Agency as Transport

Erwin Schrödinger

Schrödinger’s Bridge and Optimal Information Transport

  • Schrödinger’s bridge problem:
    • Find most likely stochastic process between two distributions
    • Minimum relative entropy solution
    • Optimal transport of probability mass
    • Discrete MaxEnt coupling: Sinkhorn / IPF

Intelligence as Optimal Control

Shannon abstracted a code as probability over symbols. The analogous move: an agent transports probability mass from \(p\) to \(q\).

  • Wasserstein: minimum ground-cost transport
  • Schrödinger bridge: maximum-entropy stochastic interpolation
  • Sinkhorn: discrete algorithm; not a fourth geometry

Three Geometries

Optimal intelligence, in this module’s voice, is a protocol under an entropic budget.

  • Fisher–Rao / Crooks — near-equilibrium; \(\langle W_{\mathrm{ex}}\rangle\ge\mathcal{L}^2/\tau\)
  • Wasserstein — minimum ground-cost on sample space
  • Schrödinger bridge — max-entropy interpolation; Sinkhorn computes it
  • No-gos: Landauer, (^2/), (I+H=C)
  • Prescriptions: Crooks geodesic; Wasserstein; Schrödinger
  • Do not collapse the three
  • Ch.~14 and 22: Schrödinger bridge and \(W_2\)
  • Discrete algorithm: IPF / Sinkhorn
  • Speed limit there is \(W_2^2/(T\tau)\), not \(\mathcal{L}^2/\tau\)
  • Finite-time Landauer is Section 22.3

Limits on Intelligence

Information-Theoretic Limits on Intelligence

  • Thermodynamics limits mechanical engines

  • Information theory limits information engines

Same kind of fundamental constraint

What Intelligent Systems Must Do

  • Acquire information (sensing)
  • Store information (memory)
  • Process information (computation)
  • Erase information (memory mgmt)
  • Act on information (output)

What is the thermodynamic cost?

Landauer’s Principle

Erasing 1 bit requires: \(Q \geq k_B T \log 2\)

  • Not engineering limitation
  • Fundamental thermodynamic bound
  • Entropy must go somewhere

At room temperature: \(\sim 3 \times 10^{-21}\) Joules/bit

A Unified View of Intelligence Through Information

  • Converging perspectives on intelligence:

    • Efficient entropy reduction (Entropy Game)
    • Energy-efficient information processing (Information Engines)
    • Path optimization in information space (Least Action)
    • Optimal probability transport (Schrödinger’s Bridge)
  • Unified core: Intelligence as optimal information processing

  • Implications:

    • Fundamental limits on intelligence
    • New metrics for AI systems
    • Principled approach to cognitive modeling
    • Information-theoretic approaches to learning

Research Directions

  • Open questions:
    • Information-theoretic intelligence metrics
    • Physical limits of intelligent systems
    • Connections to quantum information theory
    • Practical algorithms based on these principles
    • Biological implementations of information engines
  • Applications:
    • Active learning systems
    • Energy-efficient AI
    • Robust decision-making under uncertainty
    • Cognitive architectures

Purely Entropic Good Regulator

Residual outcome uncertainty is disturbance entropy minus information the regulator has about the disturbance.

  • Disturbance \(D\), response \(R\), essential variable \(E\)
  • \(H(E)\ge H(D)-I(D;R)\)
  • Useless variety appears as small \(I(D;R)\)

One shot, not an MDP: choose a policy \(\pi(a\mid s)\) to minimise outcome entropy.

  • State \(S\), action \(A\), outcome \(Z\sim p(z\mid s,a)\)
  • Objective: \(\min_\pi H(Z)\)
  • Model (weak): \(H(A\mid S)=0\), i.e. \(A=h(S)\)

Randomising between outcome-distinct actions cannot minimise \(H(Z)\).

  • At state \(s\), mix \(a_1,a_2\) with \(\psi(s,a_1)=z_1\ne z_2=\psi(s,a_2)\)
  • Transfer mass \(\alpha\): \((u,v)\mapsto(u-\alpha,v+\alpha)\)
  • \(g''(\alpha)<0\): interior mixture is not a minimum

Among \(H(Z)\)-minimisers there exists a deterministic policy: \(H(A\mid S)=0\).

  • Pick one action per \(s\) among those sharing \(z^\ast(s)\)
  • \(A=h(S)\) leaves \(p(z)\), hence \(H(Z)\), unchanged
  • Stochastic channel: same existence via extreme points

The theorem is weaker than the popular slogan.

  • Weak model: constant \(A=a_0\) also has \(H(A\mid S)=0\)
  • Low \(H(Z)\) means predictable, not desirable
  • Catastrophe with \(H(Z)=0\) is ``optimal’’ under entropy alone

IB gloss, not the theorem: keep only distinctions that matter for action.

  • \(\min I(S;A)\) subject to \(H(Z)\le\epsilon\) — course reading
  • Deterministic \(A=h(S)\) gives \(I(S;A)=H(A)\)
  • Do not attribute that optimisation to Conant and Ashby

Apply three no-gos and name which geometry you are using as the prescription.

  • Historical thread: Carnot \(\to\) Clausius \(\to\) Boltzmann \(\to\) Shannon \(\to\) Landauer \(\to\) here
  • Landauer; Crooks \(\mathcal{L}^2/\tau\); \(I+H=C\); DPI
  • GRT: exists optimal \(\pi\) with \(H(A\mid S)=0\) (earned, with caveats)
  • Human bandwidth: \(\sim 100\) bits/s — The Atomic Human

{no_gos = [‘Landauer,’ ‘Crooks L^2/tau,’ ‘I+H=C,’ ‘human bandwidth’] prescriptions = [‘Boltzmann/MaxEnt p,’ ‘Crooks geodesic,’ ‘Wasserstein plan,’ ‘Schrodinger bridge’]

Interpret This Week

  • Optimal trajectories and optimal intelligence?
  • Transport plan as an act; which geometry is Sinkhorn?
  • Why \(H(A\mid S)=0\), and why that model is weak?
  • Why low \(H(Z)\) is not the same as good?
  • Purely entropic Schottky; entropy today (last revision)

After This Lecture

  • The Week 1 question should now have a precise answer

Further Reading

  • Chapters 1–2; §4.2 (Sinkhorn, optional) of Peyré and Cuturi (2019)

  • Chapter 14 and Sections 22.2–22.3 of Welling et al. (2026)

  • the whole paper of Crooks (2007)

  • §§3–4 (optional) of Cuturi (2013)

  • the Good Regulator Theorem of Conant and Ashby (1970)

  • requisite variety; Shannon form of Ashby (1956)

  • Chapter 1 of Lawrence (2024)

Thanks!

References

Ashby, W.R., 1956. An introduction to cybernetics. Chapman & Hall, London.
Conant, R.C., Ashby, W.R., 1970. Every good regulator of a system must be a model of that system. International Journal of Systems Science 1, 89–97.
Crooks, G.E., 2007. Measuring thermodynamic length. Phys. Rev. Lett. 99, 100602. https://doi.org/10.1103/PhysRevLett.99.100602
Cuturi, M., 2013. Sinkhorn distances: Lightspeed computation of optimal transportation distances, in: Advances in Neural Information Processing Systems. pp. 2292–2300.
Lawrence, N.D., 2024. The atomic human: Understanding ourselves in the age of AI. Allen Lane.
Peyré, G., Cuturi, M., 2019. Computational optimal transport: With applications to data science, Foundations and trends in machine learning. Now Publishers. https://doi.org/10.1561/2200000073
Welling, M., Lu, S., Holdijk, L., 2026. Generative AI and stochastic thermodynamics: A tale of free energies. Cambridge University Press, Cambridge, U.K.