Shannon Entropy and the Partition Function

Neil D. Lawrence

FW26, William Gates Building

This Session

  • Shannon \(H\); Boltzmann \(S = kH\)
  • Arithmetic coding → Dasher: \(H\) as bits, \(p\) as the next letter
  • Partition function as a generating function
  • Scaffolding: \(I(X;Y)\), capacity, DPI (statement)

Lecture 1–2 counted human communication in Shannon’s bits — a bottleneck on how fast thought can leave the body.

  • Today: derive \(H\) and connect \(S = kH\)
  • The bottleneck remains; we gain the measure behind the bit

Information, entropy and intelligence course notebook setup

Shannon Entropy

Information Theory and AI

  • Claude Shannon developed information theory at Bell Labs
  • Information measured in bits, separated from context
  • Makes information fungible and comparable

Information Transfer Rates

  • Humans speaking: ~2,000 bits per minute
  • Machines communicating: ~600 billion bits per minute
  • Machines share information 300 million times faster than humans

Brownian Motion and Wiener

Betrand Russell
Albert Einstein
Norbert Wiener

Brownian Motion

Stochasticity and Control

Shannon asked: what number measures uncertainty in a discrete distribution \(p=(p_1,\ldots,p_n)\)?

  • Continuity; maximum at uniform; additive for independent parts
  • Result: \(H(p) = -\sum_i p_i \log p_i\)
  • No-go: codes cannot beat \(H\) on average; prescription: \(p\) is the code or belief

Arithmetic Coding and Dasher

Arithmetic Coding

  • Compression entails probabilistic modelling (MacKay, 2003, p. Ch.~6)
  • Predict next symbol → encode cheaply when the prediction is sharp
  • Modelling separated from the bit-string construction

The guessing game

  • Human predicts next character (MacKay Sec.~6.1)
  • Record guess-rank → skewed alphabet → easy to compress
  • Decoder: identical twin, stopped after the same number of guesses

From guesses to intervals

  • Model supplies \(P(x_n \mid x_{<n})\) — any context-dependent distribution
  • Subdivide \([0,1)\) so each symbol interval has length equal to its probability
  • Code = any bit-string whose dyadic interval ⊆ message interval
  • Length of final interval \(= P(\text{message}\mid\text{model})\)

Train a Predictive Model

Algorithm 6.3 in Code

Export for Dasher

  • Retrain on another corpus → rewrite JSON → reload Dasher
  • Box heights follow your \(P(c \mid \mathrm{context})\)

Dasher: Arithmetic Coding as Interface

  • Letters sized by \(P(\texttt{char} | \texttt{context})\). A large target = common = cheap
  • Move pointer right of centre: boxes enlarge and stream left
  • Move pointer left of centre to zoom out / unwrite
  • Character written when its box crosses the centre crosshair
  • Ease of selection \(\propto\) probability \(\propto\) 1 / information cost

DASHER screen height ∝ probability · boxes stream left across the crosshair

TYPED:

bits: 0.0 avg: b/ch H(next):

Click or Space to Go/Pause · Copy extracts text · pointer right zooms · left zooms out · Backspace unwrites · Escape resets

The Partition Function

The canonical ensemble from the bath: fix \(\beta\) and let the small system fluctuate.

  • \(Z(\beta) = \sum_i e^{-\beta E_i}\)
  • \(U = -\partial_\beta \log Z\), \(F = -\beta^{-1}\log Z\), \(S = \beta(U-F)\)
  • Equilibrium = on the \(\beta\)-manifold; finite-time driving leaves it
  • \(U = -\partial_\beta \log Z\)
  • \(F = -\beta^{-1}\log Z\)
  • Bath justifies the canonical ensemble

Week 3 derived Shannon entropy for discrete \(p=(p_1,\ldots,p_n)\).

  • \(0 \le H(p) \le \log n\) on \(n\) outcomes
  • Maximum at uniform; zero on a delta
  • Code interpretation: average length cannot beat \(H\)

Comparing two distributions needs a functional that is always sensible.

  • \(\mathrm{KL}(p\|q) = \sum_i p_i \log(p_i/q_i)\) discrete; \(\int p\log(p/q)\,dx\) continuous
  • Always \(\mathrm{KL}(p\|q) \ge 0\); zero iff \(p = q\) (same support)
  • Extra surprise when \(q\) stands in for \(p\) — not a metric, but a directed cost

Gaussian channel uses \(-\int p\log p\) — differential entropy.

  • Can be negative; not bounded below
  • Not a code-length bound — compare distributions with KL

Scaffolding, Not Outcomes

Four results we will need later — stated, not proved today.

  • Chain rule: \(H(X,Y) = H(X) + H(Y|X)\)
  • Mutual information: \(I(X;Y) = H(X)-H(X|Y) = H(X)+H(Y)-H(X,Y)\)
  • Capacity: no rate above \(C\); achieving \(C\) requires the capacity-achieving \(p(x)\)
  • Data processing: if \(X\to Y\to Z\) is Markov, \(I(X;Z)\le I(X;Y)\)
  • No-go: \(R \le C\); processing cannot create information
  • Prescription: the \(p(x)\) that achieves \(C\)
  • Week 8: prove DPI; information bottleneck as the prescription on \(I\)

Three Framings, First Pass

  • Same \(H\); three operational assumptions
  • Week 3: one bit \(\leftrightarrow\) \(k_B T \ln 2\) joules (Szilard, Landauer)
  • Synthesis is week 4

Define This Week

  • Why is entropy a sensible measure of information?
  • Equilibrium versus non-equilibrium? (first cut)
  • Channel capacity? (statement)
  • Mutual information? (definition)

After This Lecture

  • Next: Maxwell and Landauer (3 November)
  • LLM: is Shannon entropy a bound or a recipe?

Further Reading

  • Sections 1–6 of Shannon (1948)

  • Chapters 1–4; Chapter 6 of MacKay (2003)

  • Chapter 2 of Cover and Thomas (1991)

  • Sections 1.2.3–1.2.4 and 3.2.4 of Welling et al. (2026)

  • Chapter 16 of Callen (1985)

  • Chapter 7 of Cover and Thomas (1991)

  • Chapters 8–10 of MacKay (2003)

Thanks!

References

Callen, H.B., 1985. Thermodynamics and an introduction to thermostatistics, 2nd ed. Wiley, New York.
Coales, J.F., Kane, S.J., 2014. The “yellow peril” and after. IEEE Control Systems Magazine 34, 65–69. https://doi.org/10.1109/MCS.2013.2287387
Cover, T.M., Thomas, J.A., 1991. Elements of information theory. Wiley, New York.
Einstein, A., 1905. Über die von der molekularkinetischen Theorie der Wärme geforderte Bewegung von in ruhenden Flüssigkeiten suspendierten Teilchen. Annalen der Physik 322, 549–560. https://doi.org/10.1002/andp.19053220806
MacKay, D.J.C., 2003. Information theory, inference and learning algorithms. Cambridge University Press, Cambridge, U.K.
Shannon, C.E., 1948. A mathematical theory of communication. The Bell System Technical Journal 27, 379–423. https://doi.org/10.1002/j.1538-7305
Welling, M., Lu, S., Holdijk, L., 2026. Generative AI and stochastic thermodynamics: A tale of free energies. Cambridge University Press, Cambridge, U.K.