Questions
Published on 13 October. Meet the whole list in lecture 1. Do not expect to answer most of it then.
Each question has two stages.
- Define: a competent textbook answer, in the week the object is introduced.
- Interpret: this module’s own answer. Later — often week 5, or week 8.
Class tests examine the applied form of the define-stage questions, on a new example. They do not ask “What is X?”
Lecture 1 — 13 October — Motivation and Review
Define this week
- How was entropy discovered?
- Who were Carnot and Clausius, and what did each contribute?
- What is the relationship between energy and entropy?
- What is a thermodynamic bath?
- What is Schottky’s anomaly? (first cut)
- What are the information constraints on a human? (bandwidth, first cut)
The theme (return every week)
- Does entropy tell us what we cannot do, or what we should do?
- What, this week, is the no-go, and what is the prescription?
Named, not yet answered
- How is entropy understood today?
- What is the difference between equilibrium and non-equilibrium thermodynamics?
- What is the purely entropic interpretation of Schottky’s anomaly?
- What is the Legendre transform? (you have already written $F = U - TS$)
Lecture 2 — 20 October — Boltzmann and Free Energy
Quiz 1 occupies the first ten minutes.
Define this week
- What is the Gibbs–Boltzmann distribution (canonical ensemble)?
- How do internal energy, entropy, and Helmholtz free energy enter $p_i$?
- What is Schottky’s anomaly?
- What is a thermodynamic bath? (revisited)
Named, not yet answered
- How is entropy understood today?
- Shannon entropy versus differential entropy? (full contrast: week 6)
Lecture 3 — 27 October — Shannon Entropy
Define this week
- Why is entropy a sensible measure of information?
- What is the chain rule of entropy?
- What is mutual information? (definition from the chain rule)
- What is Shannon’s channel capacity theorem? (statement)
- What is the difference between equilibrium and non-equilibrium thermodynamics? (first cut)
Named, not yet answered
- What is the data processing inequality? (statement today; proof week 8)
- What is the information bottleneck? (week 8)
- How is entropy understood today? (Shannon alone is not the intended answer)
- Shannon entropy versus differential entropy? (full contrast: week 6)
Lecture 4 — 3 November — Maxwell and Landauer
Define this week
- What is Maxwell’s demon?
- How does one Shannon bit relate to thermodynamic work? (Szilard; Feynman piston chain)
- Why cannot Maxwell-demon feedback improve a car engine the way ATP synthase uses information?
- How does Landauer’s principle relate to Clausius’s second law?
- What is a Boltzmann memory (equal-depth double well), and how does it relate to an information reservoir?
- What is the thermodynamics of information? (Parrondo et al.; first cut)
- What is the relationship between information and intelligence? (first cut: Landauer)
- What are the information constraints on a human? (bandwidth and erasure)
Interpret (Worksheet 2; revisit week 8)
- Dissipation ledger (Ellis) versus information ledger (Szilard/Landauer/Bennett): same no-go, different bookkeeping — which argument matches the generality of the second law?
- Does Landauer derive the erasure cost, or assume the second law to save it? (Contrast bad textbook circularity with Bennett (1982) on logical irreversibility of erasure.)
Lecture 5 — 10 November — MaxEnt and Synthesis
Quiz 2 occupies the first ten minutes.
Define this week
- What is the maximum entropy principle?
- What is the exponential family?
- What is the Legendre transform? (Helmholtz was the first example; $H = A - \theta\cdot\eta$ is the second)
- Why Helmholtz $F$ rather than Gibbs $G$? (which variables the bath fixes: we use fixed $T$, fixed state space; chemistry often uses fixed $T,P$)
- What is the relationship between information and entropy?
- How is entropy understood today? (intended answer: LO7)
Named, not yet answered
- What if MaxEnt is over a coupling, with two prescribed marginals rather than a list of moments? (Sinkhorn; week 8)
Lecture 6 — 17 November — Fisher Metric
Define this week
- What is thermodynamic length? (Crooks: Fisher–Rao length of a path of equilibrium states)
- What is KL divergence?
- What is the difference between Shannon entropy and differential entropy?
- How does the Legendre transform produce the dual coordinates of information geometry?
- What does it mean for an exponential family to be e-flat and m-flat?
- What is dual flatness?
- What is the Pythagorean theorem for KL divergence on dual flats?
Named, not yet answered
- How do optimal trajectories in thermodynamic length relate to optimal intelligence?
Lecture 7 — 24 November — Projection and Geodesics
Quiz 3 occupies the first ten minutes.
Define this week
- Thermodynamic length as a geodesic problem (minimum-dissipation protocol)
- The exponential family, now geometrically (project in $\eta$, travel in $\theta$)
- What is maximum entropy as an $m$-projection?
- What is natural gradient descent, and why is it the geometrically correct gradient?
Named, not yet answered
- What is an alternating $m$-projection onto two constraint sets? (Sinkhorn; week 8)
Lecture 8 — 1 December — Multi-Information and Limits on Intelligence
Quiz 4 occupies the first ten minutes.
Define this week
- What is the multi-information and what is mutual information?
- What is the data processing inequality?
- What is the information bottleneck?
- What is von Neumann entropy?
- What is the matrix exponential family?
- What is a coupling of two distributions?
- What is entropy-regularized optimal transport, and why is its solution an exponential family?
- What does Sinkhorn iterate, and which of the three geometries does it compute?
Interpret this week
- How do optimal trajectories in thermodynamic length relate to optimal intelligence?
- What is the relationship between information and intelligence? (final)
- How is a transport plan an abstraction of an action, the way $p$ is an abstraction of a code?
- In entropic OT, what is the no-go and what is the prescription?
- Is “optimal intelligence” cheapest rearrangement, least-committal rearrangement, or minimum-dissipation protocol? Name the geometry.
- What is the data processing inequality? (interpret: a superintelligence that “just processes more”)
- What is the information bottleneck? (interpret: intelligence as relevant compression)
- How is entropy treated in early cybernetics?
- What is the law of requisite variety?
- What is the Good Regulator Theorem?
- Why does concavity of Shannon entropy yield a deterministic $H(Z)$-optimal policy?
- Why is $H(A\mid S)=0$ a deliberately weak definition of ``model’’?
- Why does minimising $H(Z)$ not by itself mean the outcome is desirable?
- What is the purely entropic interpretation of the Good Regulator Theorem?
- What is a viable system?
- What are the information constraints on a human?
- What is the purely entropic interpretation of Schottky’s anomaly?
- How is entropy understood today? (last revision)
The cybernetics cluster (requisite variety, Good Regulator, viable system) is used in lecture 8 as named evaluation tools for LO13. It is not a separate outcome and not the course punchline (that remains: entropy forbids; probability prescribes). Requisite variety appears in Shannon form as $H(E)\ge H(D)-I(D;R)$. The Good Regulator Theorem proper is the existence of an $H(Z)$-optimal policy with $H(A\mid S)=0$, earned from concavity of entropy (and, for stochastic channels, extreme points of the policy polytope). Caveats to teach: a constant action also satisfies $H(A\mid S)=0$; low $H(Z)$ is predictability, not desirability. The IB-shaped reading $\min I(S;A)$ subject to $H(Z)\le\epsilon$ is a course gloss after the theorem, not Conant and Ashby’s 1970 statement. The data-processing inequality and the information bottleneck are taught under LO10 in the first half of lecture 8 and used in the second half. A light inaccessible-game introduction may appear as colour, not as a separate outcome. Sinkhorn is named under LO12 as the discrete algorithm for the Schrödinger / entropic coupling, not as a fourth geometry.