FW26, William Gates Building
Fisher’s information \(\neq\) Shannon’s entropy * Shannon \(H\): uncertainty in an outcome * Fisher \(I(\theta)\): sensitivity of \(\log p\) to \(\theta\) * Same word; different operational reading
\[ I(\theta)=\mathbb{E}\!\left[(\partial_\theta\log p)^2\right] \] * Large \(I\): data distinguish nearby \(\theta\) well * Small \(I\): parameter almost invisible in samples
Week 2, revisited (name only): * Two-state: \(G(\beta)\propto C\) * Schottky peak \(=\) Fisher peak * Entropic reading: week 8
Same \(G\), three origins: \[ G(\boldsymbol{\theta}) = \nabla^2 \mathcal{A}(\boldsymbol{\theta}) = \mathrm{Cov}_{\boldsymbol{\theta}}[T(\mathbf{x})] \] * Fisher / CR / Hessian — now as Riemannian metric
Statistical Manifold: * Each point \(\boldsymbol{\theta}\) = a probability distribution * Space of all distributions = curved manifold * Fisher information = metric (ruler) on this space * Measures “closeness” between distributions
\[ \text{d}s^2 = \text{d}\boldsymbol{\theta}^\top G(\boldsymbol{\theta}) \text{d}\boldsymbol{\theta} \] * Measures information distance between distributions * Larger \(G\) = distributions more distinguishable * Smaller \(G\) = distributions harder to tell apart
Cramér–Rao (restated geometrically): \[ \text{cov}(\hat{\boldsymbol{\theta}}) \succeq G^{-1}(\boldsymbol{\theta}) \] * \(G^{-1}\) = error ellipsoid * High \(G\) → tight estimation; low \(G\) → loose
Two Roles of Fisher Information: 1. Metric → defines distances between distributions 2. In gradient → \(\nabla H = -G(\boldsymbol{\theta})\boldsymbol{\theta}\)
\[ \dot{\boldsymbol{\theta}} = \nabla H = -G(\boldsymbol{\theta})\boldsymbol{\theta} \]
Gaussian: \(G(\boldsymbol{\theta}) = \Sigma\) * Information metric = covariance * \(G^{-1} = \Sigma^{-1}\) = precision
* Information ellipsoid = probability ellipsoid * Special to Gaussians in natural parameters
Categorical: \[ G_{ij} = \delta_{ij}\pi_i - \pi_i\pi_j \] * Defines probability simplex geometry * Center of simplex: balanced information * Corners: concentrated information * Metric captures curvature
Information Geometry: * Fisher metric → Riemannian geometry * Exponential families → dually flat structure * Geodesics → shortest paths between distributions * Zero curvature → special “flat” structure * Key for constrained dynamics later
A statistical manifold: each point is a distribution \(p(x\mid\theta)\).
Before length: the exact nonequilibrium relation that the second law and the \(\mathcal{L}^2/\tau\) bound sit inside.
\[ \frac{P(W)}{P_R(-W)} = e^{\beta(W-\Delta F)} \]
For a slow protocol \(\lambda(t)\) on the equilibrium manifold, define length with the Fisher metric.
\[ \mathcal{L} = \int_0^\tau \sqrt{\dot\lambda^\top \mathcal{I}(\lambda)\,\dot\lambda}\,dt \]
Chapters 1–2 of Amari (2016)
fluctuation theorem (statement) of Crooks (1999)
equality; second-law corollary of Jarzynski (1997)
Chapter 17 (fluctuation theorem; stops before length) of Welling et al. (2026)
the whole paper of Crooks (2007)