Long Proofs
- The Quadrupole Formula for Gravitational Radiation
- Oppenheimer–Snyder Collapse: the Interior Matching
- The Cantor–Schröder–Bernstein Theorem
- The Completeness Theorem
- Arithmetization and the Incompleteness Theorems
- The Universal Turing Machine
- Completeness of the Real Numbers
- Equivalence of the Einstein–Hilbert Action in Metric and Vielbein Variables
- Wigner's Theorem on Quantum Symmetries
- The Classification of Kinematical Algebras
- Invariant Bilinear Forms on an Expanded Algebra
- Darboux's Theorem: Local Canonical Coordinates
- The Korn–Lichtenstein Theorem and the Elliptic Canonical Form
- The Cauchy–Kovalevskaya Theorem
- The Stäckel Conditions and the Eleven Separable Systems
- The Mean-Value Property Characterizes Harmonic Functions
- Kruzhkov's Doubling of Variables and the Entropy Solution
- Lévy's Continuity Theorem
- The Lindeberg–Feller Central Limit Theorem
- The Berry–Esseen Inequality
- Wilks' Theorem in the Multi-Parameter Case
- Rice's Formula for the Expected Number of Upcrossings
- The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators
- The Spectral Theorem for a Bounded Self-Adjoint Operator
- Stone's Theorem on One-Parameter Unitary Groups
- Von Neumann's Criterion for Self-Adjoint Extensions
- The Nuclear Spectral Theorem of Gelfand and Maurin
- The Implicit Function Theorem
- The General Stokes Theorem for Differential Forms
- The Poincaré–Birkhoff–Witt Theorem
- Cartan's Criterion for Semisimplicity
- Racah's Theorem on the Number of Casimir Operators
- Covering Spaces and the Fundamental Group of the Rotation Group
- The Structure Constants of $\mathfrak{su}(3)$
- Simplicity of $\mathfrak{su}(3)$
- From the Little Algebra to the Labels of a Massive Representation
- Whitehead's Lemmas and the Rigidity of Semisimple Algebras
- The Second Cohomology of the Galilei Algebra
- Central Extensions of Three Infinite-Dimensional Algebras
- The Strengthened Jacobi Condition and Conjugate Points
- Tonelli's Existence Theorem for an Integrand Depending on the Function
- Completeness of the Sturm–Liouville Eigenfunctions in the Energy Norm
- The Plateau Problem: the Architecture of Douglas' Proof
- Change of Variables in a Multiple Integral
- The Helmholtz Decomposition
- The Quotient Manifold Theorem
- The Liouville–Arnold Theorem
- Marsden–Weinstein Reduction
- The Stone–von Neumann Theorem
- Hudson's Theorem: the Pure States of Non-Negative Wigner Function
- The Jacobi Identity for the Dirac Bracket
- The Gauss–Codazzi Form of the Einstein–Hilbert Lagrangian
- The Hypersurface-Deformation Algebra
- Existence and Uniqueness of a Nilpotent BRST Charge
- Bertrand's Theorem
- Isotropic and Cubic Elastic Tensors
- Reduction of Three-Dimensional Elasticity to the Kirchhoff Plate Equation
- Hertz's Solution for the Elastic Half-Space Under an Axisymmetric Pressure
- D'Alembert's Paradox for a Body of Arbitrary Shape
- The Coefficient $6\pi$ in Stokes' Drag Law
- The Kármán Spacing Ratio of the Vortex Street
- The Joukowski Map, the Kutta Condition and the $2\pi$ Lift Slope
- The Kármán–Howarth Relation and the Four-Fifths Law
- The Korteweg–de Vries Equation from the Free-Surface Problem
- The Lorenz System from Boussinesq Convection
- Rayleigh Acoustic Streaming Above a Vibrating Surface
- The Free-Streamline Jet: Kirchhoff's Contraction Coefficient
- Prandtl's Boundary-Layer Equations and the Blasius Similarity Solution
This appendix collects the derivations that are too long to interrupt the running text. Each section opens by naming the statement it proves and closes by returning the reader to the chapter where that statement is used. The derivations are carried out in full: every algebraic step that a careful reader would want to check is displayed, and all quantities carry their SI dimensions throughout — in particular the constants \(G\), \(c\) and \(\hbar\) are kept explicit everywhere, in obedience to the SI axiom of the front matter.
A word on what in full can and cannot mean. Some of the results collected here descend, at exactly one step, from a theorem this treatise does not build — most often a convergence theorem of Lebesgue integration, of which Real Analysis constructs only the Riemann predecessor, and in four cases a substantial theorem of another subject altogether. Where that happens the section does three things rather than pass over it: it states the imported theorem precisely, as a displayed environment of its own with its citation; it derives everything else from that statement; and it carries a remark headed What is quoted here naming the import and saying why it is not proved. A derivation that hides such a step would be worth less than the pending-derivation box it replaces. Fourteen of the thirty-two sections written for Part II quote nothing at all.
The Quadrupole Formula for Gravitational Radiation
This appendix proves Phenomenon 46.1 (Chapter 46): the field equations of general relativity, linearized about flat spacetime, admit transverse wave solutions propagating at speed \(c\) with two polarizations; these waves carry energy; and a slowly moving, self-gravitating source radiates them with a power fixed by the third time derivative of its mass quadrupole moment. The chain of results proved below — wave equation, transverse–traceless gauge, geodesic-deviation observable, Isaacson energy flux, quadrupole generation formula, and the binary-inspiral application — is exactly the chain that connects Einstein's field equations to the two great experimental confirmations: the orbital decay of the binary pulsar PSR B1913+16 (Section 47.1) and the direct detection of GW150914 by the LIGO interferometers (Section 47.3).
The history deserves one sentence. Einstein obtained the linearized wave solutions in 1916 [Einstein:1916], and in 1918 published the quadrupole radiation formula in its correct form, repairing a computational slip of the earlier paper [Einstein:1918]; the modern treatments we follow — and against which every step below can be checked — are those of Misner, Thorne and Wheeler [Misner:1973], Wald [Wald:1984], and Maggiore [Maggiore:2008].
Throughout this section the metric signature is \((-,+,+,+)\), coordinates are \(x^\mu = (x^0, x^i) = (ct, \vect{x})\), Greek indices run over \(0,1,2,3\), Latin indices over the spatial values \(1,2,3\), and the flat d'Alembertian is
Setup: linearized geometry
We consider spacetimes that deviate weakly from Minkowski space, so that global coordinates exist in which
with \(\eta_{\mu\nu} = \diag(-1,+1,+1,+1)\). All indices are raised and lowered with \(\eta\), an operation that is consistent to the linear order in \(h\) at which we work. The inverse metric is
as one checks by direct multiplication: \((\eta_{\mu\rho} + h_{\mu\rho})(\eta^{\rho\nu} - h^{\rho\nu}) = \delta_\mu^{\ \nu} + h_\mu^{\ \nu} - h_\mu^{\ \nu} + O(h^2) = \delta_\mu^{\ \nu} + O(h^2)\).
Christoffel symbols.
Inserting Equations (A.2) and (A.3) into the definition of the Levi-Civita connection and discarding \(O(h^2)\),
The connection is thus itself of first order in \(h\).
Riemann tensor.
In the curvature tensor \(R^{\rho}{}_{\sigma\mu\nu} = \pp_\mu\Gamma^{\rho}{}_{\nu\sigma} - \pp_\nu\Gamma^{\rho}{}_{\mu\sigma} + \Gamma^{\rho}{}_{\mu\lambda}\Gamma^{\lambda}{}_{\nu\sigma} - \Gamma^{\rho}{}_{\nu\lambda}\Gamma^{\lambda}{}_{\mu\sigma}\) the \(\Gamma\Gamma\) terms are \(O(h^2)\) and drop. Substituting Equation (A.4),
where the two \(\pp_\mu\pp_\nu h_{\lambda\sigma}\) terms cancelled. Lowering the first index gives the fully covariant linearized Riemann tensor,
Ricci tensor, Ricci scalar, Einstein tensor.
Contracting the first and third indices of Equation (A.5) and writing \(h \equiv \eta^{\mu\nu}h_{\mu\nu}\) for the trace,
and contracting once more,
The linearized Einstein tensor is therefore
Trace reversal.
The expression Equation (A.9) simplifies decisively in terms of the trace-reversed perturbation
(The map is an involution: trace-reversing twice returns \(h_{\mu\nu}\).) Substituting the third relation of Equation (A.10) into Equation (A.7) term by term,
so that
and the Einstein tensor assembles into
the two \(\frac14\eta_{\mu\nu}\Box\bar h\) contributions having cancelled between \(R_{\mu\nu}\) and \(-\frac12\eta_{\mu\nu}R\). Every term now contains \(\bar h_{\mu\nu}\) either under a d'Alembertian or under a divergence \(\pp^\rho\bar h_{\rho\nu}\) — which is the invitation to fix a gauge.
Gauge freedom and the Lorenz gauge
The split Equation (A.2) is not unique: an infinitesimal coordinate change
preserves the form \(g = \eta + h\) with a different, equally small \(h\). Indeed, from the tensor transformation law and \(\pp x^\rho/\pp x'^\mu = \delta^\rho_{\ \mu} - \pp_\mu\xi^\rho + O(\xi^2)\),
so the perturbation transforms as
This is the gauge freedom of linearized gravity, the precise analogue of \(A_\mu \to A_\mu - \pp_\mu\chi\) in electrodynamics [Jackson:1999]. The linearized Riemann tensor Equation (A.6) is gauge invariant: substituting Equation (A.16),
because partial derivatives commute and the eight triple-derivative terms cancel in pairs. Curvature — hence everything observable — is unaffected by the choice of \(\xi^\mu\); we may therefore spend that choice on simplifying the field equations.
Taking the trace of Equation (A.16) gives \(h' = h - 2\pp_\rho\xi^\rho\), so the trace-reversed perturbation transforms as
and its divergence as
Given any \(\bar h_{\mu\nu}\), choose \(\xi_\nu\) to solve the inhomogeneous wave equation \(\Box\xi_\nu = \pp^\mu\bar h_{\mu\nu}\) — a solution always exists, for instance by the retarded Green function of \(\Box\) constructed in Generation: the quadrupole formula below. In the new coordinates the Lorenz gauge (also called harmonic or de Donder gauge) holds:
In this gauge the last three terms of Equation (A.13) vanish, and the full Einstein equations \(G_{\mu\nu} = (8\pi G/c^4)\,T_{\mu\nu}\), linearized, collapse to a flat-space wave equation with source:
Taking \(\pp^\mu\) of both sides and using Equation (A.20) shows \(\pp^\mu T_{\mu\nu} = 0\): at linear order the source moves on the flat background, its energy–momentum conserved in the special-relativistic sense. This is the self-consistency of the linear approximation [Misner:1973] [Wald:1984].
Vacuum solutions: plane waves and the transverse–traceless gauge
In vacuum (\(T_{\mu\nu} = 0\)) equation Equation (A.21) reads \(\Box\bar h_{\mu\nu} = 0\): each component obeys the massless wave equation, so disturbances of the metric propagate at the speed \(c\) — the first assertion of Phenomenon 46.1. Insert the plane-wave ansatz
(real part understood). Then \(\Box\bar h_{\mu\nu} = -k_\alpha k^\alpha\, \bar h_{\mu\nu} = 0\) forces the wave vector to be null,
and the Lorenz condition Equation (A.20) becomes the transversality constraint
Counting the physical degrees of freedom.
A symmetric \(A_{\mu\nu}\) has \(10\) independent components; Equation (A.24) removes \(4\). But the gauge is not yet exhausted: any further \(\xi^\mu\) with \(\Box\xi^\mu = 0\) preserves Equation (A.20), by Equation (A.19). Take \(\xi^\mu = B^\mu\ee^{\ii k_\alpha x^\alpha}\) with the same null \(k\); then Equation (A.18) shifts the amplitude by
The four free constants \(B_\mu\) remove four more components. Explicitly, let the wave run along \(z\), \(k^\mu = (k,0,0,k)\), \(k_\mu = (-k,0,0,k)\), and write \(\beta \equiv k_\rho B^\rho = k(B_0 + B_3)\). From Equation (A.25),
and the four quantities \((B_0 + B_3,\,B_1,\,B_2,\,B_3 - B_0)\) are independent linear combinations of the \(B_\mu\): the linear system is invertible. We may therefore impose
The remaining component \(A_{00}\) then vanishes automatically: transversality Equation (A.24) with \(\nu = 0\) reads \(k^0 A_{00} + k^i A_{i0} = 0\), and \(A_{i0} = 0\) with \(k^0 = \omega/c \neq 0\) forces \(A_{00} = 0\). The count of physical polarizations is thus
Because the trace vanishes, \(\bar h_{\mu\nu} = h_{\mu\nu}\) in this gauge: bar and no-bar coincide. The conditions
define the transverse–traceless (TT) gauge. For the wave along \(z\), transversality further gives \(k(A_{0\nu} + A_{3\nu}) = 0\), hence \(A_{3\nu} = -A_{0\nu} = 0\): only \(A_{11} = -A_{22}\) and \(A_{12} = A_{21}\) survive. Writing \(A_{11} \equiv h_+\) and \(A_{12} \equiv h_\times\),
the two physical polarizations of Phenomenon 46.1 [Misner:1973] [Maggiore:2008]. They are spin-2 objects: a rotation by \(\psi\) about the propagation axis mixes them as \(h_+ \to h_+\cos 2\psi + h_\times\sin 2\psi\), so the pattern returns to itself after half a turn.
Physical effect: geodesic deviation and the strain
A single free test mass tells us nothing: in TT coordinates a particle initially at rest stays at fixed coordinates. Its geodesic equation with initial four-velocity \(u^\mu = (c,0,0,0)\) gives
since \(h_{0\mu} = 0\) in the TT gauge, Equation (A.29). The coordinates simply ride along with the wave. What oscillates is the proper distance between two such masses. Let both sit on the \(x\) axis with coordinate separation \(L_{\mathrm{c}}\), the wave Equation (A.30) passing along \(z\). The proper distance at time \(t\) is
so the fractional length change — the strain — is
for a wave of amplitude \(h\) optimally oriented with respect to the separation.
The same result follows covariantly from geodesic deviation, which also exhibits the wave as a tidal force. For two neighbouring geodesics with separation vector \(\xi^\mu\) and common four-velocity \(u^\mu\) [Misner:1973] [Wald:1984],
For slowly moving masses \(u^\mu \simeq (c,0,0,0)\), \(\tau \simeq t\), and the only components of Riemann we need follow from Equation (A.6) with the TT conditions \(h_{0\mu} = 0\):
Then Equation (A.34) gives, in the local proper frame of the first mass,
which integrates (for \(\xi\) nearly constant) to \(\delta\xi^i = \frac12 h^{\mathrm{TT}}_{ij}\xi^j\), reproducing Equation (A.33). Note that the observable is curvature, Equation (A.35), which we proved gauge invariant: the strain is physics, not coordinates.
Applying \(\delta\xi^i = \frac12 h^{\mathrm{TT}}_{ij}\xi^j\) to a ring of free masses of radius \(L\) in the \(xy\) plane, \(\xi^j = L(\cos\varphi, \sin\varphi, 0)\): a pure \(h_+\) wave displaces \((\delta x, \delta y) = \frac12 h_+ L\,(\cos\varphi, -\sin\varphi)\), squeezing the ring into an ellipse along \(x\) while stretching it along \(y\) and vice versa each half period; a pure \(h_\times\) wave produces the same ellipse rotated by \(45^\circ\). This quadrupolar breathing pattern is precisely what a Michelson interferometer with perpendicular arms measures: the wave lengthens one arm while it shortens the other, and the differential strain \(\Delta L_x - \Delta L_y = h_+ L\) appears as an optical phase shift at the dark port. With \(L = 4\,\mathrm{km}\) arms and strain sensitivity near \(10^{-22}\) per root hertz in the band around \(100\,\mathrm{Hz}\), the Advanced LIGO instruments [Aasi:2015] resolve arm-length changes of order \(10^{-18}\,\mathrm{m}\) — the observable exploited in Section 47.3 [Abbott:2016].
Energy carried by the wave: the Isaacson tensor
That gravitational waves carry energy was disputed into the 1950s; the resolution is that the energy is not localizable at a point (the equivalence principle forbids it) but is perfectly well defined once averaged over a few wavelengths. We follow the short-wave (Isaacson) analysis [Misner:1973] [Maggiore:2008].
Expand the Ricci tensor of \(g = \eta + h\) to second order in \(h\),
where \(R^{(1)}\) is the linear expression Equation (A.7). In vacuum, at second order, the equation \(R_{\mu\nu} = 0\) says that the small residual background curvature generated at \(O(h^2)\) is sourced by \(-R^{(2)}_{\mu\nu}\): the wave gravitates. Averaging over a spacetime region several wavelengths across (denoted \(\avg{\cdot}\)) and defining
the coarse-grained field equations take the form \(G^{(1)}_{\mu\nu}[\text{background}] = (8\pi G/c^4)\,t_{\mu\nu}\): the averaged quadratic terms act as an effective stress–energy tensor of the radiation.
The straightforward if lengthy expansion of the Ricci tensor to second order gives [Maggiore:2008]
Work in the TT gauge, where \(h = 0\) and \(\pp^\beta h_{\alpha\beta} = 0\): the last six terms of Equation (A.39) vanish identically. Under the average, total derivatives are suppressed by (wavelength)/(averaging scale) and may be dropped, which licenses integration by parts inside \(\avg{\cdot}\); and the field equations \(\Box h_{\alpha\beta} = 0\) hold. Then, term by term,
Only the first two terms of Equation (A.39) survive, combining to
Its trace, \(\eta^{\mu\nu}\avg{R^{(2)}_{\mu\nu}} = \frac14\avg{h_{\alpha\beta}\Box h^{\alpha\beta}} = 0\) after one more integration by parts, vanishes on shell, so Equation (A.38) yields the Isaacson stress–energy tensor. In the TT gauge only spatial components of \(h\) are nonzero, and
(spatial indices summed; their up/down position is immaterial in the flat background). Although derived in a particular gauge, the averaged tensor is gauge invariant to the order considered [Misner:1973] [Maggiore:2008].
For the plane wave Equation (A.30) travelling along \(z\), every component is a function of \(t - z/c\), so \(\pp_z = -c^{-1}\pp_t\) acting on \(h\), and the energy density and energy flux are
using \(h^{\mathrm{TT}}_{ij}h^{\mathrm{TT}}_{ij} = 2(h_+^2 + h_\times^2)\) from Equation (A.30). The prefactor \(c^3/16\pi G\) is enormous, \(\approx 8.0\times 10^{33}\,\mathrm{W}/\mathrm{m}^{2}\) per unit \(\avg{\dot h^2}//\mathrm{s}^{2}\): even a strain of \(10^{-21}\) oscillating at \(100\,\mathrm{Hz}\) carries a flux of order \(10^{-3}\,\mathrm{W}/\mathrm{m}^{2}\) — comparable to moonlight, delivered by a distortion of geometry a thousandth of a proton radius across a kilometre.
Generation: the quadrupole formula
We now solve the sourced equation Equation (A.21).
Retarded Green function.
The retarded solution of \(\Box f = -S\), with \(\Box\) as in Equation (A.1), is
To verify, consider a single spherical wave \(f = g(t - r/c)/r\) centred on \(\vect{x}'\), \(r = \abs{\vect{x}-\vect{x}'}\). For \(r \neq 0\), the radial Laplacian gives \(\nabla^2 f = r^{-1}\pp_r^2(rf) = g''(t-r/c)/(c^2 r) = c^{-2}\pp_t^2 f\), so the homogeneous wave equation holds; near \(r = 0\), \(f \to g(t)/r\) and \(\nabla^2(1/r) = -4\pi\delta^3(\vect{x}-\vect{x}')\) supplies a point source of strength \(4\pi g(t)\). Superposing such waves with \(g = S\,\dd^3x'/4\pi\) yields Equation (A.48); causality selects the retarded rather than advanced root [Jackson:1999]. Applying this to Equation (A.21) with \(S = (16\pi G/c^4)T_{\mu\nu}\),
Far zone and slow motion.
Let the source occupy a region of size \(d\) about the origin and let \(r = \abs{\vect{x}} \gg d\), with unit line of sight \(\vect{n} = \vect{x}/r\). Then \(\abs{\vect{x}-\vect{x}'} = r - \vect{n}\cdot\vect{x}' + O(d^2/r)\), and keeping only the leading \(1/r\),
If moreover the internal motions are slow, \(v \ll c\) — equivalently the emitted wavelength \(\lambda \sim c\,d/v\) far exceeds the source size — the retardation spread \(\vect{n}\cdot\vect{x}'/c\) across the source may be dropped, and every component of \(T_{\mu\nu}\) is evaluated at the single retarded time \(t_r = t - r/c\).
From stress to quadrupole: the conservation identities.
The spatial integral of \(T^{ij}\) is not an independent datum: flat-space conservation \(\pp_\mu T^{\mu\nu} = 0\), guaranteed at this order by Equation (A.21), trades it for time derivatives of the energy density. With \(x^0 = ct\), the two components of the conservation law read
Multiply the first by \(x^i x^j\) and integrate over all space (the source has compact support, so boundary terms vanish under integration by parts):
Multiply the second of Equation (A.51) by \(x^j\), integrate, and symmetrize:
Chaining Equations (A.52) and (A.53),
where, with the mass density \(\rho \equiv T^{00}/c^2\) (the dominant part of \(T^{00}\) for slow motion), the mass quadrupole moment is
Substituting Equation (A.54) into Equation (A.50),
— the quadrupole formula for the field [Einstein:1918] [Misner:1973]. Dimensionally: \(G/c^4 \sim \mathrm{s}^{2}/\mathrm{kg}/\mathrm{m}\) and \(\ddot Q \sim \mathrm{kg}\,\mathrm{m}^{2}/\mathrm{s}^{2}\), so \(h\) is a pure number divided by \(r\) — a strain falling off as \(1/r\), as a radiation field must. The time components need not be computed separately: the Lorenz condition Equation (A.20) determines \(\bar h^{0\mu}\) from \(\bar h^{ij}\) up to stationary pieces. Physically, the monopole moment \(\int\rho\,\dd^3x\) is the conserved total mass and the dipole moment \(\int\rho\,x^i\,\dd^3x\) moves uniformly by conservation of momentum, so neither can oscillate: mass conservation kills monopole radiation, momentum conservation kills dipole radiation, and the quadrupole is the first moment free to radiate [Misner:1973]. This is the deep reason gravitational radiation is so feeble compared with the electric-dipole radiation of electromagnetism.
TT projection and the reduced quadrupole.
The physical waveform is the TT part of Equation (A.56). Define the projector transverse to the line of sight, \(P_{ij}(\vect{n}) = \delta_{ij} - n_i n_j\) (idempotent: \(P_{ik}P_{kj} = P_{ij}\)), and the TT projector
Because \(\Lambda\) annihilates any term proportional to \(\delta_{kl}\), we may freely replace \(Q_{ij}\) by its trace-free part. We adopt, here and in the luminosity below, the reduced quadrupole moment
(with \(r'^2 = x_k x_k\) the squared distance from the origin), in terms of which
Radiated power.
Insert Equation (A.59) into the flux Equation (A.47) and integrate over a large sphere. For an outgoing wave \(h \propto r^{-1}\times(\text{function of } t - r/c)\) one has \(\pp_r h = -c^{-1}\dot h + O(1/r^2)\), so the radial energy flux is \(c\,t^{0r} = (c^3/32\pi G)\avg{\dot h^{\mathrm{TT}}_{ij}\dot h^{\mathrm{TT}}_{ij}}\) and
where the projector identity of Equation (A.57) collapsed the two \(\Lambda\)'s into one. The angular integral needs only the elementary moments of the unit vector,
fixed by isotropy and by contracting both sides. For any constant symmetric trace-free \(A_{ij}\), expanding Equation (A.57) gives \(\Lambda_{ij,kl}A_{ij}A_{kl} = A_{ij}A_{ij} - 2A_{ij}A_{ik}n_j n_k + \frac12(n_i n_j A_{ij})^2\), and with Equation (A.61),
Substituting into Equation (A.60),
— Einstein's quadrupole luminosity [Einstein:1918] [Maggiore:2008], with the reduced moment Equation (A.58) and the third time derivative evaluated at retarded time. The prefactor \(G/c^5 \approx 2.75\times 10^{-53}\, \mathrm{W}^{-1}\) explains at a glance why laboratory sources are hopeless and only relativistic astrophysical masses radiate detectably; its inverse \(c^5/G \approx 3.6\times 10^{52}\,\mathrm{W}\) sets the natural luminosity scale of strong-field gravity.
Application: the circular compact binary
The cleanest astrophysical source is two point masses \(m_1\), \(m_2\) in a circular orbit — the configuration realized, to excellent approximation, by compact binaries in their late inspiral. Work in the centre-of-mass frame with total and reduced masses
separation \(a\), and Kepler angular frequency
With the centre of mass at the origin, \(\vect{x}_1 = (m_2/M)\,\vect{x}\) and \(\vect{x}_2 = -(m_1/M)\,\vect{x}\) in terms of the relative coordinate \(\vect{x}(t) = a(\cos\Omega t, \sin\Omega t, 0)\), so the quadrupole of the two \(\delta\)-function masses reduces to that of a single mass \(\mu\) on the relative orbit:
Using \(\cos^2\theta = \frac12(1+\cos 2\theta)\), \(\sin^2\theta = \frac12(1-\cos 2\theta)\), \(\sin\theta\cos\theta = \frac12\sin 2\theta\), the reduced moment Equation (A.58) has time-dependent part
with \(\mathcal{I}_{zz}\) constant. Everything oscillates at \(2\Omega\): the mass distribution returns to itself after half an orbit, so
Differentiating Equation (A.67) three times,
whence, counting \(\mathcal{I}_{xy} = \mathcal{I}_{yx}\) twice,
constant in time, so the average in Equation (A.63) is trivial:
after eliminating \(\Omega\) with Equation (A.65).
Orbital decay.
The Newtonian orbital energy is, by the virial relation \(K = \frac12\mu a^2\Omega^2 = \frac12\,G m_1 m_2/a = -\frac12 U\),
Radiation drains it, \(\dot E = -P\); differentiating Equation (A.72) and inserting Equation (A.71),
The orbit shrinks, \(\Omega\) grows, \(P \propto a^{-5}\) grows faster: a runaway — the chirp.
Chirp mass and frequency evolution.
Express the evolution in the observable \(f \equiv f_{\mathrm{GW}}\). From Equations (A.65) and (A.68), \(\pi f = \Omega = (GM)^{1/2}a^{-3/2}\), i.e. \(a = (GM)^{1/3}(\pi f)^{-2/3}\). Substituting into Equations (A.71) and (A.72) and introducing the chirp mass
both energy and power depend on the masses only through \(\mathcal{M}\):
Energy balance \(\dot f\,(\dd E/\dd f) = -P\) with \(\dd E/\dd f = -\frac13\mathcal{M}^{5/3}G^{2/3}\pi^{2/3}f^{-1/3}\) then gives
— the announced law \(\dot f \propto \mathcal{M}^{5/3}f^{11/3}\). Measuring \(f\) and \(\dot f\) at any moment of the inspiral therefore yields the chirp mass directly, with no knowledge of the distance; the distance then follows from the amplitude. That amplitude is read off Equations (A.59) and (A.67): for the optimal (face-on) orientation the two polarizations have common magnitude
using \(a^2\Omega^2 = GM/a\) and Equation (A.74).
Numerical estimate (i): a GW150914-like binary.
Take two black holes of \(30\,M_\odot\) each at luminosity distance \(r = 410\,\mathrm{Mpc}\), the parameters of the first LIGO detection to within its uncertainties (the measured values were \(36\,M_\odot\) and \(29\,M_\odot\) at \(410^{+160}_{-180}\,\mathrm{Mpc}\) [Abbott:2016]). We use \(G = 6.67430(15)\times 10^{-11}\,\mathrm{m}^{3}/\mathrm{kg}/\mathrm{s}^{2}\) [Tiesinga:2021], the exact \(c = 299792458\,\mathrm{m}/\mathrm{s}\) of the SI [BIPM:2019], and \(M_\odot = 1.989\times 10^{30}\,\mathrm{kg}\). Then \(m_1 = m_2 = 30\,M_\odot = 5.97\times 10^{31}\,\mathrm{kg}\), \(M = 60\,M_\odot = 1.193\times 10^{32}\,\mathrm{kg}\), \(\mu = 15\,M_\odot\), and
so that
Orbital geometry across the band. With \(GM = 7.97\times 10^{21}\,\mathrm{m}^{3}/\mathrm{s}^{2}\), the separation at gravitational-wave frequency \(f\) is \(a = (GM)^{1/3}(\pi f)^{-2/3}\). At the low edge of the detector band, \(f = 35\,\mathrm{Hz}\): \(a = 2.0\times 10^{7}\,\mathrm{m}/\mathrm{s}^{2/3}\times (110\,/\mathrm{s})^{-2/3} \approx 8.7\times 10^{5}\,\mathrm{m}\) — two thirty-solar-mass black holes, each of Schwarzschild radius \(2Gm/c^2 \approx 89\,\mathrm{km}\), orbiting \(870\,\mathrm{km}\) apart at an orbital frequency of \(17.5\,\mathrm{Hz}\). Near merger the separation approaches a few hundred kilometres: at \(a = 2.5\times 10^{5}\,\mathrm{m}\),
reproducing the observed sweep of GW150914 from \(35\,\mathrm{Hz}\) to about \(250\,\mathrm{Hz}\) [Abbott:2016]. Strain. At \(f = 150\,\mathrm{Hz}\) and \(r = 410\,\mathrm{Mpc} = 1.27\times 10^{25}\,\mathrm{m}\), Equation (A.77) gives, with \((G\mathcal{M}/c^2)^{5/3} = (3.85\times 10^{4})^{5/3}\,\mathrm{m}^{5/3} = 4.4\times 10^{7}\,\mathrm{m}^{5/3}\) and \((\pi f/c)^{2/3} = (1.57\times 10^{-6})^{2/3}\,\mathrm{m}^{-2/3} = 1.35\times 10^{-4}\,\mathrm{m}^{-2/3}\),
for optimal orientation; averaging over source inclination and detector antenna pattern reduces this by a factor of a few, in agreement with the observed peak strain of \(10^{-21}\) [Abbott:2016]. By Equation (A.33) the LIGO arms of \(L = 4\,\mathrm{km}\) then change length by \(\Delta L = \frac12 h L \approx 2\times 10^{-18}\,\mathrm{m}\) [Aasi:2015]. Chirp rate. At \(f = 35\,\mathrm{Hz}\), Equation (A.76) gives
so the residual time to coalescence, \(\tau \approx \frac{3}{8}f/\dot f \approx 0.2\,\mathrm{s}\), matches the observed duration of the signal in band [Abbott:2016]. All three numbers — band, amplitude, chirp — flow from Equations (A.76) and (A.77) with one chirp mass; that a single \(\mathcal{M}\) fits the entire waveform is the quantitative content of the statement that GW150914 was a compact binary coalescence.
Numerical estimate (ii): orbital decay of PSR B1913+16.
For a binary of period \(P_b\), Kepler's law \(P_b \propto a^{3/2}\) gives \(\dot P_b/P_b = \frac{3}{2}\,\dot a/a\), and inserting Equation (A.73) with \(a = (GM)^{1/3}(P_b/2\pi)^{2/3}\) yields the decay of the period. The derivation above assumed a circular orbit; for an eccentric orbit the same energy-balance computation, carried through the Keplerian ellipse harmonics, multiplies the circular result by the enhancement factor \(f(e) = \left(1 + \frac{73}{24}e^2 + \frac{37}{96}e^4\right) (1-e^2)^{-7/2}\), which we quote from the standard treatment [Maggiore:2008]. Thus
The Hulse–Taylor pulsar PSR B1913+16, discovered in 1974 [Hulse:1975], has [Weisberg:2016]
(masses themselves determined from two relativistic timing observables, so the test uses no free parameters). The pieces of Equation (A.83) are
and assembling them with \(G^{5/3} = 1.098\times 10^{-17}\) and \(c^5 = 2.422\times 10^{42}\) (SI),
a dimensionless \(\mathrm{s}/\mathrm{s}\): the orbit of period near eight hours shortens by about \(76\,\mu\mathrm{s}\) per year. The full general-relativistic prediction with the precise timing parameters is \(\dot P_b^{\mathrm{GR}} = -2.40263(5)\times 10^{-12}\), and thirty-five years of timing give an observed intrinsic \(\dot P_b = -2.398(4)\times 10^{-12}\) — a ratio of observed to predicted of \(0.9983(16)\) [Taylor:1982] [Weisberg:2016]. The same parameters give a present gravitational-wave luminosity \(P \approx 7.8\times 10^{24}\,\mathrm{W}\) from Equation (A.71) (times \(f(e)\)) at separation \(a \approx 1.95\times 10^{9}\,\mathrm{m}\) — about \(2\times 10^{-2}\) of the Sun's electromagnetic output, radiated by two neutron stars orbiting within what would fit inside the Sun. Integrating Equation (A.73), the circular estimate gives coalescence after \(a^4/(4\beta) \approx 5.2\times 10^{16}\,\mathrm{s}\) (about \(1.7\) billion years, shortened to roughly \(0.3\) billion by the eccentricity); the binary pulsar is a young GW150914 caught three hundred million years before its chirp.
The two estimates close the loop of this proof: the strain, frequency sweep and chirp rate derived here are the quantities measured directly by the interferometers for GW150914, and the period decay Equation (A.83) is the quantity measured by pulsar timing for PSR B1913+16, matching Section 47.1. The experimental record of both confrontations — apparatus, data and uncertainties — is presented in Experiment: Gravitational Waves.
Oppenheimer–Snyder Collapse: the Interior Matching
This appendix proves Theorem 45.23 (Chapter 45): the collapse of a uniform, pressureless, momentarily static ball of dust is an exact solution of the field equations, consisting of a contracting closed Friedmann–Lemaître–Robertson–Walker (FLRW) interior matched smoothly — induced metric and extrinsic curvature both continuous — to a Schwarzschild exterior across the falling surface [Oppenheimer:1939b]. The exterior half of the argument is already in the chapter: Birkhoff's theorem (Theorem 45.4) forces the outside geometry to be Equation (45.6), and the radial-plunge cycloid (Proposition 45.21) is the surface's worldline seen from outside. What remains — the interior solution, the identification of the boundary, and the junction — follows, in the treatment of [Misner:1973].
The interior: a contracting closed dust universe
Inside the ball the matter is homogeneous and isotropic about every comoving point, so the interior metric is of FLRW form; the momentarily static initial state will force the closed (positively curved) case, and we write it, in comoving coordinates \((\tau,\chi,\theta,\varphi)\) with \(x^{0}:=c\tau\),
which is Equation (45.34); \(\tau\) is proper time on every comoving worldline. Write \(g_{ij}=a^{2}\gamma_{ij}\) with \(\gamma_{ij}\) the metric of the unit 3-sphere and primes for \(\dd/\dd x^{0}\). The Christoffel symbols of Equation (A.89) are, by Equation (13.299),
with \(\tilde{\Gamma}\) the Christoffel symbols of \(\gamma_{ij}\) alone. Contracting the Riemann tensor Equation (13.306) and using the fact that the unit 3-sphere is maximally symmetric with Ricci tensor \(\tilde{R}_{ij}=2\gamma_{ij}\),
Pressureless dust comoving with these coordinates has \(T_{\mu\nu}=\rho\,u_{\mu}u_{\nu}\) with \(u_{\mu}=(-c,0,0,0)\), hence \(T_{00}=\rho c^{2}\), \(T_{ij}=0\) and trace \(T=-\rho c^{2}\). Two facts follow before any field equation is solved. First, since \(\Gamma^{\mu}{}_{00}=0\) in Equation (A.90), the comoving worldlines are geodesics — as they must be: \(\nabla_{\mu}T^{\mu\nu}=0\) for dust is exactly the statement that each grain free-falls, and it is what makes the boundary of the truncated ball a geodesic of both geometries below. Second, in the Ricci form \(R_{\mu\nu}=(8\pi G/c^{4})\left(T_{\mu\nu} -\tfrac{1}{2}g_{\mu\nu}T\right)\) (The Einstein Field Equations) the two components of Equation (A.91) give, restoring \(\tau\) via \(\dot{a}=c\,a'\),
the second obtained by eliminating \(a''\) between the \(00\) and \(ij\) equations. (These are the Friedmann equations of Evidence-Based Cosmology for dust and \(k=+1\); the derivation is kept here so that this proof is self-contained.) Differentiating Equation (A.93) and inserting Equation (A.92) yields \(\dot{\rho}/\rho=-3\dot{a}/a\), i.e.
mass conservation in a contracting volume.
“Momentarily static” means \(\dot{a}=0\) at \(\tau=0\); call the initial values \(a_{\mathrm{m}}\), \(\rho_{\mathrm{m}}\). Then Equation (A.93) at \(\tau=0\) reads
which is why the curvature had to be positive: a momentarily static homogeneous dust ball has nowhere to go but down its own cycloid. With Equations (A.94) and (A.95), Equation (A.93) becomes \(\dot{a}^{2}=c^{2}\left(a_{\mathrm{m}}/a-1\right)\), which the cycloid solves: substituting
gives, exactly as in the proof of Proposition 45.21, \(\dot{a}^{2}=c^{2}(1-\cos\eta)/(1+\cos\eta) =c^{2}\left(a_{\mathrm{m}}/a-1\right)\). The scale factor reaches \(a=0\) at \(\eta=\pi\), i.e. at proper time
the same on every comoving worldline: by homogeneity, all the dust — centre and surface alike — arrives at the singularity simultaneously in comoving time. This is part (iii) of Theorem 45.23.
The boundary, and the matching conditions
Truncate the interior at \(\chi=\chi_{0}\) (with \(\chi_{0}<\pi/2\) for a ball less than half the 3-sphere). The boundary consists of comoving dust worldlines, which are geodesics of the interior geometry by the first fact above; its areal radius — read off the angular part of Equation (A.89) — is
Outside, spherical symmetry and Theorem 45.4 force the Schwarzschild geometry Equation (45.6) with some mass parameter \(M\); the boundary, being made of freely falling dust with no pressure acting on it from either side, must follow a radial timelike geodesic of that exterior, released from rest at \(R_{0}\). By Proposition 45.21 its areal radius is the cycloid Equation (45.31),
The interior assigns the same boundary the areal radius Equation (A.98) with \(a(\tau)\) the cycloid Equation (A.96): the same shape \(R=(R_{0}/2)(1+\cos\eta)\), with proper time \(\tau=\left(a_{\mathrm{m}}/2c\right)(\eta+\sin\eta)\). The two descriptions of one worldline agree for all \(\eta\) iff
which is Equation (45.35). The mass identity follows at once: by Equation (A.95),
The Schwarzschild mass is the initial density times \(\tfrac{4}{3}\pi R_{0}^{3}\) — the Euclidean volume of areal radius \(R_{0}\), not the (larger) proper volume of the curved interior. The deficit is not an error: \(Mc^{2}\) is the total energy including the (negative) gravitational binding energy, and the curved-volume excess of rest mass is exactly what binding has removed.
It remains to show that agreement of the boundary worldline is the whole of the junction condition — that the two geometries fit with no surface layer. Two conditions must hold on the matching hypersurface: the induced metrics must agree, and the extrinsic curvatures must agree.
Induced metric. Coordinatize the boundary by \((\tau,\theta,\varphi)\). From inside, Equation (A.89) restricted to \(\chi=\chi_{0}\) gives \(-c^{2}\dd\tau^{2}+a^{2}\sin^{2}\chi_{0}\,\dd\Omega^{2} =-c^{2}\dd\tau^{2}+R(\tau)^{2}\dd\Omega^{2}\). From outside, along a radial worldline parametrized by its proper time at areal radius \(R(\tau)\), the induced metric is likewise \(-c^{2}\dd\tau^{2}+R(\tau)^{2}\dd\Omega^{2}\). Given Equation (A.100) the two functions \(R(\tau)\) coincide, so the induced metrics agree identically.
Extrinsic curvature. By spherical symmetry the extrinsic curvature has two independent components, \(K^{\tau}{}_{\tau}\) and \(K^{\theta}{}_{\theta}=K^{\varphi}{}_{\varphi}\). The \(\tau\tau\) component measures the normal component of the boundary worldlines' acceleration, and vanishes on both sides because the boundary is geodesic in both geometries — comoving dust inside, radial free fall outside. For the angular component, from inside the outward unit normal is \(n=a^{-1}\pp_{\chi}\); for a hypersurface of constant \(\chi\) in a metric with no \(\chi\)-cross terms, \(K_{ij}=\tfrac{1}{2}\,n^{\chi}\pp_{\chi}g_{ij}\), and with \(g_{\theta\theta}=a^{2}\sin^{2}\chi\),
From outside, write \(f:=1-r_{\mathrm{s}}/R\) and let \(u^{\mu}=(u^{t},u^{r})\) be the falling boundary's four-velocity, with \(u^{t}=\mathcal{E}/(fc^{2})\) and \(c^{2}(u^{r})^{2}=\mathcal{E}^{2}-fc^{4}\) (Proposition 45.21). The outward unit normal \(n^{\mu}=(n^{t},n^{r})\) is fixed by \(n_{\mu}u^{\mu}=0\) and \(n_{\mu}n^{\mu}=1\): orthogonality gives \(n^{t}=u^{r}n^{r}/(f^{2}c^{2}u^{t})\), and substituting into the normalization \(-fc^{2}(n^{t})^{2}+f^{-1}(n^{r})^{2}=1\),
using \(f^{3}c^{2}(u^{t})^{2}=f\,\mathcal{E}^{2}/c^{2}\) in the first step and \(\mathcal{E}^{2}-c^{2}(u^{r})^{2}=fc^{4}\) in the second; so \(n^{r}=\mathcal{E}/c^{2}\). Since the only angular Christoffel symbol involved is \(\Gamma^{\theta}{}_{r\theta}=1/R\) (Equation (45.2)),
The two expressions Equations (A.102) and (A.103) agree iff \(\cos\chi_{0}=\sqrt{1-r_{\mathrm{s}}/R_{0}}\), i.e. iff \(r_{\mathrm{s}}=R_{0}\sin^{2}\chi_{0} =a_{\mathrm{m}}\sin^{3}\chi_{0}\) — precisely the condition Equation (A.100) already imposed by the worldline. The junction is therefore smooth, with no surface shell, exactly when Equation (45.35) holds; parts (i) and (ii) of Theorem 45.23 are proved.
The global picture
The matched spacetime settles the questions the vacuum extension of the chapter left open. The surface crosses \(r=r_{\mathrm{s}}\) at the finite comoving time given by \(\eta_{\mathrm{h}}\) with \(\cos\eta_{\mathrm{h}}=2r_{\mathrm{s}}/R_{0}-1\); from that moment the exterior region contains trapped surfaces (Remark 45.26) and the vacuum part of the spacetime is the regions I and II of the Kruskal diagram (Remark 45.20) — the white-hole region III and the second universe IV are excised, their place taken by the matter-filled interior, which is why they are artefacts of vacuum eternity rather than predictions about collapse. A distant observer, meanwhile, receives the exponentially redshifted last light of Proposition 45.22: the “frozen star” outside, the finite-time singularity inside, with no contradiction between them [Oppenheimer:1939b] [Misner:1973].
The Cantor–Schröder–Bernstein Theorem
This appendix proves Theorem 3.72 of Logic, Sets, and Maps: two sets that inject into each other are equipotent. It is what makes the comparison of cardinals a genuine order relation rather than a pair of unrelated inequalities, and it is used immediately in Corollary A.3 to identify the cardinal of the continuum with that of the power set of \(\N\). The proof is elementary but not obvious — neither injection need be surjective anywhere, and the bijection must be manufactured out of the two of them.
Let \(X\) and \(Y\) be sets and let \(f:X\longrightarrow Y\) and \(g:Y\longrightarrow X\) be injective. Then there exists a bijection \(h:X\longrightarrow Y\). Rests on Definition 3.45 and Proposition 3.48.
Derives Theorem A.1. The obstruction to using \(g^{-1}\) as the bijection is the set of points of \(X\) outside the image of \(g\), where \(g^{-1}\) is undefined. Track where those points propagate under repeated application of \(g\circ f\), and use \(f\) there instead.
Define a sequence of subsets of \(X\) by
and let
Now define \(h:X\longrightarrow Y\) by
\(h\) is well defined. The only question is the second branch. If \(x\notin C\) then in particular \(x\notin C_{0}=X\setminus g(Y)\), so \(x\in g(Y)\); and since \(g\) is injective there is exactly one \(y\in Y\) with \(g(y)=x\). Write \(g^{-1}(x)\) for it.
\(h\) is injective. Each branch is injective on its own domain: \(f\) by hypothesis, and \(g^{-1}\) because \(g\) is a map. It remains to check that the two branches never collide, i.e. that
Suppose not, and let \(f(c)=g^{-1}(x)\) with \(c\in C\) and \(x\notin C\). Applying \(g\) to both sides gives \(x=g(f(c))\). Since \(c\in C\), there is an \(n\) with \(c\in C_{n}\), whence
contradicting \(x\notin C\). This proves Equation (A.107), and with it the injectivity of \(h\).
\(h\) is surjective. Let \(y\in Y\) and consider the point \(g(y)\in X\). There are two cases, and they are exhaustive by Equation (3.25).
If \(g(y)\notin C\), then by the second branch of Equation (A.106),
so \(y\) is attained.
If \(g(y)\in C\), then \(g(y)\in C_{n}\) for some \(n\). That \(n\) cannot be \(0\), because \(g(y)\) lies in \(g(Y)\) while \(C_{0}=X\setminus g(Y)\) does not meet \(g(Y)\). So \(n=m+1\) for some \(m\geq0\) and \(g(y)\in C_{m+1}=g(f(C_{m}))\), i.e. \(g(y)=g(f(c))\) for some \(c\in C_{m}\). Injectivity of \(g\) gives \(y=f(c)\), and since \(c\in C_{m}\subset C\) the first branch of Equation (A.106) yields \(h(c)=f(c)=y\). Again \(y\) is attained.
Being injective and surjective, \(h\) is a bijection by Proposition 3.48.
∎The construction Equations (A.104), (A.105) and (A.106) is explicit: given \(f\) and \(g\), the bijection \(h\) is determined, with no appeal to Axiom 3.78. This matters, because the companion statement that any two cardinals are comparable — that at least one of the two injections always exists — is equivalent to the axiom of choice. CSB says that comparability, where it holds, is antisymmetric; it does not say that any two sets are comparable.
\(\abs{\R}=\abs{\mathcal{P}(\N)}\), that is \(\mathfrak{c}=2^{\aleph_{0}}\). Rests on Theorem A.1 and Corollary 3.68.
Derives Corollary A.3. We exhibit injections in both directions and invoke Theorem A.1.
An injection \(\mathcal{P}(\N)\longrightarrow\R\). Send \(A\subset\N\) to the real number whose base-\(3\) expansion has digit \(1\) at place \(n\) when \(n\in A\) and digit \(0\) otherwise,
If \(A\neq B\), let \(k\) be the least element of their symmetric difference, say \(k\in A\setminus B\). Then
since the digits omitted from \(B\) beyond place \(k\) contribute at most the full geometric tail. Hence \(\Phi(A)\neq\Phi(B)\) and \(\Phi\) is injective. Using digits \(0\) and \(1\) in base \(3\) — rather than base \(2\) — is exactly what makes the tail strictly smaller than the leading term, so that the two expansions cannot coincide.
An injection \(\R\longrightarrow\mathcal{P}(\N)\). Send \(x\in\R\) to its Dedekind cut \(\Psi(x):=\set{q\in\Q\mid q<x}\), injective because \(\Q\) is dense in \(\R\): if \(x<y\) there is a rational strictly between them, lying in \(\Psi(y)\) but not in \(\Psi(x)\). This lands in \(\mathcal{P}(\Q)\) rather than \(\mathcal{P}(\N)\), but \(\abs{\Q}=\aleph_{0}\) by Corollary 3.68, and a bijection \(\N\longrightarrow\Q\) induces a bijection \(\mathcal{P}(\Q)\longrightarrow\mathcal{P}(\N)\) by taking preimages. Composing gives the required injection.
∎The Cantor–Schröder–Bernstein Theorem discharges the proof obligation of Theorem 3.72, and Corollary A.3 supplies the identification \(\mathfrak{c}=2^{\aleph_{0}}\) used in Section 3.5.3 when the continuum hypothesis is stated. The next appendix treats the other half of the foundational material of that chapter.
The Completeness Theorem
This appendix proves Theorem 3.84 of Logic, Sets, and Maps: for first-order logic, semantic consequence and syntactic derivability coincide,
The left-to-right direction is the substantive one and is equivalent to a statement about existence: every consistent theory has a model. That is what the Henkin construction builds, and it builds it out of nothing but the syntax — the elements of the model are the closed terms of the language itself. The theorem is Gödel's doctoral work [Goedel:1930]; the proof given here is Henkin's later and simpler one.
Throughout, the logic is first-order logic with equality, so the equality axioms — reflexivity \(t=t\) and the congruence schemas asserting that equals may be substituted in any function or relation symbol — are part of the logical apparatus and available in every derivation. The language is assumed countable, which every language in this book is; the uncountable case needs only Zorn's lemma (Axiom 3.78) in place of the enumeration in Lemma A.6.
Soundness
If \(T\vdash\varphi\) then \(T\vDash\varphi\). Rests on Axiom 3.34, Definition 3.14 and Equation (3.10).
Derives Theorem A.4. Induction on the length of the derivation, using Axiom 3.34 at the meta-level. Let \(M\) be any structure with \(M\vDash T\), and let \(\psi_{1},\ldots,\psi_{k}\) be a derivation of \(\varphi\) from \(T\). We show every \(\psi_{i}\) is true in \(M\).
If \(\psi_{i}\) is a member of \(T\), it is true in \(M\) by assumption. If it is a logical axiom, it is true in every structure whatever: the propositional axioms are tautologies in the sense of Definition 3.14, and their truth tables were computed once and for all in Section 3.1.2; the quantifier axiom \(\forall x\,\psi(x)\rightarrow\psi(t)\) holds because if \(\psi\) is satisfied by every element it is satisfied by the one denoted by \(t\); and the equality axioms hold because equality is interpreted as genuine identity on the domain.
If \(\psi_{i}\) follows by modus ponens from \(\psi_{j}\) and \(\psi_{j}\rightarrow\psi_{i}\) with \(j<i\), then both are true in \(M\) by the induction hypothesis, and row two of Equation (3.10) — the only row in which a conditional is false — is thereby excluded, so \(\psi_{i}\) is true. If \(\psi_{i}\) is \(\forall x\,\psi\) obtained by generalization from \(\psi\), where \(x\) does not occur free in \(T\), then \(\psi\) is true in \(M\) under every assignment to \(x\), which is what \(\forall x\,\psi\) asserts.
Every line is therefore true in \(M\), in particular the last, so \(M\vDash\varphi\). As \(M\) was an arbitrary model of \(T\), \(T\vDash\varphi\).
∎Soundness is the direction that makes derivations worth performing; it is also the direction used, in Theorem A.30, to conclude that \(Q\) does not prove a sentence false in \(\N\). The remaining direction occupies the rest of this section.
Maximal consistent sets
A set \(S\) of sentences is maximal consistent if it is consistent in the sense of Definition 3.82 and no proper superset of it is consistent. Rests on Definition 3.82.
Every consistent set of sentences extends to a maximal consistent set. Rests on Definition A.5 and Proposition 3.67.
Derives Lemma A.6. The language is countable, so its sentences can be listed \(\varphi_{0},\varphi_{1},\varphi_{2},\ldots\) — there are countably many finite strings over a countable alphabet, by Proposition 3.67. Given a consistent \(S\), define an increasing chain by \(S_{0}:=S\) and
Each \(S_{n}\) is consistent, by induction on \(n\). The base case is the hypothesis on \(S\). For the step, suppose \(S_{n}\) is consistent but both \(S_{n}\cup\set{\varphi_{n}}\) and \(S_{n}\cup\set{\neg\varphi_{n}}\) are not. Inconsistency of the first gives \(S_{n}\vdash\neg\varphi_{n}\) and of the second \(S_{n}\vdash\varphi_{n}\) — in each case by discharging the assumption, which is the deduction theorem — so \(S_{n}\) proves both a sentence and its negation, contradicting its consistency. Hence at least one of the two branches of Equation (A.110) is consistent, and \(S_{n+1}\) is.
Let \(S^{*}:=\bigcup_{n\geq0}S_{n}\). It is consistent: a derivation of a contradiction is a finite object and mentions only finitely many members of \(S^{*}\), all of which lie in some single \(S_{n}\) — the chain is increasing — and that \(S_{n}\) was shown consistent. It is maximal: every sentence of the language is some \(\varphi_{n}\), and Equation (A.110) placed either it or its negation in \(S_{n+1}\), so no consistent sentence can be added.
∎Let \(S^{*}\) be maximal consistent. Then for all sentences \(\varphi,\psi\):
-
exactly one of \(\varphi\) and \(\neg\varphi\) belongs to \(S^{*}\);
-
\(S^{*}\) is deductively closed: if \(S^{*}\vdash\varphi\) then \(\varphi\in S^{*}\);
-
\(\varphi\wedge\psi\in S^{*}\) if and only if both \(\varphi\in S^{*}\) and \(\psi\in S^{*}\);
-
\(\varphi\vee\psi\in S^{*}\) if and only if \(\varphi\in S^{*}\) or \(\psi\in S^{*}\);
-
\(\varphi\rightarrow\psi\in S^{*}\) if and only if \(\varphi\notin S^{*}\) or \(\psi\in S^{*}\).
Rests on Definition A.5 and Proposition 3.22.
Derives Lemma A.7. (1) Not both, by consistency. At least one: if neither is in \(S^{*}\), then by maximality \(S^{*}\cup\set{\varphi}\) is inconsistent, giving \(S^{*}\vdash\neg\varphi\), and likewise \(S^{*}\vdash\varphi\); that contradicts consistency.
(2) If \(S^{*}\vdash\varphi\) then \(S^{*}\cup\set{\varphi}\) is consistent — adding what is already derivable cannot produce a new contradiction — so maximality forces \(\varphi\in S^{*}\).
(3)–(5) Each follows from (1) and (2) together with the propositional tautologies of Proposition 3.22. For (5): if \(\varphi\notin S^{*}\) then \(\neg\varphi\in S^{*}\) by (1), and \(\neg\varphi\vdash \varphi\rightarrow\psi\), so \(\varphi\rightarrow\psi\in S^{*}\) by (2); if \(\psi\in S^{*}\) the same conclusion follows from \(\psi\vdash\varphi\rightarrow\psi\). Conversely if \(\varphi\rightarrow\psi\in S^{*}\) and \(\varphi\in S^{*}\), modus ponens and (2) give \(\psi\in S^{*}\). The other two are proved the same way.
∎Henkin witnesses
A maximal consistent set decides every sentence, but it may assert \(\exists x\,\varphi(x)\) without containing any instance \(\varphi(t)\) — and then no model can be read off it, because there is no term available to serve as the witness. The remedy is to put the witnesses into the language by hand.
Let \(S\) be a consistent set of \(L\)-sentences. Let \(L'\) extend \(L\) by a new constant \(c_{\varphi}\) for each \(L\)-formula \(\varphi(x)\) with exactly one free variable, and let
be the corresponding Henkin axiom. Then \(S\cup\set{H_{\varphi}\mid\varphi}\) is consistent as a set of \(L'\)-sentences. Rests on Definition 3.82.
Derives Lemma A.8. A derivation is finite, so if the extended set were inconsistent, finitely many Henkin axioms would already suffice; it is enough to show that adding one at a time preserves consistency. So suppose \(S'\) is consistent, \(S'\) does not mention the constant \(c:=c_{\varphi}\), and \(S'\cup\set{H_{\varphi}}\) is inconsistent. Then
In particular \(S'\vdash\neg\varphi(c)\).
Now the crucial step. The constant \(c\) occurs nowhere in \(S'\), so it plays no role in the derivation beyond being a name; replacing every occurrence of \(c\) throughout the derivation by a variable \(y\) that appears nowhere in it yields again a correct derivation, and therefore \(S'\vdash\neg\varphi(y)\) with \(y\) free and unconstrained. Generalization is then legitimate — its side condition is exactly that the variable is not free in the premises — and gives
But \(S'\vdash\exists x\,\varphi(x)\) was established above, so \(S'\) is inconsistent, contrary to hypothesis. Hence \(S'\cup\set{H_{\varphi}}\) is consistent.
∎Adding constants creates new formulas, which need witnesses of their own, so one application does not suffice. Iterate: set \(L_{0}:=L\) and \(S_{0}:=S\), let \(L_{n+1}\) and \(S_{n+1}\) be the result of applying Lemma A.8 to \(L_{n}\) and \(S_{n}\), and put \(L_{\infty}:=\bigcup_{n}L_{n}\) and \(S_{\infty}:=\bigcup_{n}S_{n}\). Every \(L_{\infty}\)-formula lies in some \(L_{n}\) and so has its witness constant in \(L_{n+1}\); and \(S_{\infty}\) is consistent because any derivation of a contradiction from it is finite and lives inside some consistent \(S_{n}\). Finally apply Lemma A.6 to \(S_{\infty}\) inside \(L_{\infty}\), giving a set
Extending to a maximal set adds sentences but removes none, so the Henkin axioms survive.
The term model
Everything needed to build a structure is now inside \(T^{*}\). The domain will be the closed terms of \(L_{\infty}\) — expressions such as \(0\), \(f(c_{3})\), \(g(c_{1},f(c_{2}))\) — with terms identified exactly when \(T^{*}\) says they are equal.
Let \(\mathrm{CT}\) be the set of closed \(L_{\infty}\)-terms, and define
The structure \(M\) has domain \(\mathrm{CT}/\!\sim\), the set of classes \([t]\), and interprets
Rests on Definition A.5 and Equation (A.112).
\(\sim\) is an equivalence relation, \(\mathrm{CT}\) is non-empty, and Equations (A.115) and (A.116) do not depend on the representatives chosen. Rests on Definition A.9, Lemma A.7, Definition 3.58 and Theorem 3.60.
Derives Lemma A.10. Reflexivity: \(t=t\) is a logical axiom, so it lies in \(T^{*}\) by Lemma A.7(2). Symmetry and transitivity: the equality congruence schemas give \(\vdash(t=s)\rightarrow(s=t)\) and \(\vdash\left((t=s)\wedge(s=u)\right)\rightarrow(t=u)\), and deductive closure transfers these to membership in \(T^{*}\). So \(\sim\) satisfies the three clauses of Definition 3.58 and, by Theorem 3.60, partitions \(\mathrm{CT}\) into classes.
\(\mathrm{CT}\) is non-empty because \(L_{\infty}\) contains the Henkin constants, of which there is at least one; a structure must have a non-empty domain, and this is where that is secured.
Well-definedness: suppose \(t_{i}\sim s_{i}\) for each \(i\), i.e.\ \((t_{i}=s_{i})\in T^{*}\). The congruence schema for the function symbol \(f\) gives
so by Lemma A.7(2),(3) the conclusion lies in \(T^{*}\) and \([f(\vect{t})]=[f(\vect{s})]\). The congruence schema for the relation symbol \(R\) gives the same conclusion for Equation (A.116).
∎The truth lemma
For every \(L_{\infty}\)-sentence \(\varphi\),
Rests on Definition A.9, Lemma A.10, Lemma A.8 and Lemma A.7.
Derives Lemma A.11. Induction on the number of connectives and quantifiers in \(\varphi\). The statement being proved is about all sentences of that complexity at once, which is what makes the quantifier case go through.
Atomic. A closed atomic sentence is \(R(t_{1},\ldots,t_{n})\) or \(t=s\). For the first, Equation (A.116) says \(M\vDash R(t_{1},\ldots,t_{n})\) precisely when \(R(t_{1},\ldots,t_{n})\in T^{*}\). For the second, \(M\vDash t=s\) means \([t]=[s]\), which by Equation (A.113) means \((t=s)\in T^{*}\). Note that the interpretation of a closed term \(t\) in \(M\) is \([t]\) itself, by induction on the term using Equations (A.114) and (A.115).
Negation. \(M\vDash\neg\varphi\) iff \(M\nvDash\varphi\), iff \(\varphi\notin T^{*}\) by the induction hypothesis, iff \(\neg\varphi\in T^{*}\) by Lemma A.7(1).
Conditional. \(M\vDash\varphi\rightarrow\psi\) iff \(M\nvDash\varphi\) or \(M\vDash\psi\), which by the induction hypothesis is \(\varphi\notin T^{*}\) or \(\psi\in T^{*}\), which by Lemma A.7(5) is \(\varphi\rightarrow\psi\in T^{*}\). The cases of \(\wedge\) and \(\vee\) are identical, using clauses (3) and (4).
Existential. Every element of the domain is \([t]\) for some closed term \(t\), so
Each \(\varphi(t)\) has fewer quantifiers than \(\exists x\,\varphi(x)\), so the induction hypothesis applies to it and the right-hand side is equivalent to \(\varphi(t)\in T^{*}\) for some closed \(t\).
If \(\exists x\,\varphi(x)\in T^{*}\), the Henkin axiom Equation (A.111) for \(\varphi\) is in \(T^{*}\) by Equation (A.112), and modus ponens with Lemma A.7(2) puts \(\varphi(c_{\varphi})\in T^{*}\); that is the required witness. Conversely if \(\varphi(t)\in T^{*}\) for some closed \(t\), then since \(\vdash\varphi(t)\rightarrow\exists x\,\varphi(x)\) is a logical axiom, deductive closure gives \(\exists x\,\varphi(x)\in T^{*}\).
Universal. \(\forall x\,\varphi\) is treated as \(\neg\exists x\,\neg\varphi\), already covered.
∎Completeness, and compactness
Every consistent set of sentences has a model. Rests on Lemma A.11, Definition A.9 and Equation (A.112).
Derives Theorem A.12. Given consistent \(S\), build \(T^{*}\) as in Equation (A.112) and let \(M\) be the term structure of Definition A.9. By Lemma A.11, \(M\) satisfies exactly the sentences of \(T^{*}\), and \(S\subset T^{*}\), so \(M\vDash S\). The model is a structure for \(L_{\infty}\); forgetting the interpretations of the added constants leaves a model of \(S\) in the original language.
∎\(T\vDash\varphi\) implies \(T\vdash\varphi\). With Theorem A.4, this establishes Equation (A.109). Rests on Theorems A.4 and A.12.
Derives Theorem A.13. Contrapositive. Suppose \(T\nvdash\varphi\). Then \(T\cup\set{\neg\varphi}\) is consistent — an inconsistency would give \(T\vdash\varphi\) by reductio — so by Theorem A.12 it has a model \(M\). That \(M\) satisfies \(T\) and falsifies \(\varphi\), so \(T\nvDash\varphi\).
∎If every finite subset of \(T\) has a model, then \(T\) has a model. Rests on Theorems A.4 and A.12.
Derives Corollary A.14. If \(T\) had no model it would, by Theorem A.12, be inconsistent, so some derivation would produce a contradiction from it. A derivation is a finite object and cites only finitely many members of \(T\); those form a finite inconsistent subset, which by Theorem A.4 has no model. This contradicts the hypothesis.
∎Compactness is the corollary a physicist is most likely to meet, because it manufactures models nobody intended. Take the theory of the ordered field \(\R\) (Real Analysis), adjoin a new constant \(\epsilon\) and the infinitely many axioms \(0<\epsilon<1/\overline{n}\), one for each natural number \(n\). Any finite subset mentions only finitely many of them and is satisfied in \(\R\) itself, by choosing \(\epsilon\) small enough. Compactness then hands over a model of the entire set: an ordered field elementarily equivalent to \(\R\) — satisfying every first-order sentence that \(\R\) satisfies — but containing an element smaller than every positive rational. That is a rigorous infinitesimal, and the non-standard analysis built on it is a legitimate alternative foundation for the calculus. The same argument shows no first-order theory can pin down \(\N\) up to isomorphism, which is the model-theoretic shadow of Theorem 3.87.
The Completeness Theorem discharges the proof obligation of Theorem 3.84, and it is worth restating why this sits beside Theorem 3.87 without tension. Completeness is a statement about the logic: the calculus derives everything that follows from the premises, whatever the premises are. Incompleteness is a statement about a particular theory: the axioms of arithmetic do not, among themselves, settle every arithmetical sentence. A complete calculus applied to incomplete axioms yields exactly what Equation (3.88) describes, a sentence that neither follows nor fails to follow — and by Theorem A.12 there are then models of the axioms in which it holds and models in which it does not.
Arithmetization and the Incompleteness Theorems
This appendix supplies the technical machinery behind Sections 3.8 and 3.9 of Logic, Sets, and Maps: the coding of syntax into arithmetic, the diagonal lemma (Lemma 3.86), Rosser's strengthening of the first incompleteness theorem, the derivability conditions from which the second theorem follows, and Church's theorem on the undecidability of first-order validity (Theorem 3.98). Throughout, \(T\) is an effectively axiomatized first-order theory in the language of arithmetic \(\set{0,S,+,\cdot}\) of Definition 3.83 that contains Robinson arithmetic \(Q\): the successor, addition and multiplication axioms of that definition, without the induction schema, together with the one axiom that induction would otherwise supply,
PA proves Equation (A.118) by induction, so once the schema is dropped it must be assumed outright; without it the bounded-quantifier lemmas used below all fail. \(Q\) is therefore finitely axiomatized, which Theorem A.30 depends on, and it is already strong enough for everything except the derivability conditions of The derivability conditions and the second theorem, which need PA itself. The order is the defined relation
Everything below applies verbatim to PA and, through the standard translation of arithmetic into set theory, to ZFC.
Gödel numbering
Fix an injective assignment of positive integers to the finitely many symbols of the language, and code a finite string \(\sigma_{1}\sigma_{2}\cdots\sigma_{k}\) of symbols by
where \(p_{i}\) is the \(i\)-th prime and \(c(\sigma)\) the code of the symbol \(\sigma\). Unique factorization makes Equation (A.120) injective and makes decoding a matter of computing prime exponents. A finite sequence of strings — a derivation — is coded the same way, one prime per member. We write \(\overline{n}\) for the numeral of \(n\), the closed term \(S^{n}(0)\) of the object language.
The choice of scheme is immaterial, as Remark 3.39 warned: any injective coding whose encoding and decoding operations are computable serves, and all the results below are invariant under changing it.
Under Equation (A.120) the following relations on natural numbers are decidable in the sense of Definition 3.93: “\(n\) codes a formula”, “\(n\) codes a sentence”, “\(n\) codes an axiom of \(T\)”, “\(m\) codes a \(T\)-derivation of the formula coded by \(n\)”, and the function \(\mathrm{sub}(m,n)\) returning the code of the result of substituting the numeral \(\overline{n}\) for the free variable of the formula coded by \(m\). Rests on Equation (A.120), Definition 3.93 and Definition 3.81.
Derivation. Derives Proposition A.15. Each is settled by a bounded search through the prime factorization of the argument, composed with the syntactic tests of Definition 3.80, which are finite string comparisons. Axiomhood is decidable precisely because \(T\) was assumed effectively axiomatized (Definition 3.81); this is the only hypothesis of the theorem that is not automatic, and it is where the assumption earns its place. Being a derivation is the conjunction of finitely many such tests, one per line of the coded sequence, each of which asks whether a line is an axiom or follows from earlier lines by one of the finitely many rules.
∎Representability
Computability of a relation is a statement in the metatheory; what the incompleteness proofs need is that \(T\) can talk about the relation and prove the right instances.
A relation \(R\subset\N^{k}\) is representable in \(T\) if there is a formula \(\rho(x_{1},\ldots,x_{k})\) such that for all \(n_{1},\ldots,n_{k}\in\N\),
A function is representable if its graph is, together with a provable uniqueness clause. Rests on Definition 3.80.
Every decidable relation and every computable total function is representable in \(Q\), and hence in every theory containing \(Q\). Moreover \(Q\) is \(\Sigma_{1}\)-complete: every true sentence of the form \(\exists x_{1}\cdots\exists x_{k}\,\delta\), with \(\delta\) containing only bounded quantifiers, is provable in \(Q\). Rests on Definition A.16, Theorem A.25, Lemma A.20 and Definition 3.93.
The proof occupies the rest of this subsection. It is the one genuinely long construction in the chapter, and it is worth saying at the outset why it cannot be shortened: \(Q\) has no induction, so nothing may be proved inside \(Q\) by induction on a variable. Every induction below is performed in the metatheory, on a concrete numeral, and what it produces is not one theorem of \(Q\) but a recipe yielding a separate \(Q\)-derivation for each numeral. That is exactly the strength the incompleteness proofs need, and no more.
Bounded and existential formulas
A bounded quantifier is one of the forms \(\forall x\,(x\leq t\rightarrow\cdots)\) or \(\exists x\,(x\leq t\wedge\cdots)\), abbreviated \(\forall x\leq t\) and \(\exists x\leq t\), where \(t\) is a term not containing \(x\). A formula is \(\Delta_{0}\) if all its quantifiers are bounded, and \(\Sigma_{1}\) if it has the form \(\exists x_{1}\cdots\exists x_{k}\,\delta\) with \(\delta\) of class \(\Delta_{0}\). Rests on Equation (A.119).
The point of the restriction is that a \(\Delta_{0}\) sentence can be checked by a finite computation: every quantifier ranges over an explicitly bounded set of numbers.
For all \(m,n\in\N\):
-
if \(m+n=p\) then \(Q\vdash\overline{m}+\overline{n}=\overline{p}\), and similarly \(Q\vdash\overline{m}\cdot\overline{n}=\overline{q}\) when \(mn=q\);
-
if \(m\neq n\) then \(Q\vdash\overline{m}\neq\overline{n}\);
-
if \(m\leq n\) then \(Q\vdash\overline{m}\leq\overline{n}\), and if \(m>n\) then \(Q\vdash\neg(\overline{m}\leq\overline{n})\);
-
\(Q\) proves the bounded expansion
\begin{equation} \tag{A.123} \forall z\,\bigl(z\leq\overline{m}\rightarrow z=\overline{0}\vee z=\overline{1}\vee\cdots\vee z=\overline{m}\bigr)\ep \end{equation}
Rests on Definition 3.83, Equation (A.118) and Equation (A.119).
Derives Lemma A.19. (1) Induction in the metatheory on \(n\). For \(n=0\) the axiom \(x+0=x\) gives \(\overline{m}+\overline{0}=\overline{m}\). If \(Q\vdash\overline{m}+\overline{n}=\overline{m+n}\), then the axiom \(x+Sy=S(x+y)\) gives \(\overline{m}+\overline{n+1}=S(\overline{m}+\overline{n}) =S\overline{m+n}=\overline{m+n+1}\). Multiplication is the same argument using \(x\cdot0=0\) and \(x\cdot Sy=x\cdot y+x\). Note what happened: for each fixed pair \((m,n)\) this is a finite chain of equalities, hence a genuine \(Q\)-derivation; the induction that generated the chain was ours, not \(Q\)'s.
(2) Suppose \(m\neq n\), say \(m<n\). Metatheoretic induction on \(m\). If \(m=0\) then \(\overline{n}\) is \(S\overline{n-1}\) and the axiom \(Sx\neq0\) applies directly. If \(m>0\), both numerals are successors, and the axiom \(Sx=Sy\rightarrow x=y\) contraposed reduces the claim for \((m,n)\) to the claim for \((m-1,n-1)\), which the induction hypothesis supplies.
(3) If \(m\leq n\), take \(z:=\overline{n-m}\); by (1), \(Q\vdash\overline{n-m}+\overline{m}=\overline{n}\), so \(Q\vdash\exists z\,(z+\overline{m}=\overline{n})\), which is \(\overline{m}\leq\overline{n}\) by Equation (A.119). If \(m>n\), we must refute \(\exists z\,(z+\overline{m}=\overline{n})\). By (4), applied with \(\overline{n}\), any such \(z\) would satisfy \(z=\overline{j}\) for some \(j\leq n\); but then \(z+\overline{m}=\overline{j+m}\) by (1), and \(j+m\geq m>n\), so \(\overline{j+m}\neq\overline{n}\) by (2). As \(j\) ranges over finitely many values, this is a finite case analysis inside \(Q\).
(4) Metatheoretic induction on \(m\). For \(m=0\): suppose \(z\leq\overline{0}\), i.e. \(w+z=\overline{0}\) for some \(w\). If \(z\neq\overline{0}\) then Equation (A.118) gives \(z=Sv\) for some \(v\), whence \(w+Sv=S(w+v)\) by the addition axiom, so \(\overline{0}\) is a successor, contradicting \(Sx\neq0\). Hence \(z=\overline{0}\). This is the step that fails without Equation (A.118): nothing else in \(Q\) rules out an element that is neither \(0\) nor a successor.
For the step, assume Equation (A.123) for \(m\) and suppose \(z\leq\overline{m+1}\), say \(w+z=\overline{m+1}=S\overline{m}\). If \(z=\overline{0}\) we are done. Otherwise \(z=Sv\) by Equation (A.118), so \(S(w+v)=S\overline{m}\), and injectivity of \(S\) gives \(w+v=\overline{m}\), i.e. \(v\leq\overline{m}\). The induction hypothesis then yields \(v=\overline{j}\) for some \(j\leq m\), so \(z=Sv=\overline{j+1}\) with \(j+1\leq m+1\).
∎Every true \(\Sigma_{1}\) sentence is provable in \(Q\). Rests on Definition A.18, Lemma A.19 and Proposition 3.22.
Derives Lemma A.20. First the \(\Delta_{0}\) case, by induction in the metatheory on the structure of the formula. Atomic sentences are equations and inequations between closed terms; evaluating the terms and applying Lemma A.19(1)–(3) settles them, and a true one is proved while a false one has its negation proved. The connectives are immediate: if the parts are decided, Lemma A.7-style propositional reasoning decides the compound, using only the tautologies of Proposition 3.22.
For a bounded quantifier, let \(\forall z\leq\overline{m}\,\psi(z)\) be true. Then \(\psi(\overline{j})\) is true for every \(j\leq m\), and each is proved by the induction hypothesis; Equation (A.123) converts the finitely many instances into the universally quantified statement, since any \(z\) below \(\overline{m}\) is provably one of the \(\overline{j}\). If instead \(\exists z\leq\overline{m}\,\psi(z)\) is true, some witness \(\overline{j}\) with \(j\leq m\) has \(\psi(\overline{j})\) true and provable, and Lemma A.19(3) supplies \(\overline{j}\leq\overline{m}\).
Now the general case. Let \(\exists x_{1}\cdots\exists x_{k}\,\delta\) be true, and let \(n_{1},\ldots,n_{k}\) be witnesses. Then \(\delta(\overline{n_{1}},\ldots,\overline{n_{k}})\) is a true \(\Delta_{0}\) sentence, hence provable, and existential generalization gives the \(\Sigma_{1}\) sentence.
∎Coding sequences: the $\beta$-function
Representing a recursively defined function requires quantifying over the sequence of intermediate values, and arithmetic has no sequences. Gödel's device turns a sequence of arbitrary length into a pair of numbers.
Let \(m_{0},\ldots,m_{n}\) be pairwise coprime positive integers and let \(r_{0},\ldots,r_{n}\) satisfy \(0\leq r_{i}<m_{i}\). Then some \(a\) has \(a\equiv r_{i}\pmod{m_{i}}\) for every \(i\).
Derives Lemma A.21. Put \(P:=m_{0}m_{1}\cdots m_{n}\) and \(P_{i}:=P/m_{i}\). Since the moduli are pairwise coprime, \(P_{i}\) is coprime to \(m_{i}\), so there is \(u_{i}\) with \(P_{i}u_{i}\equiv1\pmod{m_{i}}\). Set \(a:=\sum_{i}r_{i}P_{i}u_{i}\). Modulo \(m_{i}\) every term with \(j\neq i\) vanishes, because \(m_{i}\) divides \(P_{j}\), and the remaining term is \(r_{i}P_{i}u_{i}\equiv r_{i}\).
∎the remainder of \(a\) on division by \(1+(i+1)b\).
For every finite sequence \(s_{0},s_{1},\ldots,s_{n}\) of natural numbers there exist \(a,b\) with \(\beta(a,b,i)=s_{i}\) for all \(i\leq n\). Rests on Definition A.22 and Lemma A.21.
Derives Lemma A.23. Let \(c:=\max\set{n,s_{0},\ldots,s_{n}}\) and put \(b:=c!\). Write \(m_{i}:=1+(i+1)b\) for \(i=0,\ldots,n\).
The moduli are pairwise coprime. Suppose a prime \(p\) divides both \(m_{i}\) and \(m_{j}\) with \(i<j\leq n\). Then \(p\) divides their difference,
so \(p\) divides \((j-i)\) or \(p\) divides \(b\). If \(p\) divided \(b\) then, since \(p\) also divides \(m_{i}=1+(i+1)b\), it would divide the difference \(m_{i}-(i+1)b=1\), which is impossible. Hence \(p\) divides \(j-i\). But \(0<j-i\leq n\leq c\), so \(j-i\) is one of the factors of \(c!=b\), giving \(p\mid b\) — the case just excluded. No such prime exists, so \(\gcd(m_{i},m_{j})=1\).
The residues fit. Each \(s_{i}\leq c\leq c!=b<1+(i+1)b=m_{i}\), so \(s_{i}\) is a legitimate remainder modulo \(m_{i}\).
Lemma A.21 now supplies \(a\) with \(a\equiv s_{i}\pmod{m_{i}}\) for every \(i\leq n\), and since \(0\leq s_{i}<m_{i}\) that congruence says exactly \(\mathrm{rem}(a,m_{i})=s_{i}\), i.e. \(\beta(a,b,i)=s_{i}\).
∎The relation \(\beta(a,b,i)=v\) is defined by the \(\Delta_{0}\) formula
and \(Q\) proves that the value is unique: for each numeral triple, at most one \(v\) satisfies it. Rests on Definition A.22, Definition A.18 and Lemma A.19.
Derives Lemma A.24. Equation (A.125) is the division algorithm written out: \(v\) is the remainder exactly when \(a=q\cdot m+v\) with \(v<m\). The quotient \(q\) is at most \(a\), so bounding its quantifier by \(a\) loses nothing, and every quantifier in Equation (A.125) is therefore bounded. For uniqueness at numeral arguments: if \(v,v'\) both satisfied it with \(v<v'\), subtracting gives \((q-q')m=v'-v\) with \(0<v'-v<m\), impossible for a multiple of \(m\); and by Lemma A.19 this finite arithmetic is carried out inside \(Q\) for each numeral instance.
∎Representing the computable functions
Every primitive recursive function is representable in \(Q\). Rests on Definition A.16, Lemma A.19, Lemma A.23, Lemma A.24 and Lemma A.20.
Derives Theorem A.25. Induction in the metatheory on the construction of the function.
Initial functions. The zero function is represented by \(y=\overline{0}\); the successor by \(y=Sx\); the projection \(\pi^{n}_{i}\) by \(y=x_{i}\). In each case Equation (A.121) holds because the defining equation is provable for numerals by Lemma A.19(1), and Equation (A.122) because a wrong value is refuted by Lemma A.19(2).
Composition. Let \(f(\vect{x})=h(g_{1}(\vect{x}),\ldots, g_{k}(\vect{x}))\) with \(h,g_{1},\ldots,g_{k}\) represented by \(\eta,\gamma_{1},\ldots,\gamma_{k}\). Take
At numeral arguments the \(z_{i}\) are provably the values of the \(g_{i}\), by the uniqueness clause of Definition A.16, and \(\eta\) then pins \(y\); the negative clause follows likewise.
Primitive recursion. Let
with \(g,h\) represented by \(\gamma,\eta\). The obstruction is that \(f\)'s value at \(n\) depends on the whole computation \(f(\vect{x},0),\ldots, f(\vect{x},n)\), which is a sequence. Encode it with \(\beta\): say that \(y\) is the value if there is a pair \((a,b)\) coding a sequence whose \(0\)th entry is \(g(\vect{x})\), whose entries step according to \(h\), and whose \(n\)th entry is \(y\). Formally,
For numeral arguments \(\vect{m},n\) with true value \(p\), Lemma A.23 supplies \(a,b\) coding the actual computation sequence, every conjunct of Equation (A.127) is then a true \(\Delta_{0}\) or \(\Sigma_{1}\) statement about numerals, and Lemma A.20 makes it provable — giving Equation (A.121). For Equation (A.122), if \(y=\overline{p'}\) with \(p'\neq p\), then \(B(a,b,\overline{n},y)\) together with the uniqueness of Lemma A.24 and the induction hypothesis on \(\eta\) forces \(\overline{p'}=\overline{p}\), refuted by Lemma A.19(2).
∎Proof of Theorem A.17. Derives Theorem A.17. A decidable relation has a computable indicator function (Definition 3.93), so it suffices to treat total computable functions.
By Kleene's normal form, every total computable \(f\) can be written \(f(\vect{x})=V\bigl(\mu z\,[\,C(\vect{x},z)=0\,]\bigr)\) with \(V\) and \(C\) primitive recursive, where \(\mu z\) denotes the least \(z\) making the bracket true — \(C\) tests whether \(z\) codes a halting computation of the machine for \(f\) on \(\vect{x}\), which by Proposition A.15 is a primitive recursive test, and \(V\) reads the output off it. Since \(f\) is total, such a \(z\) exists for every \(\vect{x}\). Let \(\chi\) and \(\nu\) represent \(C\) and \(V\), which Theorem A.25 provides, and set
At numeral arguments, let \(z_{0}\) be the true least witness. Then \(\chi(\vect{\overline{m}},\overline{z_{0}},\overline{0})\) is provable, each of the finitely many \(\neg\chi(\vect{\overline{m}},\overline{w}, \overline{0})\) with \(w<z_{0}\) is provable, and Equation (A.123) assembles them into the bounded universal clause. So Equation (A.121) holds. If \(y\) is given a wrong numeral value, the uniqueness clauses for \(\chi\) and \(\nu\) refute it, giving Equation (A.122).
The \(\Sigma_{1}\)-completeness half of the theorem is Lemma A.20. Every step used only the seven axioms of \(Q\); no induction inside the theory was required, because each claim was proved separately for each numeral by an induction carried out here rather than there.
Axiom 3.94 entered nowhere as a step. It is needed only to identify the informal notion of “computable procedure” with the formal one in the statement of the theorem; the proof itself concerns Turing machines and primitive recursion throughout.
∎This discharges the last outstanding obligation of the chapter. It is what licenses the substitution formula of Theorem A.26, the provability predicate Equation (A.129), and the \(\Sigma_{1}\) step in both Theorem 3.87 and Theorem A.30.
By Proposition A.15 and Theorem A.17 there is a formula \(\mathrm{Proof}_{T}(y,x)\) representing “\(y\) codes a \(T\)-derivation of the formula coded by \(x\)”, and we set
which is the formula promised in Equation (3.86). Note that \(\mathrm{Prov}_{T}\) is \(\Sigma_{1}\) but not decidable: the existential quantifier is unbounded, and searching for a proof is the semi-decidable procedure of Section 3.9.3.
The diagonal lemma
Let \(\psi(x)\) be any formula with one free variable. Then there is a sentence \(\sigma\) with
Rests on Proposition A.15 and Theorem A.17.
Derives Theorem A.26. By Proposition A.15 the substitution function \(\mathrm{sub}\) is computable, so by Theorem A.17 it is representable: there is a formula \(\mathrm{Sub}(x,y,z)\) such that for all \(m,n\),
Given \(\psi\), define the auxiliary formula in the single free variable \(x\)
let \(k:=\ulcorner\theta\urcorner\) be its own Gödel number, and put
The point of the construction is the arithmetical identity
which holds because substituting the numeral \(\overline{k}\) into the formula coded by \(k\) — namely \(\theta\) — produces exactly \(\sigma\).
Now instantiate Equation (A.131) at \(m=n=k\) and use Equation (A.134):
Substituting this equivalence into Equation (A.133) gives
and the right-hand side is provably equivalent to \(\psi(\ulcorner\sigma\urcorner)\), since a universally quantified variable constrained to a single value may be replaced by that value. This is Equation (A.130).
∎No circularity is involved. The sentence \(\sigma\) was written down explicitly in Equation (A.133); it mentions only the numeral \(\overline{k}\), a number computed before \(\sigma\) existed. Self-reference is achieved without self-mention, and this is the whole trick.
Rosser's form of the first theorem
Theorem 3.87 was proved in Logic, Sets, and Maps using soundness to rule out \(T\vdash\neg G_{T}\). Gödel's own proof replaced soundness by \(\omega\)-consistency — no formula \(\varphi\) has \(T\vdash\exists x\,\varphi(x)\) together with \(T\vdash\neg\varphi (\overline{n})\) for every \(n\) — and Rosser removed even that [Rosser:1936].
Let \(T\supset Q\) be effectively axiomatized and merely consistent. Then there is a sentence \(R\) with \(T\nvdash R\) and \(T\nvdash\neg R\). Rests on Theorem A.26, Theorem A.17 and Lemma A.19.
Derives Theorem A.27. Let \(\mathrm{neg}(x)\) be the computable function taking the code of a formula to the code of its negation, represented as in Theorem A.17. Apply Theorem A.26 to
obtaining \(R\) with \(T\vdash R\leftrightarrow\psi(\ulcorner R\urcorner)\). In words, \(R\) says: for every proof of me there is a no-longer proof of my negation.
\(T\nvdash R\). Suppose \(T\vdash R\), with a derivation coded by \(m\). Then \(\mathrm{Proof}_{T}(\overline{m},\ulcorner R\urcorner)\) is true, hence provable by Equation (A.121). Combining with \(R\) itself, \(T\) proves \(\exists z\leq\overline{m}\ \mathrm{Proof}_{T}(z,\ulcorner\neg R\urcorner)\). But \(T\) is consistent, so \(T\nvdash\neg R\), so for each \(j\leq m\) the statement \(\neg\mathrm{Proof}_{T}(\overline{j},\ulcorner\neg R\urcorner)\) is true and hence provable by Equation (A.122). Since \(Q\) proves the bounded expansion \(\forall z\,(z\leq\overline{m}\rightarrow z=\overline{0}\vee\cdots\vee z=\overline{m})\), the theory proves the negation of the displayed existential. So \(T\) is inconsistent — contradiction.
\(T\nvdash\neg R\). Suppose \(T\vdash\neg R\), with derivation coded by \(m\); then \(T\vdash\mathrm{Proof}_{T}(\overline{m},\ulcorner\neg R\urcorner)\), and therefore
taking \(z=\overline{m}\) as the witness. On the other hand \(\neg R\) is provably equivalent to \(\exists y\,\bigl(\mathrm{Proof}_{T}(y,\ulcorner R\urcorner)\wedge \neg\exists z\leq y\,\mathrm{Proof}_{T}(z,\ulcorner\neg R\urcorner)\bigr)\), so with Equation (A.136) and the provable trichotomy of \(\leq\), the theory proves \(\exists y<\overline{m}\ \mathrm{Proof}_{T}(y,\ulcorner R\urcorner)\). But consistency gives \(T\nvdash R\), so each \(\neg\mathrm{Proof}_{T}(\overline{j},\ulcorner R\urcorner)\) with \(j<m\) is true and provable, and the bounded expansion again refutes the existential. So \(T\) is inconsistent — contradiction.
∎Bare consistency therefore suffices, and \(R\) is undecided by \(T\) exactly as Equation (3.88) asserts.
The derivability conditions and the second theorem
For \(T\supset\) PA with the coding of Gödel numbering, the three conditions Equations (3.93), (3.94) and (3.95) hold. Rests on Theorem A.17 and Equation (A.129).
Derivation. Derives Proposition A.28. Equation (3.93) is \(\Sigma_{1}\)-completeness (Theorem A.17): a derivation of \(\varphi\) is a concrete finite object, so \(\mathrm{Prov}_{T}(\ulcorner\varphi\urcorner)\) is a true \(\Sigma_{1}\) sentence and is therefore provable. Equation (3.94) formalizes the observation that two derivations, one of \(\varphi\rightarrow\chi\) and one of \(\varphi\), can be concatenated and extended by one application of modus ponens to give a derivation of \(\chi\); the concatenation is a computable operation on codes and PA proves that it does what it does, by induction on the length of the derivations. Equation (3.95) is Equation (3.93) formalized inside PA, and is the laborious one: it requires proving in PA the general \(\Sigma_{1}\)-completeness statement rather than using it instance by instance, which is done by induction on the structure of \(\Sigma_{1}\) formulas.
∎If \(T\supset\) PA is effectively axiomatized and consistent, then \(T\nvdash\mathrm{Con}_{T}\). Rests on Proposition A.28, Theorem 3.87 and Equation (3.89).
Derives Theorem A.29. Let \(G\) be the Gödel sentence of Equation (3.89),
We derive Equation (3.92) from the conditions.
From Equation (A.137), \(T\vdash G\rightarrow\neg\mathrm{Prov}_{T} (\ulcorner G\urcorner)\), equivalently \(T\vdash\mathrm{Prov}_{T}(\ulcorner G\urcorner)\rightarrow\neg G\). Applying Equation (3.93) to this theorem and then Equation (3.94),
By Equation (3.95),
and chaining Equation (A.139) with Equation (A.138),
A theory proving both a sentence and its negation proves everything, and this fact is itself formalizable, giving \(T\vdash\bigl(\mathrm{Prov}_{T}(\ulcorner G\urcorner)\wedge \mathrm{Prov}_{T}(\ulcorner\neg G\urcorner)\bigr)\rightarrow \mathrm{Prov}_{T}(\ulcorner 0=1\urcorner)\). With Equation (A.140),
whose contrapositive, recalling \(\mathrm{Con}_{T}=\neg\mathrm{Prov}_{T}(\ulcorner 0=1\urcorner)\) from Equation (3.90), reads
Combining Equation (A.142) with Equation (A.137) gives \(T\vdash\mathrm{Con}_{T}\rightarrow G\), which is Equation (3.92).
Finally, suppose \(T\vdash\mathrm{Con}_{T}\). Modus ponens would yield \(T\vdash G\), contradicting the first half of Theorem 3.87, which requires only consistency. Hence \(T\nvdash\mathrm{Con}_{T}\).
∎Church's theorem
The set of valid sentences of first-order logic is undecidable [Church:1936b] [Turing:1937]. Rests on Theorems 3.84, 3.96 and A.17.
Derives Theorem A.30. The proof has two moves: reduce validity to provability in a finitely axiomatized theory, then reduce halting to that.
Validity to provability in \(Q\). Robinson arithmetic \(Q\) has finitely many axioms; let \(\chi\) be their conjunction, a single sentence. By the deduction theorem, for every sentence \(\varphi\),
the last step by Theorem 3.84. A decision procedure for validity would therefore decide provability in \(Q\). This is where finite axiomatizability is essential; the argument does not run for PA, whose induction schema is infinite.
Halting to provability in \(Q\). The halting set
is semi-decidable: simulate \(M\) on \(w\) with Theorem 3.95 and halt when it halts. Every semi-decidable set is definable by a \(\Sigma_{1}\) formula — the existential quantifier ranging over codes of terminating computations, whose verification is decidable by Proposition A.15 — so fix \(\eta(x,y)\) of that form with
Then:
-
If \((m,w)\in H\), the sentence \(\eta(\overline{m},\overline{w})\) is a true \(\Sigma_{1}\) sentence, so \(Q\vdash\eta(\overline{m},\overline{w})\) by \(\Sigma_{1}\)-completeness (Theorem A.17).
-
If \((m,w)\notin H\), the sentence is false in \(\N\). Since \(\N\) is a model of \(Q\), soundness (Theorem 3.84) gives \(Q\nvdash\eta(\overline{m},\overline{w})\).
Hence \(Q\vdash\eta(\overline{m},\overline{w})\) if and only if \(M\) halts on \(w\). A decision procedure for \(Q\)-provability would decide the halting problem, contradicting Theorem 3.96; by Equation (A.143) a decision procedure for validity would supply one. Therefore no decision procedure for validity exists.
∎Arithmetization and the Incompleteness Theorems discharges the proof obligations of Lemma 3.86, Theorem 3.88 and Theorem 3.98, and supplies the sharpened consistency-only form of Theorem 3.87. The one technical input those arguments rest on, the representability theorem Theorem A.17, is standard and long — and is proved here in full rather than cited, in obedience to the same editorial rule that governs every physical derivation in this book, so no link in the chain stands on an unproved result.
The Universal Turing Machine
This appendix proves Theorem 3.95 of Logic, Sets, and Maps: there is a single Turing machine \(U\) which, given the code of any machine \(M\) together with an input \(w\), reproduces exactly what \(M\) does on \(w\). The theorem is the reason a computer is a general-purpose device rather than a machine built for one task, and it is the technical core of [Turing:1937]. It is also what makes Theorem 3.96 possible: a machine that can run any machine can be asked to run itself.
The machine of Definition 3.92 has one tape. Building \(U\) directly in that format is possible but obscures the argument in bookkeeping, so we proceed in two steps: first show that extra tapes buy no extra power (Lemma A.32), then construct \(U\) with three of them.
Coding a machine as a string
Fix a machine \(M=(Q,\Gamma,\sqcup,\Sigma,\delta,q_{0})\). Number its states \(q_{0},q_{1},\ldots,q_{r}\) and its tape symbols \(a_{0}=\sqcup,a_{1},\ldots,a_{s}\), and write \(\mathrm{bin}(i)\) for the binary numeral of \(i\). Each transition \(\delta(q_{i},a_{j})=(q_{k},a_{l},D)\) is coded by the block
and the code of the whole machine is the concatenation of the blocks for all
pairs on which \(\delta\) is defined, separated by the marker ; and
enclosed by \#:
The alphabet of the code is the fixed finite set \(\set{0,1,,,;,\#, L,R}\), the same for every \(M\) however large \(Q\) and \(\Gamma\) are — which is what allows one machine \(U\), with one fixed alphabet, to read them all.
Two properties are immediate from Equation (A.145) and are used
below. First, whether a string is a well-formed code is decidable in the
sense of Definition 3.93: a machine scans left to right and checks
that the punctuation alternates as the format requires and that every block
has five fields. Second, a machine can look up a transition: given
\(\mathrm{bin}(i)\) and \(\mathrm{bin}(j)\) written somewhere on its tape, it
walks the blocks in order, comparing the first two fields of each with the
target pair symbol by symbol, and either finds the matching block or reaches
the closing \#. Both are bounded searches through a finite string.
Extra tapes buy no extra power
A \(k\)-tape Turing machine has \(k\) two-way unbounded tapes with independent heads, and a transition function
reading the \(k\) scanned symbols at once and writing, and moving each head left, right or not at all. Input is on tape 1, output read from tape 1; undefined \(\delta\) means halt. Rests on Definition 3.92.
For every \(k\)-tape machine \(N\) there is a single-tape machine \(N'\), in the sense of Definition 3.92, computing the same partial function. If \(N\) halts on \(w\) in \(n\) steps, \(N'\) halts on \(w\) in \(O(n^{2})\) steps. Rests on Definitions 3.92 and A.31.
Derives Lemma A.32. The construction is to lay the \(k\) tapes out as \(2k\) parallel tracks of one tape. A tape whose cells each carry a \(2k\)-tuple of symbols is nothing more than a tape over the alphabet
which is a finite set — it has \((\abs{\Gamma})^{k}2^{k}\) elements — so \(N'\) is a legitimate single-tape machine over a legitimate finite alphabet. Of each pair of tracks, the first holds the contents of simulated tape \(i\) and the second holds a single \(1\) marking where head \(i\) stands, all other cells of that track carrying \(0\).
\(N'\) keeps the simulated state of \(N\) in its own finite control, which is possible because \(Q\) is finite. One simulated step is performed by two sweeps.
Gathering sweep. \(N'\) walks right from the left end of the used region to its right end, remembering in its control which of the \(k\) head markers it has passed and which symbol stood beneath each. There are \(k\) markers and each carries one of \(\abs{\Gamma}\) symbols, so the information to be remembered is one of \((\abs{\Gamma})^{k}\) possibilities — a finite number, hence storable in the control. At the end of the sweep \(N'\) knows the full \(k\)-tuple of scanned symbols, which is exactly the argument \(\delta\) of Equation (A.146) requires.
Writing sweep. \(N'\) consults \(\delta\) — a fixed finite table built into its control — obtaining the new state, the \(k\) symbols to write and the \(k\) head movements. It then walks back left, and at each marker it passes it writes the new symbol on that track and moves the marker one cell left, right, or not at all as prescribed. It records the new state in its control. If \(\delta\) was undefined on the tuple, \(N'\) halts instead.
The used region grows by at most one cell per simulated step in each direction, so after \(n\) simulated steps it has length \(O(n)\) and each sweep costs \(O(n)\) moves; \(n\) simulated steps therefore cost \(O(n^{2})\) moves of \(N'\). Finally, the input arrives on a single ordinary tape, so \(N'\) begins with a preparatory pass converting it to the track format, and ends with one converting track 1 back; each is \(O(n)\).
By construction the track contents after the \(n\)-th simulated step encode the tapes of \(N\) after its \(n\)-th step, and \(N'\) halts exactly when \(N\) does, so the two compute the same partial function.
∎Construction of $U$
There is a Turing machine \(U\) such that for every machine \(M\) and every input \(w\), \(U\) on input \((\ulcorner M\urcorner,w)\) halts if and only if \(M\) halts on \(w\), and then leaves the same output. Rests on Equation (A.145), Lemma A.32 and Definition 3.92.
Derives Theorem A.33. We build \(U\) with three tapes and appeal to Lemma A.32 at the end.
Layout. Tape 1 holds \(\ulcorner M\urcorner\) and is never altered.
Tape 2 holds the simulated tape of \(M\), with each symbol \(a_{j}\) of \(\Gamma\)
represented by the block \(\mathrm{bin}(j)\) and blocks separated by
,; the simulated head position is marked by a dot on the first
symbol of one block, which costs only a doubling of the tape alphabet.
Tape 3 holds \(\mathrm{bin}(i)\) for the current simulated state \(q_{i}\).
Initialization. \(U\) first verifies that tape 1 carries a well-formed code, as described in Coding a machine as a string, and halts rejecting if not. It then converts the input \(w=w_{1}\cdots w_{p}\) into block form on tape 2, marks the first block, and writes \(\mathrm{bin}(0)\) on tape 3. Tape 2 now encodes the initial configuration of \(M\) on \(w\).
The step loop. \(U\) repeats the following.
-
Read the marked block on tape 2 into the control — it is \(\mathrm{bin}(j)\) for the currently scanned symbol \(a_{j}\) — and read \(\mathrm{bin}(i)\) from tape 3. Neither can be held in the finite control as a whole, since \(\abs{Q}\) and \(\abs{\Gamma}\) are unbounded across machines \(M\); instead \(U\) keeps its head positions and compares symbol by symbol, which is why the lookup is a scan rather than a table jump.
-
Walk tape 1 from the left
\#, comparing the first two fields of each block with \(\mathrm{bin}(i)\) and \(\mathrm{bin}(j)\). Comparison is performed by moving the two heads in step and matching one binary digit at a time. -
If the closing
\#is reached with no match, then \(\delta\) is undefined at \((q_{i},a_{j})\); \(U\) erases its scaffolding, converts tape 2 back to plain symbols, copies it to the output tape and halts. -
If a block matches, read its remaining three fields \(\mathrm{bin}(k)\), \(\mathrm{bin}(l)\), \(d\). Copy \(\mathrm{bin}(k)\) over tape 3. Overwrite the marked block of tape 2 with \(\mathrm{bin}(l)\), shifting the rest of the tape if the two blocks differ in length. Move the mark to the neighbouring block in the direction \(d\), appending a fresh block \(\mathrm{bin}(0)\) for \(\sqcup\) if the mark would leave the written region.
Faithfulness. We claim that for every \(n\), if \(M\) has not halted within \(n\) steps then after \(n\) passes through the loop, tape 2 encodes the tape of \(M\) after \(n\) steps with the mark on the scanned cell, and tape 3 holds the numeral of \(M\)'s state after \(n\) steps. The proof is induction on \(n\) (Axiom 3.34).
For \(n=0\) this is what initialization established. Assume it for \(n\). Then at step 1 of pass \(n+1\), \(U\) reads exactly the state \(q_{i}\) and scanned symbol \(a_{j}\) that \(M\) has after \(n\) steps. Step 2 locates the block Equation (A.144) for that pair, and by Equation (A.145) there is at most one such block, since \(\delta\) is a function. If none exists, \(M\) halts at step \(n+1\) by Definition 3.92, and step 3 makes \(U\) halt with tape 2 decoding to \(M\)'s final tape. If one exists, step 4 writes \(a_{l}\) in place of \(a_{j}\), moves the mark one cell in the direction \(D\) and records \(q_{k}\) — which is precisely the action Definition 3.92 prescribes for step \(n+1\). So the claim holds for \(n+1\).
Consequently \(U\) performs a pass for each step of \(M\) and halts exactly when \(M\) halts, with the same output; and if \(M\) never halts, neither branch of step 3 is ever taken and \(U\) loops forever as well. Finally, Lemma A.32 converts this three-tape \(U\) into a single-tape machine with the same behaviour, which is the machine promised by Theorem 3.95.
∎Each pass of the loop scans a code of fixed length \(\abs{\ulcorner M\urcorner}\) and a simulated tape of length \(O(n)\) after \(n\) steps, so \(U\) simulates \(n\) steps of \(M\) in \(O(n^{2}\abs{\ulcorner M\urcorner})\) of its own, and the reduction to one tape squares that again. The overhead is polynomial, and that is the point: universality costs a slowdown, never a loss of what can be computed at all. This is the precise sense in which Axiom 3.94 is unaffected by the choice of machine model.
The Universal Turing Machine discharges the proof obligation of Theorem 3.95. Its consequence is the one exploited throughout Section 3.9: a program is a string, a string is data, and data can be fed to a program — including to the program that produced it. Theorem 3.96 is that observation turned against itself, and Theorem 3.97 generalizes the damage.
Completeness of the Real Numbers
This appendix proves the statement on which all of real and complex analysis rests (Real Analysis): there exists a complete ordered field, unique up to isomorphism — the real numbers. It is the deepest proof of Parts I and II: every convergence theorem of analysis (monotone convergence, Bolzano–Weierstrass, the Cauchy criterion, the intermediate and extreme value theorems, and through them the fundamental theorem of calculus and Cauchy's integral theorem) descends from the single property proven here. The construction is Dedekind's; the rationals \(\Q\), with their arithmetic and order, are taken as already constructed (Logic, Sets, and Maps and Algebraic Structures).
Statement
There exists an ordered field \(\R\) containing \(\Q\) as an ordered subfield, with the least-upper-bound property: every nonempty subset \(S \subset \R\) that is bounded above has a least upper bound \(\sup S \in \R\). Moreover \(\R\) is Archimedean, \(\Q\) is dense in it, and every Cauchy sequence in \(\R\) converges. Rests on Definition 4.30.
The proof occupies the rest of this section: the construction (The construction: Dedekind cuts), the order and the least-upper-bound property (Order and the least-upper-bound property) — proven before the arithmetic, because it is the property everything else serves — the field operations (The field operations), and the consequences (Consequences).
The construction: Dedekind cuts
The idea is geometric: a real number is identified with the set of all rationals to its left. What must be checked is that the totality of such left rays carries an order, an arithmetic, and no gaps.
A cut is a set \(\alpha \subset \Q\) such that
-
\(\alpha \neq \varnothing\) and \(\alpha \neq \Q\) (nontriviality);
-
if \(p \in \alpha\) and \(q < p\) then \(q \in \alpha\) (downward closure);
-
\(\alpha\) has no largest element: for every \(p \in \alpha\) there is \(r \in \alpha\) with \(p < r\) (openness).
The set of all cuts is denoted \(\R\). Every rational \(q\) embeds as the cut \(q^{*} = \set{p \in \Q \mid p < q}\); properties (1)–(3) are immediate for \(q^{*}\), and \(p < q \iff p^{*} \subsetneq q^{*}\), so the embedding preserves order.
Let \(\alpha\) be a cut, \(p \in \alpha\), \(q \notin \alpha\). Then \(p < q\); and every rational \(r > q\) satisfies \(r \notin \alpha\). Rests on Definition A.36.
Derives Lemma A.37. If \(q \le p\) then downward closure (\(q < p\)) or membership itself (\(q = p\)) would put \(q \in \alpha\), a contradiction; so \(p < q\). If \(r > q\) and \(r \in \alpha\), downward closure would give \(q \in \alpha\) again — contradiction.
∎Order and the least-upper-bound property
For cuts \(\alpha, \beta\): \(\alpha < \beta\) if and only if \(\alpha \subsetneq \beta\); and \(\alpha \le \beta\) iff \(\alpha \subseteq \beta\). Rests on Definition A.36.
For any cuts \(\alpha, \beta\), exactly one of \(\alpha < \beta\), \(\alpha = \beta\), \(\beta < \alpha\) holds. Rests on Definition A.38 and Lemma A.37.
Derives Lemma A.39. At most one holds, since the three are mutually exclusive by construction. Suppose \(\alpha \not< \beta\) and \(\alpha \neq \beta\); we must show \(\beta < \alpha\), i.e. \(\beta \subsetneq \alpha\). Since \(\alpha \not\subseteq \beta\), pick \(p \in \alpha \setminus \beta\). Let \(q \in \beta\) be arbitrary. Applying Lemma A.37 to the cut \(\beta\) (with \(q \in \beta\), \(p \notin \beta\)) gives \(q < p\); downward closure of \(\alpha\) then gives \(q \in \alpha\). Hence \(\beta \subseteq \alpha\), and the inclusion is strict because \(p \notin \beta\).
∎Let \(A \subset \R\) be nonempty and bounded above. Then \(\gamma = \bigcup_{\alpha \in A} \alpha\) is a cut, and \(\gamma = \sup A\). Rests on Definitions A.36 and A.38.
Derives Theorem A.40. \(\gamma\) is a cut. Nontriviality: \(\gamma \supseteq \alpha_0 \neq \varnothing\) for any \(\alpha_0 \in A\); and if \(\beta\) is an upper bound of \(A\) then every \(\alpha \in A\) satisfies \(\alpha \subseteq \beta\), so \(\gamma \subseteq \beta \subsetneq \Q\). Downward closure: if \(p \in \gamma\) then \(p \in \alpha\) for some \(\alpha \in A\), and any \(q < p\) lies in \(\alpha \subseteq \gamma\). Openness: the same \(\alpha\) contains some \(r > p\), and \(r \in \gamma\).
\(\gamma\) is an upper bound. \(\alpha \subseteq \gamma\) for every \(\alpha \in A\) by construction, i.e. \(\alpha \le \gamma\).
\(\gamma\) is the least one. Let \(\delta\) be any upper bound of \(A\). Then \(\alpha \subseteq \delta\) for all \(\alpha \in A\), hence \(\gamma = \bigcup_{\alpha \in A}\alpha \subseteq \delta\), i.e.\ \(\gamma \le \delta\).
∎Note how little the proof costs once the construction is right: the supremum is literally the union. The gaps of \(\Q\) — the cut \(\set{p \in \Q \mid p \le 0 \text{ or } p^2 < 2}\) has no rational supremum — are filled by fiat, because the union of cuts is again a cut.
The field operations
\(\alpha + \beta = \set{p + q \mid p \in \alpha,\ q \in \beta}\), with neutral element \(0^{*}\) and \(-\alpha = \set{p \in \Q \mid \exists\, r > 0:\ -p - r \notin \alpha}\). Rests on Definition A.36.
\(\alpha + \beta\) and \(-\alpha\) are cuts; addition is associative and commutative with \(\alpha + 0^{*} = \alpha\) and \(\alpha + (-\alpha) = 0^{*}\); and \(\alpha \le \beta\) implies \(\alpha + \gamma \le \beta + \gamma\). Rests on Definition A.41, Lemma A.37 and Definition A.38.
Derives Lemma A.42. \(\alpha+\beta\) is a cut. It is nonempty; it omits \(p' + q'\) whenever \(p' \notin \alpha\), \(q' \notin \beta\) (by Lemma A.37 any element \(p + q < p' + q'\)); it is downward closed, since \(s < p + q\) writes as \(s = (s - q) + q\) with \(s - q < p\), so \(s - q \in \alpha\); and it has no largest element, since \(p\) may be enlarged inside \(\alpha\). Associativity and commutativity are inherited from \(\Q\) elementwise, and \(\alpha + 0^{*} = \alpha\) follows from downward closure and openness: \(p + n < p\) for \(n < 0\) gives \(\alpha + 0^{*} \subseteq \alpha\), while any \(p \in \alpha\) has \(r \in \alpha\), \(r > p\), whence \(p = r + (p - r) \in \alpha + 0^{*}\).
Inverse. \(-\alpha\) is a cut (checks as above). For \(\alpha + (-\alpha) \subseteq 0^{*}\): if \(p \in \alpha\) and \(q \in -\alpha\) with witness \(r > 0\) (so \(-q - r \notin \alpha\)), then \(p < -q - r\) by Lemma A.37, so \(p + q < -r < 0\). Conversely let \(s < 0\); put \(t = -s/2 > 0\). By the Archimedean property of \(\Q\) there is an integer \(n\) with \(n t \in \alpha\) and \((n+1)t \notin \alpha\) (walk upward in steps of \(t\) from inside \(\alpha\); the walk exits because \(\alpha \neq \Q\)). Take \(p = n t \in \alpha\) and \(q = s - n t\). Then \(-q - t = -s + n t - t = n t + 2t - t = (n+1)t \notin \alpha\), so \(q \in -\alpha\) with witness \(t\), and \(p + q = s\). Monotonicity of \(+\) is elementwise.
∎For \(\alpha, \beta > 0^{*}\):
extended to all signs by \(\alpha\, 0^{*} = 0^{*}\) and the sign rules \(\alpha\beta = -\left((-\alpha)\beta\right)\) etc., with unit \(1^{*}\) and, for \(\alpha > 0^{*}\), the reciprocal \(\alpha^{-1} = \set{p \in \Q \mid p \le 0} \cup \set{p > 0 \mid \exists\, r > 0:\ (p^{-1} - r)^{-1} \notin \alpha}\). Rests on Definitions A.36 and A.41.
With Definitions A.41 and A.43, \(\R\) satisfies the ordered-field axioms (Algebraic Structures), and the embedding \(q \mapsto q^{*}\) is a field homomorphism. Rests on Definition A.41, Definition A.43, Lemma A.42 and Definition 4.30.
Derives Lemma A.44. The verifications are of exactly the same kind as in Lemma A.42 — elementwise inheritance from \(\Q\) on the positive cone, then bookkeeping of the sign cases for associativity, commutativity, distributivity, and \(\alpha\alpha^{-1} = 1^{*}\) (the reciprocal argument mirrors the additive-inverse argument with the multiplicative Archimedean walk \(p, 2p, 4p, \dots\)). They are long but routine, and no idea beyond those already displayed enters; we omit the repetition. Compatibility of order and product — \(\alpha, \beta > 0^{*} \implies \alpha\beta > 0^{*}\) — is immediate from Equation (A.148).
∎Consequences
For \(\alpha, \beta \in \R\) with \(\alpha > 0^{*}\) there is \(n \in \N\) with \(n\alpha > \beta\); and between any two reals lies a rational. Rests on Theorem A.40 and Definition A.36.
Derives Corollary A.45. If \(n\alpha \le \beta\) for all \(n\), the set \(A = \set{n\alpha \mid n \in \N}\) is bounded above, so \(\gamma = \sup A\) exists (Theorem A.40). Then \(\gamma - \alpha < \gamma\) is not an upper bound: some \(n\alpha > \gamma - \alpha\), whence \((n+1)\alpha > \gamma\) — contradicting that \(\gamma\) bounds \(A\). Density: if \(\alpha < \beta\), pick \(q \in \beta \setminus \alpha\); openness of \(\beta\) gives \(r \in \beta\), \(r > q\), and then \(\alpha \le q^{*} < r^{*} \le \beta\), so the embedded rationals \(q^{*}, r^{*}\) separate \(\alpha\) from \(\beta\).
∎A nondecreasing sequence \((a_n)\) in \(\R\) that is bounded above converges to \(\sup_n a_n\). Rests on Theorem A.40.
Derives Corollary A.46. Let \(s = \sup_n a_n\), which exists by Theorem A.40. Given \(\varepsilon > 0\), \(s - \varepsilon\) is not an upper bound, so \(a_N > s - \varepsilon\) for some \(N\); monotonicity gives \(s - \varepsilon < a_n \le s\) for all \(n \ge N\), i.e.\ \(\abs{a_n - s} < \varepsilon\).
∎Every Cauchy sequence \((a_n)\) in \(\R\) converges in \(\R\). Rests on Corollary A.46 and Theorem A.40.
Derives Corollary A.47. A Cauchy sequence is bounded (take \(\varepsilon = 1\): beyond some \(N\) all terms lie within \(1\) of \(a_N\), and only finitely many terms remain). Define \(b_n = \inf_{k \ge n} a_k\), which exists because \(\set{a_k \mid k \ge n}\) is bounded below and Theorem A.40 applied to the negatives provides infima. The sequence \((b_n)\) is nondecreasing and bounded above, so by Corollary A.46 it converges to \(L = \sup_n b_n = \liminf_n a_n\). Given \(\varepsilon > 0\) choose \(N\) with \(\abs{a_m - a_n} < \varepsilon/2\) for \(m, n \ge N\). Then for \(n \ge N\) every \(a_k\) with \(k \ge n\) lies in \((a_N - \varepsilon/2,\ a_N + \varepsilon/2)\), so \(b_n \in [a_N - \varepsilon/2,\ a_N + \varepsilon/2]\) and hence \(\abs{L - a_N} \le \varepsilon/2\); combining, \(\abs{a_n - L} \le \abs{a_n - a_N} + \abs{a_N - L} < \varepsilon\) for \(n \ge N\).
∎Proof of Theorem A.35. Derives Theorem A.35. Assembling the section. The set of cuts of Definition A.36, ordered by inclusion, is a totally ordered set containing an order-embedded copy of \(\Q\) (Lemmas A.37 and A.39); it has the least-upper-bound property by Theorem A.40; addition and multiplication of Definitions A.41 and A.43 make it an ordered field by Lemmas A.42 and A.44. The three remaining assertions are Corollary A.45 (Archimedean, and \(\Q\) dense) and Corollary A.47 (Cauchy sequences converge). Existence is therefore established by construction, which is what the theorem claims.
∎Any two complete ordered fields are isomorphic as ordered fields: the isomorphism sends each element to the supremum of the embedded rationals below it, and Corollary A.45 (which followed from completeness alone) makes the map well defined, bijective, and operation-preserving. Thus the real numbers are characterized by the axioms, not by the construction; Cantor's alternative construction by Cauchy sequences of rationals yields the same field. This is why Real Analysis may take Theorem A.35 as its single axiom for \(\R\) and never mention cuts again.
Equivalence of the Einstein–Hilbert Action in Metric and Vielbein Variables
This appendix proves the equivalence asserted in Remark 13.159 (Differentiable Manifolds, Tensors, and Curvature, Equation (13.317)) and used to set up the action of Geometric Formulation of Gravity: in \(D=p+q\) dimensions, the Einstein–Hilbert Lagrangian density built from the vielbein and the Cartan curvature two-form coincides, up to a fixed combinatorial constant, with \(\sqrt{\abs{g}}\,R\). All conventions (frame indices \(a,b,\ldots\), the structure equations, the definition \(R^{a}{}_{b}=\frac12 R^{a}{}_{b\mu\nu}\,\dd x^{\mu}\wedge\dd x^{\nu}\)) are those of Sections 13.1, 13.8 and 13.12.
Statement
Let \((M,g)\) be a \(D\)-dimensional pseudo-Riemannian manifold with vielbein \(e^{a}\), torsion-free spin connection \(\omega^{a}{}_{b}\), and curvature two-form \(R^{a}{}_{b}\). Then, as \(D\)-forms,
with \(\epsilon_{a_{1}\ldots a_{D}}\) the numerical (frame-index) Levi-Civita symbol, \(\epsilon_{1\ldots D}=+1\), and \(R\) the Ricci scalar of Definition 13.153. Equivalently, defining \(S_{\mathrm{EH}}=\dfrac{1}{2\kappa}\displaystyle\int\dd^{D}x\, \sqrt{\abs{g}}\,R\) (Geometric Formulation of Gravity),
Rests on Lemma A.50, Definition 13.153, Equation (13.264), Proposition 13.123 and Equation (13.310).
The proof needs one combinatorial lemma about the Levi-Civita symbol, proven in A contraction lemma for the Levi-Civita symbol, and a component computation carried out at a conveniently chosen point of \(M\), in Proof of the theorem.
A contraction lemma for the Levi-Civita symbol
Let \(\epsilon_{i_{1}\ldots i_{D}}\) be the totally antisymmetric numerical symbol in \(D\) indices ranging over \(1,\ldots,D\), normalized by \(\epsilon_{1\ldots D}=+1\), and write \(\epsilon^{i_{1}\ldots i_{D}}\) for the same numerical values with the indices held upstairs. Then
summed over \(k_{3},\ldots,k_{D}=1,\ldots,D\).
Derives Lemma A.50. Both sides of Equation (A.151) are, as functions of the four free indices \(i,j,l,m\), antisymmetric under \(i\leftrightarrow j\) and under \(l\leftrightarrow m\) separately: for the right side this is immediate; for the left side, swapping \(i\leftrightarrow j\) swaps the first two slots of \(\epsilon^{ij k_{3}\ldots}\), which flips its sign by total antisymmetry, and the swap \(l\leftrightarrow m\) acts the same way on the second factor. Hence both sides vanish whenever \(i=j\) or \(l=m\), and, by the double antisymmetry, it suffices to verify Equation (A.151) for the case \(i=l\), \(j=m\), \(i\neq j\) (every other nonvanishing assignment of \(\set{i,j}=\set{l,m}\) then follows by the antisymmetries just established, and every assignment with \(\set{i,j}\neq\set{l,m}\) gives \(0=0\): the left side vanishes because no choice of \(k_{3},\ldots,k_{D}\) can make both \(\epsilon\) factors nonzero when \(\set{i,j}\ne\set{l,m}\), since each factor forces its own two free indices to be distinct from the shared \(k_{3},\ldots,k_{D}\) and from each other, and a nonzero product further forces the sets \(\set{i,j}\) and \(\set{l,m}\) to be identical, as both equal the complement of \(\set{k_{3},\ldots,k_{D}}\) in \(\set{1,\ldots,D}\)).
At \(i=l\), \(j=m\), \(i\ne j\): the right side of Equation (A.151) is \((D-2)!\,(1-0)=(D-2)!\). For the left side, \(\epsilon^{ijk_{3}\ldots k_{D}}\epsilon_{ijk_{3}\ldots k_{D}}\) (no sum yet on \(i,j\)) is nonzero only when \((k_{3},\ldots,k_{D})\) is a permutation of the \(D-2\) elements of \(\set{1,\ldots,D}\setminus\set{i,j}\), and each such term contributes \(\left(\epsilon^{ijk_{3}\ldots k_{D}}\right)^{2}=1\); there are exactly \((D-2)!\) such permutations, so the sum equals \((D-2)!\). The two sides agree.
∎Proof of the theorem
Proof of Theorem A.49. Derives Theorem A.49. Fix a point \(P\in M\). Expand the curvature two-form and the vielbeins in the coordinate basis,
so that, writing \(\Omega\) for the left side of Equation (A.149),
Choice of point-adapted frame.
Both sides of Equation (A.149) are built — with no further differentiation — from the algebraic data \(\left(g_{\mu\nu}(P),\,R^{a}{}_{b\mu\nu}(P),\,e^{a}{}_{\mu}(P)\right)\) and the numerical symbols \(\epsilon,\eta\); consequently, once verified in one admissible frame and coordinate system at \(P\), the identity holds in every frame and coordinate system at \(P\), by the tensorial (respectively, local-Lorentz-covariant) transformation laws of \(g\), \(R^{a}{}_{b\mu\nu}\) and \(e^{a}{}_{\mu}\) established in Section 13.5.2, Section 13.8 and Theorem 13.156 — both sides of Equation (A.149) transform as the components of a \(D\)-form, i.e. by the same Jacobian determinant, so an equality of components in one frame is an equality of components in every frame. It therefore suffices to prove Equation (A.149) at \(P\) in a coordinate system and frame chosen so that
Such a choice exists: pick any coordinates first, so that \(g_{\mu\nu}(P)\) is some fixed nondegenerate symmetric matrix of signature \((p,q)\); a constant linear change of coordinates diagonalizes it to \(\eta_{\mu\nu}\) at \(P\) (Sylvester's law of inertia, Algebraic Structures), after which the vielbein \(e^{a}{}_{\mu}=\delta^{a}_{\mu}\) satisfies \(\eta_{ab}\,\delta^{a}_{\mu}\delta^{b}_{\nu}=\eta_{\mu\nu}=g_{\mu\nu}(P)\) (Equation (13.264)), and every other vielbein compatible with \(g_{\mu\nu}(P)\) differs from this one by a local Lorentz transformation (Proposition 13.123), so this choice is always reachable.
Component computation at $P$.
Under Equation (A.152), \(e^{a_{i}}{}_{\lambda_{i}}(P) =\delta^{a_{i}}_{\lambda_{i}}\) collapses each frame sum \(a_{3},\ldots,a_{D}\) against the epsilon symbol into the corresponding coordinate index, and \(R^{a_{1}a_{2}}{}_{\mu\nu}(P)\) becomes numerically the doubly-raised Riemann tensor \(g^{a_{1}\rho}(P)g^{a_{2}\sigma}(P)R_{\rho\sigma\mu\nu}(P)\), so that at \(P\),
where the \(\rho,\sigma\) sum has simply been relabelled from \(a_{1},a_{2}\). Reordering the wedge of coordinate differentials (Section 13.6.3), \(\dd x^{\mu}\wedge\dd x^{\nu}\wedge\dd x^{\lambda_{3}} \wedge\cdots\wedge\dd x^{\lambda_{D}} =\varepsilon^{\mu\nu\lambda_{3}\ldots\lambda_{D}}\, \dd x^{1}\wedge\cdots\wedge\dd x^{D}\), so
By Lemma A.50 (with \(D-2\) contracted indices \(\lambda_{3},\ldots,\lambda_{D}\)), the bracket equals \((D-2)!\left(\delta^{\mu}_{\rho}\delta^{\nu}_{\sigma} -\delta^{\mu}_{\sigma}\delta^{\nu}_{\rho}\right)\), giving
Renaming the summed indices \(\rho\leftrightarrow\sigma\) in the second term and using the antisymmetry of the curvature two-form in its upper pair, \(R^{\rho\sigma}{}_{\mu\nu}=-R^{\sigma\rho}{}_{\mu\nu}\) (inherited from \(\omega^{ab}=-\omega^{ba}\), Equation (13.310)),
so the bracket is \(2\,R^{\rho\sigma}{}_{\rho\sigma}(P)\) and
Identification with the Ricci scalar.
It remains to recognize \(R^{\rho\sigma}{}_{\rho\sigma}(P)\). Lowering both upper indices with \(g^{-1}(P)\),
Contract the inner index pair first: by Definition 13.153, \(R_{\mu\nu}=R^{\lambda}{}_{\mu\lambda\nu}=g^{\lambda\kappa} R_{\kappa\mu\lambda\nu}\), i.e. contracting the first index of the (fully lowered) Riemann tensor with the third, via \(g^{-1}\), produces the Ricci tensor in its second and fourth slots. Applying this with \(\kappa\to\rho,\,\mu\to\beta,\,\lambda\to\alpha,\,\nu\to\sigma\) gives \(g^{\rho\alpha}R_{\alpha\beta\rho\sigma}=R_{\beta\sigma}\), so
the last equality being Equation (13.307). Hence \(R^{\rho\sigma}{}_{\rho\sigma}(P) = R(P)\), and Equation (A.153) reads
Finally, under Equation (A.152), \(g(P)=\det\eta_{\mu\nu} =(-1)^{q}\), so \(\sqrt{\abs{g(P)}}=1\) and \(\dd x^{1}\wedge\cdots\wedge\dd x^{D}=\sqrt{\abs{g(P)}}\,\dd^{D}x\) at \(P\). Substituting,
which is Equation (A.149) at \(P\). Since \(P\in M\) was arbitrary and, by the covariance argument above, the identity transported to any frame at \(P\), Equation (A.149) holds identically on \(M\). Dividing by \(2\kappa\,(D-2)!\) and integrating gives Equation (A.150).
∎For \(D=2\) there are no vielbein factors left in Equation (A.149) (\((D-2)!=0!=1\), and the product \(e^{a_{3}}\wedge\cdots\wedge e^{a_{D}}\) is empty), so the claim reduces to \(\epsilon_{ab}R^{ab}=R\sqrt{\abs{g}}\,\dd^{2}x\). Using \(R^{12}=-R^{21}\) and \(\epsilon_{12}=-\epsilon_{21}=1\), the left side is \(2R^{1}{}_{2}\), and \(R^{1}{}_{2}=K\,e^{1}\wedge e^{2}\) with \(K\) the Gaussian curvature of Section 13.3.3; together with \(R=2K\) in \(D=2\) (Definition 13.153) this gives \(2R^{1}{}_{2}=2K\,e^{1}\wedge e^{2}=R\sqrt{\abs{g}}\,\dd^{2}x\), matching the theorem with the correct constant \((D-2)!=1\) and independently confirming Lemma A.50 in its simplest nontrivial case — see Section 13.3.3 for the surface-theory computation this reproduces.
Appendix A.8 closes the derivation anticipated in Differentiable Manifolds, Tensors, and Curvature; the action Equation (A.150) is put to use, and its variation carried out, in Geometric Formulation of Gravity and The Einstein Field Equations.
Wigner's Theorem on Quantum Symmetries
This appendix proves Theorem 94.3 (Chapter 94): every bijection of the unit rays of a Hilbert space that preserves transition probabilities is implemented by an operator that is either unitary or antiunitary, unique up to a phase. The theorem is stated in Wigner's 1931 monograph [Wigner:1931]; the proof below is Bargmann's [Bargmann:1964], streamlined but complete. Nothing in it assumes the Hilbert space separable: the basis index \(k\) may run over an arbitrary set.
Throughout, \(\mathcal{H}\) is a complex Hilbert space with \(\dim\mathcal{H} \ge 2\), the unit ray of a normalized vector is \([\psi] = \set{\ee^{\ii\alpha}\psi \mid \alpha \in \R}\), and \(T\) is a bijection of the set of unit rays with
(representatives may be chosen arbitrarily: the modulus does not depend on the choice).
Orthonormal bases map to orthonormal bases
Let \(\set{e_k}\) be an orthonormal basis of \(\mathcal{H}\) and choose any unit representatives \(\psi_k \in T[e_k]\). Then \(\set{\psi_k}\) is an orthonormal basis. Rests on Equation (A.154).
Derives Lemma A.52. Orthonormality is Equation (A.154): \(\abs{\braket{\psi_j}{\psi_k}} = \abs{\braket{e_j}{e_k}} = \delta_{jk}\). Completeness: let the unit vector \(\chi\) be orthogonal to every \(\psi_k\). Since \(T\) is surjective, \([\chi] = T[\varphi]\) for some unit \(\varphi\), and then \(\abs{\braket{\varphi}{e_k}} = \abs{\braket{\chi}{\psi_k}} = 0\) for every \(k\), contradicting completeness of \(\set{e_k}\) and \(\norm{\varphi} = 1\).
∎Phase alignment along a basis
Fix the orthonormal basis \(\set{e_k}\), single out one index, written \(k = 1\), and abbreviate the normalized two- and three-component test vectors
The representatives \(\psi_k\) of Lemma A.52 can be rephased, \(f_1 = \psi_1\) and \(f_k = \ee^{\ii\gamma_k}\psi_k\), so that the first relation below holds for every \(k \ge 2\) and the remaining two for all \(j \neq k\) with \(j, k \ge 2\):
Rests on Lemma A.52, Equation (A.154) and Equation (A.155).
Derives Lemma A.53. Expand a unit representative of \(T[u_k]\) in the basis \(\set{\psi_j}\): by Equation (A.154) tested against each \([e_j]\), its coefficients have moduli \(1/\sqrt2\) at \(j \in \set{1,k}\) and \(0\) elsewhere, so the representative is \((\ee^{\ii\alpha}\psi_1 + \ee^{\ii\beta}\psi_k)/\sqrt2\); multiplying by \(\ee^{-\ii\alpha}\) (same ray) makes it \((\psi_1 + \ee^{\ii\gamma_k}\psi_k)/\sqrt2\) with a uniquely determined \(\gamma_k\). Define \(f_k = \ee^{\ii\gamma_k}\psi_k\); the first relation of Equation (A.156) holds by construction, and \(f_k \in T[e_k]\) still. For the second: a representative of \(T[u_{jk}]\) likewise reads \((f_1 + c_j f_j + c_k f_k)/\sqrt3\) with \(\abs{c_j} = \abs{c_k} = 1\) after normalizing the \(f_1\)-coefficient; testing against \(T[u_j]\) and \(T[u_k]\) with Equation (A.154), \(\abs{1 + c_j} = \abs{1 + 1} = 2\) forces \(c_j = 1\), and similarly \(c_k = 1\). For the third: a representative of \(T[(e_j+e_k)/\sqrt2]\) is \((\ee^{\ii\mu}f_j + \ee^{\ii\nu}f_k)/\sqrt2\); testing against \(T[u_{jk}]\), \(\abs{\ee^{-\ii\mu} + \ee^{-\ii\nu}} = 2\) forces \(\mu = \nu\), and the common phase is dropped.
∎The induced scalar map
For every \(k \ge 2\) and \(\alpha \in \C\) there is a unique \(\chi_k(\alpha) \in \C\) with
and \(\chi_k(\alpha) \in \set{\alpha, \bar\alpha}\). Moreover a single alternative holds for all \(\alpha\) and all \(k\): either \(\chi_k(\alpha) = \alpha\) for every \(k\) and \(\alpha\), or \(\chi_k(\alpha) = \bar\alpha\) for every \(k\) and \(\alpha\). Write \(\chi\) for the common map. Rests on Lemma A.53 and Equation (A.154).
Derives Lemma A.54. As in Lemma A.53, a representative of the left side has components only along \(f_1, f_k\) with moduli \(1/\sqrt{1+\abs{\alpha}^2}\) and \(\abs{\alpha}/\sqrt{1+\abs{\alpha}^2}\); normalizing the \(f_1\)-component to be real positive determines the representative and hence \(\chi_k(\alpha)\) uniquely, with \(\abs{\chi_k(\alpha)} = \abs{\alpha}\). Testing against \(T[u_k]\) via Equation (A.154) gives \(\abs{1 + \chi_k(\alpha)} = \abs{1 + \alpha}\); together with equal moduli this forces \(\Re\chi_k(\alpha) = \Re\alpha\), hence \(\chi_k(\alpha) \in \set{\alpha,\bar\alpha}\).
No mixing within one \(k\): suppose \(\chi_k(\alpha) = \alpha\) and \(\chi_k(\beta) = \bar\beta\) with \(\Im\alpha \ne 0 \ne \Im\beta\). Test the rays Equation (A.157) for \(\alpha\) and \(\beta\) against each other: Equation (A.154) demands \(\abs{1 + \bar\alpha\beta} = \abs{1 + \bar\alpha\bar\beta}\), i.e.\ \(\Re\left(\bar\alpha\beta\right) = \Re\left(\bar\alpha\bar\beta\right)\), i.e. \(\Im\alpha\,\Im\beta = 0\) — a contradiction. Real \(\alpha\) satisfy both alternatives, so \(\chi_k\) is the identity on all of \(\C\) or conjugation on all of \(\C\).
No mixing across \(k\): let \(j \ne k\) and suppose \(\chi_j = \mathrm{id}\) while \(\chi_k\) is conjugation, and pick \(\alpha = \beta = \ii\). A representative of \(T[(e_1 + \alpha e_j + \beta e_k)/\sqrt3]\), normalized so that its \(f_1\)-coefficient is real positive, reads \((f_1 + c_j f_j + c_k f_k)/\sqrt3\) with \(\abs{c_j} = \abs{\alpha}\) and \(\abs{c_k} = \abs{\beta}\) (tests against each \([e_m]\)). Testing against the family Equation (A.157) with target \(j\), at the two values \(\gamma = 1\) and \(\gamma = \ii\), gives \(\abs{1 + \overline{\chi_j(\gamma)}\,c_j} = \abs{1 + \bar\gamma\,\alpha}\), which pins the real and the imaginary part of \(c_j\) and forces \(c_j = \chi_j(\alpha)\) — the within-target dichotomy just proved fixes \(\chi_j\) globally, so \(\chi_j(\gamma)\) is known at both test values — and likewise \(c_k = \chi_k(\beta)\). Testing the resulting representative against \(T[(e_j + e_k)/\sqrt2]\) — the third relation of Equation (A.156) — yields \(\abs{\chi_j(\alpha) + \chi_k(\beta)} = \abs{\alpha + \beta} = 2\), whereas \(\chi_j(\ii) + \chi_k(\ii) = \ii + \overline{\ii\,} = 0\): a contradiction. (When \(\dim\mathcal{H} = 2\) there is a single \(k\) and this step is vacuous.)
∎Proof of the theorem
Define, on finite linear combinations and extended by continuity,
If \(\chi = \mathrm{id}\), \(U\) is linear and maps the orthonormal basis \(\set{e_k}\) to the orthonormal basis \(\set{f_k}\) (Lemma A.52): it is unitary. If \(\chi\) is conjugation, \(U\) is antilinear with \(\braket{U\varphi}{U\psi} = \overline{\braket{\varphi}{\psi}}\): antiunitary.
Proof that \(U\) implements \(T\). Derives Theorem 94.3. Let \(\psi = \sum_k a_k e_k\) be an arbitrary unit vector and \(\varphi = \sum_k b_k f_k\) a unit representative of \(T[\psi]\). Testing against each \([e_k]\) gives \(\abs{b_k} = \abs{a_k}\). If only one \(a_n \neq 0\), then \([\psi] = [e_n]\) and \(T[\psi] = [f_n] = [U\psi]\) directly. Otherwise fix an index \(n\) with \(a_n \neq 0\), taking \(n = 1\) whenever \(a_1 \neq 0\), and rephase \(\varphi\) so that \(b_n = \chi(a_n)\) (possible since \(\abs{b_n} = \abs{a_n}\)). Components with \(a_k = 0\) are settled at once: \(\abs{b_k} = \abs{a_k} = 0\) gives \(b_k = 0 = \chi(a_k)\); in particular, if \(n \neq 1\) then \(a_1 = 0\) by the choice of \(n\), and only targets \(k \ge 2\), \(k \neq n\), remain to be pinned. For every such target and every \(\gamma \in \C\) we claim
— with the same global \(\chi\). For \(n = 1\) this is Lemma A.54 itself. For \(n \ge 2\), the first two paragraphs of the proof of Lemma A.54, run with anchor \(n\) in place of \(1\) — the needed alignment inputs, the pair \((f_n + f_k)/\sqrt2\) and the triples based at \(n\), are supplied by Equation (A.156) — yield Equation (A.159) with some uniform alternative \(\chi^{(n)}_{k} \in \set{\mathrm{id},\ \text{conjugation}}\) in place of \(\chi\). To identify the alternative, test the ray \([(e_n + \ii e_k)/\sqrt2]\) against \(T[(e_1 + e_n + \ii e_k)/\sqrt3] = [(f_1 + f_n + \chi(\ii)\,f_k)/\sqrt3]\) — the latter expansion established exactly as in the no-mixing step of Lemma A.54 (anchor \(1\), targets \(n\) and \(k\)): Equation (A.154) demands
and since \(\chi^{(n)}_{k}(\ii) = \pm\ii\) and \(\chi(\ii) = \pm\ii\), the left side equals \(2\) when \(\chi^{(n)}_{k}(\ii) = \chi(\ii)\) and \(0\) otherwise: hence \(\chi^{(n)}_{k}(\ii) = \chi(\ii)\), and the uniform dichotomy forces \(\chi^{(n)}_{k} = \chi\), proving the claim. Applying Equation (A.154) to the ray Equation (A.159) against \([\psi]\):
Square both sides and cancel the equal terms \(\abs{a_n}^2 + \abs{\gamma}^2\abs{a_k}^2\):
For \(\chi = \mathrm{id}\), running \(\gamma\) over \(1\) and \(\ii\) recovers both real and imaginary parts of \(\bar a_n a_k\) and \(\bar a_n b_k\) and forces \(b_k = a_k = \chi(a_k)\); for \(\chi\) conjugation, the same two tests force \(b_k = \bar a_k = \chi(a_k)\). Hence \(\varphi = \sum_k \chi(a_k) f_k = U\psi\), so \(T[\psi] = [U\psi]\).
∎Uniqueness up to phase, and exclusivity. Derives Theorem 94.3. Let \(U\) and \(V\) both implement \(T\) and set \(M = U^{-1}V\), an operator preserving every unit ray: \([M\psi] = [\psi]\). On the basis, \(M e_k = \lambda_k e_k\) with \(\abs{\lambda_k} = 1\).
If \(U\) and \(V\) have the same linearity character, \(M\) is linear: \(M(e_1 + e_k) = \lambda_1 e_1 + \lambda_k e_k\) must be proportional to \(e_1 + e_k\), so \(\lambda_k = \lambda_1\) for all \(k\) and \(M = \lambda_1\identity\): \(V = \lambda_1 U\).
If they had opposite character, \(M\) would be antilinear. Then \(M(e_1 + e_k) = \lambda_1 e_1 + \lambda_k e_k \propto e_1 + e_k\) gives \(\lambda_k = \lambda_1\), while \(M(e_1 + \ii e_k) = \lambda_1 e_1 - \ii\lambda_k e_k \propto e_1 + \ii e_k\) gives \(\lambda_k = -\lambda_1\) — impossible. Hence for \(\dim\mathcal{H}\ge2\) a symmetry transformation is implemented either by unitaries or by antiunitaries, never both, and within its class the implementation is unique up to an overall phase.
∎If \(\dim\mathcal{H} = 1\) there is a single unit ray, \(T\) is the identity, and it is implemented both by \(\identity\) (unitary) and by complex conjugation in any basis (antiunitary): the dichotomy of Theorem 94.3 genuinely requires \(\dim\mathcal{H} \ge 2\).
Appendix A.9 discharges the proof obligation of Theorem 94.3; it is put to work in Corollary 94.4 — symmetries of a connected group act by honest unitaries — and through it in the whole classification of Particles as Poincaré Representations. The antiunitary branch is not an idle alternative: time reversal must take it, as Part XI will use.
The Classification of Kinematical Algebras
This appendix proves Theorem 14.90 (Chapter 14): under isotropy, the two discrete automorphisms, and noncompactness of the boosts, there are exactly eleven kinematical Lie algebras, organized as displayed in Table 14.1. The result is Bacry and Lévy-Leblond's [Bacry:1968]; the derivation below is self-contained, and it is written out here because what it yields is not only a count but the parametrization that makes the whole family intelligible — three constants, two Jacobi constraints, and a lattice of limits that turns out to be exactly the contraction lattice.
The generators and the hypotheses
A kinematical algebra \(\mathfrak{k}\) is a real ten-dimensional Lie algebra spanned by
with \(i=1,2,3\), subject to the three hypotheses of Theorem 14.90. Written out, they are the following.
(H1) Isotropy.
The \(J_{i}\) close into \(\mathfrak{so}(3)\), under which \(K_{i}\) and \(P_{i}\) are vectors and \(H\) is a scalar:
(H2) Parity and time reversal are automorphisms.
The linear maps
are automorphisms of \(\mathfrak{k}\). These are the algebra-level statements of \(\vect{x}\mapsto-\vect{x}\) and \(t\mapsto-t\): a boost velocity reverses under either, a position reverses only under the first, and the generator of translation in time reverses only under the second.
(H3) Noncompact boosts.
Each \(K_{i}\) generates a one-parameter subgroup isomorphic to \(\R\) and not to a circle. Only the sign of one structure constant will turn out to be at stake, and (H3) is used exactly twice, in the last two cases of Step four: the enumeration.
Step one: the hypotheses reduce the algebra to five constants
Hypotheses (H1) and (H2) force the brackets not already fixed by Equation (A.161) to take the form
for five real constants \(\alpha,\beta,\gamma,\mu,\nu\). Rests on Equations (A.160), (A.161), (A.162) and (A.163).
Derives Lemma A.56. Each unknown bracket is an \(\mathfrak{so}(3)\)-covariant expression in the generators, so it may be expanded on the covariant objects available: the scalar \(H\), the vectors \(J_{k},K_{k},P_{k}\), and nothing else. Under (H2) each generator carries a definite pair of signs, and a bracket carries the product of the signs of its two entries:
| $J_{i}$ | $K_{i}$ | $P_{i}$ | $H$ | |
|---|---|---|---|---|
| parity $\Pi$ | $+$ | $-$ | $-$ | $+$ |
| time reversal $\Theta$ | $+$ | $-$ | $+$ | $-$ |
The four cases are then immediate.
The bracket \(\comm{K_{i}}{K_{j}}\). It is antisymmetric in \(ij\), hence a vector, hence of the form \(\epsilon_{ijk}V_{k}\) with \(V\in\set{J,K,P}\). Its \(\Pi\)-sign is \((-)(-)=+\), which excludes \(K\) and \(P\); only \(J\) survives. Its \(\Theta\)-sign is \((-)(-)=+\), and \(J\) is \(\Theta\)-even, so no further restriction arises. This gives the first of Equation (A.164). The bracket \(\comm{P_{i}}{P_{j}}\) is identical in structure, with \(\Pi\)-sign \(+\) and \(\Theta\)-sign \((+)(+)=+\), giving the second.
The bracket \(\comm{K_{i}}{P_{j}}\). As a rank-two tensor it decomposes into a trace, an antisymmetric part and a symmetric traceless part, of spins \(0\), \(1\) and \(2\). There is no spin-\(2\) generator in Equation (A.160), so the symmetric traceless part vanishes identically. The trace part is \(\gamma\,\delta_{ij}H\) and the antisymmetric part is \(\delta\,\epsilon_{ijk}V_{k}\). The \(\Pi\)-sign is \((-)(-)=+\), which leaves \(H\) and \(J\); the \(\Theta\)-sign is \((-)(+)=-\), and \(H\) is \(\Theta\)-odd while \(J\) is \(\Theta\)-even, so the \(J\) term is forbidden and \(\delta=0\). This is the first of Equation (A.165), and it is worth pausing on: the reason a kinematical algebra cannot have \(\comm{K_{i}}{P_{j}}\propto\epsilon_{ijk}J_{k}\) is time reversal alone.
The brackets with \(H\). \(\comm{K_{i}}{H}\) is a vector of \(\Pi\)-sign \((-)(+)=-\), excluding \(J\), and of \(\Theta\)-sign \((-)(-)=+\), excluding \(K\) (which is \(\Theta\)-odd); only \(P\) survives. \(\comm{P_{i}}{H}\) is a vector of \(\Pi\)-sign \(-\), again excluding \(J\), and of \(\Theta\)-sign \((+)(-)=-\), excluding \(P\); only \(K\) survives. These are the second of Equation (A.165) and Equation (A.166).
∎Step two: the Jacobi identity eliminates two of the five
The brackets of Lemma A.56 satisfy the Jacobi identity if and only if
A kinematical algebra is therefore determined by the three constants \((\gamma,\mu,\nu)\) alone. Rests on Lemma A.56 and Equation (A.161).
Derives Lemma A.57. Every triple containing a \(J\) is satisfied automatically, because Equations (A.164), (A.165) and (A.166) were constructed to be \(\mathfrak{so}(3)\)-covariant. Of the remaining triples, those built from three copies of one vector generator are also automatic: for \((K_{i},K_{j},K_{l})\) the cyclic sum is
by three applications of \(\epsilon_{kab}\epsilon_{kcd}=\delta_{ac}\delta_{bd}-\delta_{ad}\delta_{bc}\), the six resulting Kronecker terms cancelling in pairs; the triple \((P_{i},P_{j},P_{l})\) is the same computation with \(\alpha\to\beta\) and \(K\to P\). Four independent triples remain.
The triple \((K_{i},K_{j},P_{l})\). Using \(\comm{J_{k}}{P_{l}}=\epsilon_{klm}P_{m}\) and then the same \(\epsilon\)-identity,
The sum is \(\left(\alpha+\gamma\mu\right)\left(\delta_{il}P_{j}-\delta_{jl}P_{i}\right)\), and the two tensors \(\delta_{il}P_{j}\) and \(\delta_{jl}P_{i}\) are linearly independent, so \(\alpha=-\gamma\mu\).
The triple \((P_{i},P_{j},K_{l})\). Identically,
whose sum is \(\left(\beta-\gamma\nu\right)\left(\delta_{il}K_{j}-\delta_{jl}K_{i}\right)\), so \(\beta=\gamma\nu\).
The triples containing \(H\). For \((K_{i},K_{j},H)\) the three terms are \(\alpha\epsilon_{ijk}\comm{J_{k}}{H}=0\) by Equation (A.161), together with \(\mu\comm{P_{j}}{K_{i}}-\mu\comm{P_{i}}{K_{j}} =\mu\gamma\delta_{ij}H-\mu\gamma\delta_{ij}H=0\); the identity holds with no condition. The triple \((P_{i},P_{j},H)\) is the same with \(\mu\to\nu\). For \((K_{i},P_{j},H)\),
so \(\alpha\nu+\beta\mu=0\) — which is implied by Equation (A.167), since \(-\gamma\mu\nu+\gamma\nu\mu=0\). No further condition arises, and the two constraints of Equation (A.167) are the whole content of the Jacobi identity.
∎Step three: rescaling reduces the constants to their signs
Two kinematical algebras describe the same kinematics if they differ only by a choice of units for boosts, lengths and times. The corresponding transformations
leave Equation (A.161) and both automorphisms of (H2) intact, and they exhaust that freedom: \(J\) is fixed by the normalization of Equation (A.161).
Under Equation (A.168) the three constants transform as
Consequently:
-
which of \(\gamma,\mu,\nu\) vanish is invariant;
-
every nonvanishing constant can be scaled to \(\pm1\), and all of them simultaneously; and
-
the signs of the surviving constants are determined only up to a simultaneous flip of all three, so that the invariants are the pairwise products \(\sgn(\gamma\mu)\), \(\sgn(\gamma\nu)\) and \(\sgn(\mu\nu)\), of which two are independent.
Rests on Lemma A.57 and Equation (A.168).
Derives Lemma A.58. Equation (A.169) follows by substituting Equation (A.168) into Equations (A.165) and (A.166) and re-expressing the right-hand sides in the new generators; claim (1) is immediate from it, each constant being multiplied by a nonzero number. For (2), take logarithms of the absolute values: the exponents of \((\lambda,\rho,\sigma)\) form the matrix
so the system can be solved for \(\log\abs{\lambda},\log\abs{\rho},\log\abs{\sigma}\) whatever shifts are required; when only some constants are nonzero the corresponding subsystem is solvable a fortiori, the matrix having full rank on any subset of its rows. For (3), write \(s_{\lambda},s_{\rho},s_{\sigma}\) for the signs of the scale factors. Each of the three coefficients in Equation (A.169) has the same sign, namely \(s=s_{\lambda}s_{\rho}s_{\sigma}\), because each is a product of all three scale factors with one of them inverted and inversion does not change a sign. So all three constants are multiplied by one common \(s\in\set{\pm1}\), which may be chosen freely. The pairwise products are unchanged by that flip, and their own product is \(\left(\gamma\mu\nu\right)^{2}>0\), so only two of them are independent.
∎Step four: the enumeration
By Lemma A.58(1) the classification splits into the eight cases labelled by which of \((\gamma,\mu,\nu)\) vanish; within each case, Lemma A.58(2) and (3) reduce the surviving constants to a sign pattern modulo one simultaneous flip. Writing \(\ast\) for a nonzero constant, the count is as follows.
-
[\((0,0,0)\) — one algebra.] Every bracket among \(K,P,H\) vanishes, and with them \(\alpha\) and \(\beta\). This is the static algebra.
-
[\((\ast,0,0)\) — one algebra.] \(\alpha=\beta=0\), and the single surviving constant has a sign that the flip removes. Here \(\comm{K_{i}}{P_{j}}=\delta_{ij}H\) with abelian boosts: the Carroll algebra.
-
[\((0,\ast,0)\) — one algebra.] Likewise one algebra, \(\comm{K_{i}}{H}=P_{i}\) with every other new bracket zero: the Galilei algebra. Note that \(\gamma=0\) here, so \(\comm{K_{i}}{P_{j}}=0\) — the mass of Equation (25.15) is not in this algebra, which is exactly why it can only enter as a central extension.
-
[\((0,0,\ast)\) — one algebra.] The para-Galilei algebra, \(\comm{P_{i}}{H}=K_{i}\).
-
[\((0,\ast,\ast)\) — two algebras.] \(\alpha=\beta=0\) and the invariant is \(\sgn(\mu\nu)\): the Newton–Hooke algebras \(\mathrm{NH}_{+}\) and \(\mathrm{NH}_{-}\), whose difference is exponential against oscillatory behaviour of the free motion.
-
[\((\ast,0,\ast)\) — two algebras.] \(\alpha=0\) but \(\beta=\gamma\nu\neq0\), and the invariant is \(\sgn\beta\). The set \(\set{J_{i},P_{i}}\) closes into \(\mathfrak{so}(3,1)\) for \(\beta<0\) and into \(\mathfrak{so}(4)\) for \(\beta>0\), while \(\set{K_{i},H}\) is an abelian ideal in both cases. These are para-Poincaré and the inhomogeneous \(\SO(4)\). The boosts are abelian here, so (H3) is satisfied and neither is excluded.
-
[\((\ast,\ast,0)\) — one algebra.] \(\beta=0\) but \(\alpha=-\gamma\mu\neq0\), and the invariant is \(\sgn\alpha\); now \(\set{J_{i},K_{i}}\) closes into \(\mathfrak{so}(3,1)\) for \(\alpha<0\) and into \(\mathfrak{so}(4)\) for \(\alpha>0\). In the second case the boosts generate a compact subgroup, so (H3) excludes it and only \(\alpha<0\) survives: the Poincaré algebra \(\mathfrak{t}(4)\,\bar{\oplus}\,\mathfrak{so}(3,1)\) of Equation (14.99).
-
[\((\ast,\ast,\ast)\) — two algebras.] Two independent sign invariants give four candidates; (H3) again forces \(\alpha=-\gamma\mu<0\), halving them to two, distinguished by \(\sgn(\gamma\nu)\). These are the two simple algebras of Equations (14.109) and (14.110), \(\mathfrak{so}(4,1)\) and \(\mathfrak{so}(3,2)\): de Sitter and anti-de Sitter.
Adding the cases gives
which proves Theorem 14.90. The assignment of the traditional names to the classes is Bacry and Lévy-Leblond's [Bacry:1968]; what is established above is the count, the parametrization, and the structure of each class.
Two corollaries the enumeration makes visible
Each of the three constants can be sent to zero by an Inönü–Wigner contraction, in the sense of Definition 14.83, that leaves the other two fixed:
in each case with \(\varepsilon\to0\). The three contractions commute, so from an algebra with all three constants nonzero they generate exactly the eight vanishing patterns of Step four: the enumeration. The kinematical cube is therefore the lattice of subsets of \(\set{\gamma,\mu,\nu}\): its eight vertices are the eight cases, not eight of the eleven algebras. Three of the eleven — one de Sitter, one Newton–Hooke, and the inhomogeneous \(\SO(4)\) — are opposite-sign partners of a vertex rather than vertices of their own. Rests on Lemma A.58 and Definition 14.83.
Derives Corollary A.59. Apply Equation (A.169) to each substitution. For Equation (A.171), \(\lambda=1\) and \(\rho=\sigma=\varepsilon\) give \(\gamma\mapsto\gamma\), \(\mu\mapsto\mu\), \(\nu\mapsto\varepsilon^{2}\nu\). For Equation (A.172), \(\lambda=\rho=\varepsilon\) and \(\sigma=1\) give \(\gamma\mapsto\varepsilon^{2}\gamma\) with \(\mu\) and \(\nu\) fixed. For Equation (A.173), \(\lambda=\sigma=\varepsilon\) and \(\rho=1\) give \(\mu\mapsto\varepsilon^{2}\mu\) with \(\gamma\) and \(\nu\) fixed. In each case no structure constant diverges, so the limit exists and is a contraction. Commutativity is clear, the three substitutions acting as commuting diagonal matrices on \((\lambda,\rho,\sigma)\).
∎The linear map \(\Sigma\) exchanging boosts and space translations,
carries the algebra with constants \((\gamma,\mu,\nu)\) to the algebra with constants \((-\gamma,\nu,\mu)\), and so exchanges the cases \((\ast,\ast,0)\) and \((\ast,0,\ast)\), and the cases \((0,\ast,0)\) and \((0,0,\ast)\). It therefore identifies, as abstract Lie algebras,
and identifies the inhomogeneous \(\SO(4)\) with the compact-boost algebra that (H3) excluded from the case \((\ast,\ast,0)\). Rests on Lemmas A.56 and A.57.
Derives Corollary A.60. Under \(\Sigma\) the bracket \(\comm{K_{i}}{P_{j}}=\gamma\delta_{ij}H\) becomes \(\comm{P_{i}}{K_{j}}=\gamma\delta_{ij}H\), that is \(\comm{K_{j}}{P_{i}}=-\gamma\delta_{ij}H\), so \(\gamma\mapsto-\gamma\); and \(\comm{K_{i}}{H}=\mu P_{i}\) becomes \(\comm{P_{i}}{H}=\mu K_{i}\), so \(\nu\mapsto\mu\) and, symmetrically, \(\mu\mapsto\nu\). The constrained constants follow from Equation (A.167): \(\alpha'=-\gamma'\mu'=\gamma\nu =\beta\) and \(\beta'=\gamma'\nu'=-\gamma\mu=\alpha\), so \(\Sigma\) exchanges the two quadratic constants as well. Applying this to Poincaré, which is \((\gamma,\mu,0)\) with \(\alpha<0\), gives \((-\gamma,0,\mu)\) with \(\beta'=\alpha<0\) — the case \((\ast,0,\ast)\) with negative \(\beta\), which is para-Poincaré. The same substitution applied to \((0,\mu,0)\) gives \((0,0,\mu)\), and applied to the excluded \(\alpha>0\) member of \((\ast,\ast,0)\) gives the \(\beta>0\) member of \((\ast,0,\ast)\), the inhomogeneous \(\SO(4)\).
∎Corollary A.60 is the reason the eleven of Theorem 14.90 must be counted as eleven kinematics and not as eleven isomorphism classes of Lie algebras. The map \(\Sigma\) is an isomorphism of Lie algebras, but it is not an equivalence of kinematics: it fails to commute with the time reversal Equation (A.163), since it sends a \(\Theta\)-odd generator to a \(\Theta\)-even one, and it exchanges the physical roles of a boost and a translation. Two of the eleven are therefore abstractly the same algebra as another member of the list while describing a different physics, and Table 14.1 lists them separately for that reason. A statement of the theorem phrased as “eleven algebras up to isomorphism” is wrong, and the error is an easy one to make.
Appendix A.10 discharges the proof obligation of Theorem 14.90. Its yield for the rest of the book is Equation (A.167) and the three contractions of Corollary A.59. The second of them, Equation (A.172), is the \(c\to\infty\) limit carrying the Poincaré algebra to the Galilei algebra, and the fact that it sets \(\gamma=0\) — so that \(\comm{K_{i}}{P_{j}}\) vanishes identically in the contracted algebra — is precisely what leaves room for the mass to reappear there as the central charge of Proposition 25.14.
Invariant Bilinear Forms on an Expanded Algebra
This appendix proves Theorem 15.54 (Expansions of Lie Algebras): on \(\mathfrak{g}_{A}=A\otimes\mathfrak{g}\) every invariant symmetric bilinear form has the factorised shape Equation (15.61), so for \(k=2\) the family constructed there is already the whole of it. This is the case the book uses — it is what makes Theorem 15.58 a statement about all invariant pairings and not merely about the ones written down. It is also the only case in which the statement is true: Proposition 15.55 exhibits an invariant symmetric \(4\)-linear form that is not factorised, and Remark A.66 at the end says precisely which step of the argument below fails for \(k\ge3\) and why no repair is possible.
The hypothesis on \(\mathfrak{g}\) needs care, and the care is not academic: the Lorentz algebra of the observed \(3{+}1\) dimensions is exactly the case where the naive hypothesis fails.
A real Lie algebra \(\mathfrak{g}\) is absolutely simple if its complexification \(\mathfrak{g}\otimes_{\R}\C\) is a simple complex Lie algebra. Every absolutely simple algebra is simple; the converse fails, the standard counterexample being a complex simple algebra regarded as a real one.
Let \(\mathfrak{g}\) be absolutely simple with Killing form \(\kappa_{\mathfrak{g}}\). Then every invariant bilinear form on \(\mathfrak{g}\) is a real multiple of \(\kappa_{\mathfrak{g}}\); in particular the space of such forms is one-dimensional and every one of them is symmetric. Rests on Definition A.62.
Derivation. Derives Lemma A.63. Let \(B\) be an invariant bilinear form, that is
Define \(\beta:\mathfrak{g}\to\mathfrak{g}^{*}\) by \(\beta(X)=B(X,\cdot)\). Equation (A.176) says exactly that \(\beta\) intertwines the adjoint representation on \(\mathfrak{g}\) with the coadjoint representation on \(\mathfrak{g}^{*}\). Since \(\mathfrak{g}\) is simple it is semisimple, so \(\kappa_{\mathfrak{g}}\) is nondegenerate by Cartan's criterion and the induced map \(\kappa_{\mathfrak{g}}^{\flat}:\mathfrak{g}\to\mathfrak{g}^{*}\) is an isomorphism of representations. Hence
the algebra of endomorphisms commuting with the adjoint action, and \(B=\kappa_{\mathfrak{g}}(T\,\cdot,\cdot)\).
It remains to show \(\operatorname{End}_{\mathfrak{g}}(\mathfrak{g})=\R\). Extending scalars,
since forming the commutant of a set of operators commutes with extension of scalars. By hypothesis \(\mathfrak{g}_{\C}\) is simple, so its adjoint representation is irreducible over \(\C\), and Schur's lemma over an algebraically closed field gives \(\operatorname{End}_{\mathfrak{g}_{\C}}(\mathfrak{g}_{\C})=\C\). Hence \(\dim_{\R}\operatorname{End}_{\mathfrak{g}}(\mathfrak{g})=1\), so \(T=\lambda\,\id\) and \(B=\lambda\,\kappa_{\mathfrak{g}}\), which is symmetric.
∎Let \(\mathfrak{g}\) be absolutely simple and let \(A\) be a finite-dimensional commutative associative unital \(\R\)-algebra. Then for every invariant symmetric bilinear form \(\Omega\) on \(\mathfrak{g}_{A}\) there is a unique linear functional \(\phi:A\to\R\) with
Equivalently \(\Omega=\avg{\cdot,\cdot}_{\kappa_{\mathfrak{g}},\phi}\) in the notation of Equation (15.61), and \(\phi\mapsto\avg{\cdot,\cdot}_{\kappa_{\mathfrak{g}},\phi}\) is a linear isomorphism from \(A^{*}\) onto the space of invariant symmetric bilinear forms on \(\mathfrak{g}_{A}\), which therefore has dimension \(\dim A\). Rests on Definition A.62, Lemma A.63, Lemma 15.50, Theorem 15.51 and Equation (15.2).
Derivation. Derives Theorem A.64. Step 1: for fixed \(a,b\) the form is a multiple of the Killing form. Fix \(a,b\in A\) and define the bilinear form
on \(\mathfrak{g}\). Invariance of \(\Omega\), applied with the element \(1_{A}\otimes Z\) and using \(\comm{1_{A}\otimes Z}{a\otimes X}=a\otimes\comm{Z}{X}\) from Equation (15.2), gives
that is, \(\omega_{a,b}\) satisfies Equation (A.176). By Lemma A.63 there is a unique real number \(\Psi(a,b)\) with
Since \(\Omega\) is bilinear in its two arguments and \(\kappa_{\mathfrak{g}}\neq0\), the function \(\Psi:A\times A\to\R\) is bilinear; since \(\Omega\) and \(\kappa_{\mathfrak{g}}\) are symmetric, \(\Psi\) is symmetric.
Step 2: \(\Psi\) is a \(2\)-trace. Now use invariance of \(\Omega\) with a general element \(c\otimes Z\). By Equation (15.2), \(\comm{c\otimes Z}{a\otimes X}=(ca)\otimes\comm{Z}{X}\), so
which by Equation (A.179) reads
The Killing form is itself invariant, so \(\kappa_{\mathfrak{g}}(X,\comm{Z}{Y}) =-\kappa_{\mathfrak{g}}(\comm{Z}{X},Y)\), and Equation (A.180) collapses to
The Killing factor cannot vanish identically: a simple algebra is perfect, so the elements \(\comm{Z}{X}\) span \(\mathfrak{g}\), and \(\kappa_{\mathfrak{g}}\) is nondegenerate, so some choice of \(X,Y,Z\) makes \(\kappa_{\mathfrak{g}}(\comm{Z}{X},Y)\neq0\). Hence
which is Equation (15.58) at \(k=2\): \(\Psi\) is a \(2\)-trace on \(A\).
Step 3: a \(2\)-trace is a linear functional of the product. By Lemma 15.50 with \(k=2\) there is a unique linear \(\phi:A\to\R\), namely \(\phi(x)=\Psi(x,1_{A})\), with \(\Psi(a,b)=\phi(ab)\). Substituting into Equation (A.179) gives Equation (A.177).
Step 4: uniqueness and bijectivity. If \(\phi\) and \(\phi'\) both satisfy Equation (A.177) then \(\left(\phi(ab)-\phi'(ab)\right)\kappa_{\mathfrak{g}}(X,Y)=0\); choosing \(X,Y\) with \(\kappa_{\mathfrak{g}}(X,Y)\neq0\) and \(b=1_{A}\) gives \(\phi=\phi'\). Conversely every \(\phi\in A^{*}\) produces an invariant symmetric bilinear form by the first part of Theorem 15.51, and the assignment is visibly linear in \(\phi\), while Equation (A.177) says it is onto. The two spaces are therefore isomorphic.
∎Lemma A.63 needs \(\mathfrak{g}\) absolutely simple, not merely simple, and the difference is not a technicality in this book. The Lorentz algebra of four-dimensional spacetime, \(\mathfrak{so}(3,1)\), is simple as a real Lie algebra but not absolutely simple: its complexification is \(\mathfrak{sl}_{2}(\C)\oplus\mathfrak{sl}_{2}(\C)\), which is not simple — this is the self-dual/anti-self-dual splitting. Accordingly it carries a two-dimensional space of invariant symmetric bilinear forms, spanned by
the first proportional to the Killing form and the second available only because \(D=4\) supplies a rank-four invariant alternating symbol. Both are invariant, and the two Chern–Weil \(4\)-forms they generate are the two topological densities of four-dimensional gravity: the Killing form gives \(R^{\mu\nu}\wedge R_{\mu\nu}\), the Pontryagin density, and \(\epsilon_{\mu\nu\rho\sigma}\) gives \(\epsilon_{\mu\nu\rho\sigma}R^{\mu\nu}\wedge R^{\rho\sigma}\), the Euler density of the Gauss–Bonnet term. That a four-dimensional gauge theory of the Lorentz group admits two such terms where a general dimension admits one is exactly this failure of absolute simplicity. For such a \(\mathfrak{g}\) the conclusion of Theorem A.64 must be replaced by a sum over a basis of the invariant forms of \(\mathfrak{g}\), one functional \(\phi\) for each. The algebras to which Theorem 15.29 attaches the three signs of the cosmological constant — \(\mathfrak{so}(3,2)\) and \(\mathfrak{so}(4,1)\) in \(3{+}1\) — are absolutely simple, their complexifications being \(\mathfrak{sp}(4,\C)\) and \(\mathfrak{so}(5,\C)\), so Theorem A.64 applies to them verbatim.
It is worth being explicit about where the proof above uses \(k=2\), because the manuscript this chapter follows [Gonzalez:2026] states the converse for every \(k\), and that statement is false.
Step 1 generalises without difficulty: evaluating invariance on \(1_{A}\otimes Z\) shows, for any \(k\), that \((X_{1},\dots,X_{k})\mapsto \Omega(a_{1}\otimes X_{1},\dots,a_{k}\otimes X_{k})\) is an invariant \(k\)-linear form on \(\mathfrak{g}\). Step 2 does not. It works here because the invariance relation for a bilinear form has exactly two terms, and the invariance of \(\kappa_{\mathfrak{g}}\) itself collapses them into the single factor Equation (A.181); for \(k\ge3\) the same manipulation leaves a sum of \(k-1\) terms with no reason to vanish separately. That is not a defect of the argument: by Proposition 15.55 there is an invariant symmetric \(4\)-linear form on \(\mathfrak{so}(2,1)\gen{1}=\mathfrak{iso}(2,1)\) — the square of the standard pairing \(\avg{J_{a},P_{b}}=\eta_{ab}\) — which is not factorised at all, so no argument could have succeeded.
The structural reason is Theorem 15.53: the factorised forms are exactly the balanced ones, and a product of two invariant bilinear forms is not balanced, because \(\phi(a_{1}a_{2})\phi(a_{3}a_{4})\) is not a function of \(a_{1}a_{2}a_{3}a_{4}\). At \(k=2\) no product is available, since \(\mathrm{Inv}^{1}(\mathfrak{g})=(\mathfrak{g}^{*})^{\mathfrak{g}}\) vanishes for perfect \(\mathfrak{g}\) — which is exactly what Theorem A.64 exploits without saying so, and is the whole of the difference between the two cases.
Invariant Bilinear Forms on an Expanded Algebra discharges the proof obligation of Theorem 15.54. Its yield for the rest of the book is that Theorem 15.58 classifies all invariant pairings on a resonant subalgebra and not merely the constructed ones, which is what makes the statement that the Galilei algebra admits none, while the Bargmann algebra does, a theorem rather than a failure to find one. Its other yield is negative and equally worth having: together with Proposition 15.55 it fixes the exact reach of the factorisation, which is \(k=2\) and no further.
Darboux's Theorem: Local Canonical Coordinates
This appendix proves the geometric statement announced in Remark 5.116 of Linear Algebra and Representation Theory: a closed non-degenerate two-form on a manifold can be brought, in a neighbourhood of any point, to the constant normal form \(\sum_{i}\dd q^{i}\wedge\dd p_{i}\). Symplectic geometry therefore has no local invariants whatever — no analogue of the curvature that distinguishes one Riemannian metric from another near a point — so that the absence of a signature established pointwise in Proposition 5.115 propagates to the manifold.
What is assumed as already available is exactly the differential-form machinery of Differentiable Manifolds, Tensors, and Curvature: \(k\)-forms (Definition 13.98), the wedge product (Definition 13.102), the exterior derivative with its nilpotency and Leibniz rule (Definition 13.103, Equation (13.245) and Equation (13.246)), closed and exact forms (Definitions 13.105 and 13.106), the Lie derivative (Equation (13.267) and Proposition 13.127) and Cartan's magic formula (Proposition 13.128); from analysis, the mean value theorem (Theorem 7.35), the Cauchy criterion (Theorem 7.8), uniform continuity on a compact set (Theorem 7.25) and the Picard–Lindelöf theorem with its explicit contraction estimate (Theorem 9.8 and Corollary 9.9). Everything else — the pullback of a \(k\)-form, the smoothness of the flow of a time-dependent vector field, and the local converse of the Poincaré lemma — is built here, because the treatise does not carry it. The pointwise input is Proposition 5.115, and the argument that converts it into a statement about a varying form is the homotopy argument known as Moser's trick.
Statement
A symplectic manifold is a pair \((M,\omega)\) with \(M\) a smooth manifold (Definition 13.48) and \(\omega\) a smooth two-form on \(M\) (Definition 13.98) which is
-
closed, \(\dd\omega = 0\) (Definition 13.105); and
-
non-degenerate at every point: for each \(P\in M\) the matrix \(\omega_{ij}(P)\) of components in any chart is invertible, equivalently \(\omega_{P}(u,v)=0\) for all \(v\in T_{M}(P)\) forces \(u=0\).
Let \((M,\omega)\) be a symplectic manifold and \(P\in M\). Then \(\dim M = 2m\) is even, and there is a chart \((u^{1},\ldots,u^{2m})\) of \(M\) defined on a neighbourhood of \(P\), with \(u(P)=0\), in which
Equivalently, the components of \(\omega\) in this chart are the constant block array Equation (5.144). Rests on Definition A.67, Proposition 5.115 and Equation (5.144).
Any two symplectic manifolds of the same dimension are locally symplectomorphic: given \(P\in M\) and \(P'\in M'\) there are neighbourhoods \(U\ni P\), \(U'\ni P'\) and a diffeomorphism \(f:U\longrightarrow U'\) with \(f^{\ast}\omega' = \omega\). In particular no function of \(\omega\) and its derivatives at a point can be a symplectic invariant. Rests on Theorem A.68.
Derives Corollary A.69. Let \(u\) and \(u'\) be canonical charts around \(P\) and \(P'\) furnished by Theorem A.68, both with image containing a ball \(B\subset\R^{2m}\) about the origin. Put \(f = (u')^{-1}\circ u\) on \(U = u^{-1}(B)\). Both \(\omega\) and \(\omega'\) read as the same constant form Equation (A.184) in their own charts, and the pullback of a form is computed chartwise (Definition A.70 below), so \(f^{\ast}\omega' = \omega\). Any local invariant would have to take the same value at \(P\) and \(P'\), and \(\omega\), \(\omega'\) were arbitrary.
∎The proof of Theorem A.68 occupies the rest of this section. The pullback of a $k$-form builds the pullback of a \(k\)-form and its two structural properties; Flows of time-dependent vector fields proves that the flow of a time-dependent vector field exists and is smooth, and differentiates a pullback along it; The local converse of the Poincaré lemma proves the local converse of the Poincaré lemma on a star-shaped set; and Moser's homotopy argument assembles the three into Moser's homotopy argument.
Throughout, \(\abs{\cdot}\) is the Euclidean length on \(\R^{n}\) (Equation (6.12)) and, for a real \(n\times n\) matrix \(A\),
the two inequalities following from the Cauchy–Schwarz inequality of Linear Algebra and Representation Theory applied row by row.
The pullback of a $k$-form
Differentiable Manifolds, Tensors, and Curvature defines the pullback of a covector; the extension to a \(k\)-form is forced by the requirement that it act on \(k\) pushed-forward vectors.
Let \(f:U\longrightarrow V\) be a smooth map between open subsets of \(\R^{m}\) and \(\R^{n}\), with components \(f^{j}\), and let \(\alpha\) be a \(k\)-form on \(V\) with components \(\alpha_{j_{1}\ldots j_{k}}\). The pullback \(f^{\ast}\alpha\) is the \(k\)-form on \(U\) with components
For \(k=1\) this is the covector pullback of Section 13.5.2, and for \(k=0\) it is \(f^{\ast}g = g\circ f\). The right-hand side is totally antisymmetric in \(i_{1},\ldots,i_{k}\) because \(\alpha\) is antisymmetric in \(j_{1},\ldots,j_{k}\) and the two index groups are contracted in order, so \(f^{\ast}\alpha\) is again a \(k\)-form (Definition 13.98). Rests on Definitions 7.65, 13.83 and 13.98.
With \(f\) as above, and \(\alpha\), \(\beta\) forms on \(V\) of degrees \(k\) and \(l\):
the last for a further smooth \(g\) defined on the target of the form. Rests on Definition A.70, Equation (13.246) and Equation (13.245).
Derives Lemma A.71. Composition. Equation (A.189) is the chain rule Equation (7.52) inserted \(k\) times into Equation (A.186): \(\pp_{i}\left(g\circ f\right)^{j} = \pp_{l}g^{j}(f(x))\,\pp_{i}f^{l}(x)\).
Scalars and their differentials. For \(k=0\), \(f^{\ast}g=g\circ f\) and \(\pp_{i}(g\circ f) = \pp_{j}g(f(x))\,\pp_{i}f^{j}\), which is Equation (A.186) for the one-form \(\dd g\); hence
Wedge. It suffices to check Equation (A.187) on the components. Both sides are multilinear in \(\alpha\) and \(\beta\), so it is enough to take \(\alpha = \dd y^{a_{1}}\wedge\cdots\wedge\dd y^{a_{k}}\) and \(\beta = \dd y^{b_{1}}\wedge\cdots\wedge\dd y^{b_{l}}\) with the \(y\) the coordinates on \(V\). By Definition 13.102 and Equation (A.186), the pullback of a wedge of coordinate differentials is the wedge of their pullbacks: for two of them,
and the same cancellation of index pairs, term by term over the permutations, gives \(f^{\ast}\left(\dd y^{a_{1}}\wedge\cdots\wedge\dd y^{a_{k}}\right) = \dd f^{a_{1}}\wedge\cdots\wedge\dd f^{a_{k}}\), where \(\dd f^{a} = f^{\ast}\dd y^{a}\) by Equation (A.190). Equation (A.187) follows because both sides then read \(\dd f^{a_{1}}\wedge\cdots\wedge\dd f^{a_{k}}\wedge \dd f^{b_{1}}\wedge\cdots\wedge\dd f^{b_{l}}\).
Exterior derivative. Write \(\alpha = \frac{1}{k!}\alpha_{a_{1}\ldots a_{k}}\, \dd y^{a_{1}}\wedge\cdots\wedge\dd y^{a_{k}}\). By the two paragraphs just proved,
Applying \(\dd\) and using the Leibniz rule Equation (13.246) repeatedly, every term in which \(\dd\) falls on some \(\dd f^{a}\) vanishes by nilpotency Equation (13.245), so
the last equality by Equation (A.190) applied to the component functions and one more use of the two paragraphs above.
∎Flows of time-dependent vector fields
Moser's argument needs a family of diffeomorphisms generated by a vector field that changes with the parameter, and needs it to be smooth in the point as well as in the parameter. Ordinary Differential Equations and Sturm–Liouville Theory supplies existence and uniqueness (Theorem 9.8); the dependence on the initial point is proved here.
Let \(u:[0,T]\longrightarrow[0,\infty)\) be continuous and suppose
with constants \(c\ge0\), \(L\ge0\). Then \(u(t)\le c\,\ee^{Lt}\). Rests on Theorem 7.43 and Definition 7.53.
Derives Lemma A.72. Let \(S=\max_{[0,T]}u\), finite by Theorem 7.24. We show by induction that for every \(N\ge0\)
For \(N=0\), Equation (A.191) and \(u\le S\) give \(u(t)\le c + LSt\). Assuming Equation (A.192) at \(N\) and substituting it into Equation (A.191), term-by-term integration gives
which is the statement at \(N+1\). The last term of Equation (A.192) tends to \(0\) as \(N\to\infty\), because the exponential series converges (Definition 7.53), and the sum tends to \(\ee^{Lt}\).
∎Let \(K\subset\R^{n}\) be open, \(F:[0,1]\times K\longrightarrow\R\) continuous with \(\pp F/\pp x^{i}\) existing and continuous on \([0,1]\times K\) for each \(i\). Then \(G(x)=\int_{0}^{1}F(t,x)\,\dd t\) has continuous partial derivatives on \(K\) and
Rests on Theorem 7.35, Theorem 7.25 and Definition 7.65.
Derives Lemma A.73. Fix \(x\in K\) and \(r>0\) with \(\overline{B}_{r}(x)\subset K\). For \(0<\abs{s}\le r\) the mean value theorem Theorem 7.35 applied to \(\lambda\mapsto F(t,x+\lambda\vect{e}_{i})\) gives \(\theta = \theta(t,s)\in(0,1)\) with
The set \([0,1]\times\overline{B}_{r}(x)\) is compact, so \(\pp_{i}F\) is uniformly continuous there (Theorem 7.25): given \(\varepsilon>0\) there is \(\delta>0\) with \(\abs{\pp_{i}F(t,y)-\pp_{i}F(t,x)}<\varepsilon\) whenever \(\abs{y-x}<\delta\), uniformly in \(t\). Hence for \(\abs{s}<\delta\) the difference quotient of \(G\) differs from \(\int_{0}^{1}\pp_{i}F(t,x)\dd t\) by at most \(\varepsilon\), which is Equation (A.193). Continuity of the right-hand side in \(x\) follows from the same uniform continuity.
∎Let \(U\subseteq\R^{n}\) be open, let \(X:[0,1]\times U\longrightarrow\R^{n}\) be smooth — all partial derivatives in \((t,x)\) of every order exist and are continuous — and let \(p\in U\) satisfy \(X_{t}(p)=0\) for every \(t\in[0,1]\). Then there is an open ball \(B\ni p\) with \(\overline{B}\subset U\) and a smooth map \(\psi:[0,1]\times B\longrightarrow U\) with
and each \(\psi_{t}=\psi(t,\cdot)\) is a diffeomorphism of \(B\) onto an open subset of \(U\). Rests on Theorem 9.8, Corollary 9.9 and Lemma A.72.
Derives Theorem A.74. A Lipschitz box. Choose \(r>0\) with \(\overline{B}_{2r}(p)\subset U\) and put \(Q = [0,1]\times\overline{B}_{2r}(p)\), a compact set. On \(Q\) the continuous functions \(X\) and \(\pp_{j}X^{i}\) are bounded (Theorem 7.24); set \(M=\max_{Q}\abs{X}\) and
with \(D_{x}X\) the matrix \(\pp_{j}X^{i}\) and \(\norm{\cdot}\) as in Equation (A.185). Since \(\overline{B}_{2r}(p)\) is convex, Theorem 7.35 applied componentwise along the segment from \(z\) to \(y\) bounds each component of the difference by \(\abs{\nabla_{x}X^{i}}\,\abs{y-z}\le\norm{D_{x}X}\abs{y-z}\), and summing the \(n\) squares gives
Existence on all of \([0,1]\). The constant curve \(y\equiv p\) solves Equation (A.194) with \(y(0)=p\), because \(X_{t}(p)=0\); by uniqueness in Theorem 9.8 it is the solution through \(p\), which is the third assertion of Equation (A.194). Let now \(x\in B := B_{\delta}(p)\) with \(\delta = r\,\ee^{-L}\), and let \(\psi_{t}(x)\) be the maximal solution. So long as it remains in \(\overline{B}_{2r}(p)\), the integral form of the equation and Equation (A.196) give
so Lemma A.72 yields
The solution therefore never approaches the boundary of \(\overline{B}_{2r}(p)\). Since the Picard step \(h=\min\set{1,\;r/M}\) of Theorem 9.8 is the same at every starting time in \([0,1]\) and every starting point of \(\overline{B}_{r}(p)\), the solution can be continued in steps of that fixed length and reaches \(t=1\) after finitely many of them.
Lipschitz dependence. For \(x,y\in B\) the same computation with \(p\) replaced by \(y\), using Equation (A.196) (both curves stay in \(\overline{B}_{r}(p)\) by Equation (A.197)) and Lemma A.72, gives
The variational equation. Put \(A(t,x) = D_{x}X\bigl(t,\psi_{t}(x)\bigr)\), continuous on \([0,1]\times B\) with \(\norm{A}\le L\), and let \(\Phi(\cdot,x)\) be the unique solution of the linear system
which exists on the whole of \([0,1]\) by Corollary 9.9. Its integral form and Lemma A.72 give \(\norm{\Phi(t,x)}\le\sqrt{n}\,\ee^{L}\).
\(\Phi\) is continuous in \(x\). Subtracting the integral forms of Equation (A.199) at \(x\) and at \(y\),
The second integral is bounded by \(\sqrt{n}\,\ee^{L}\,\eta(x,y)\) with \(\eta(x,y)=\max_{s}\norm{A(s,x)-A(s,y)}\), and \(\eta(x,y)\to0\) as \(y\to x\): \(D_{x}X\) is uniformly continuous on the compact \(Q\) (Theorem 7.25) and \(\abs{\psi_{s}(x)-\psi_{s}(y)}\le\ee^{L}\abs{x-y}\) by Equation (A.198). Lemma A.72 then gives \(\norm{\Phi(t,x)-\Phi(t,y)}\le\sqrt{n}\,\ee^{2L}\eta(x,y)\).
\(\Phi\) is the derivative. Fix \(x\in B\) and let \(h\) be small enough that \(x+h\in B\). Put \(\Delta(t) = \psi_{t}(x+h)-\psi_{t}(x)\) and \(R(t) = \Delta(t)-\Phi(t,x)h\), so \(R(0)=0\) and
where \(\rho(t) = X_{t}(z+\Delta)-X_{t}(z)-D_{x}X(t,z)\Delta\) with \(z=\psi_{t}(x)\). Applying Theorem 7.35 to \(\lambda\mapsto X^{i}_{t}(z+\lambda\Delta)\) gives \(\theta_{i}\in(0,1)\) with \(\rho^{i}(t) = \left[\pp_{j}X^{i}(t,z+\theta_{i}\Delta) -\pp_{j}X^{i}(t,z)\right]\Delta^{j}\), whence \(\abs{\rho(t)}\le w\bigl(\abs{\Delta}\bigr)\abs{\Delta}\), where \(w\) is \(\sqrt{n}\) times a modulus of uniform continuity of \(D_{x}X\) on the compact set \(Q\) (Theorem 7.25), so that \(w(\lambda)\to0\) as \(\lambda\to0\). With \(\abs{\Delta}\le\ee^{L}\abs{h}\) from Equation (A.198), the integral form of the equation for \(R\) and Lemma A.72 give
so \(\abs{R(t)}/\abs{h}\to0\) as \(h\to0\). By Definition 7.67, \(\psi_{t}\) is differentiable at \(x\) with differential \(\Phi(t,x)\); the previous paragraph makes the partial derivatives continuous, so \(\psi_{t}\) is \(C^{1}\) (Definition 7.66).
Derivatives of every order. We show by induction on \(k\ge1\): if \(X\) is \(C^{k}\) then \(\psi\) is \(C^{k}\) in \(x\). The case \(k=1\) is what was just proved. Let \(X\) be \(C^{k+1}\) and consider the enlarged system on \(U\times\R^{n\times n}\),
whose right-hand side is \(C^{k}\) in \((y,\Phi)\) because \(X\) is \(C^{k+1}\). Its solution is exactly \(\bigl(\psi(t,x),\Phi(t,x)\bigr)\), and the induction hypothesis applied to Equation (A.200) makes it \(C^{k}\) in the initial data, hence in \(x\). So \(\pp\psi/\pp x = \Phi\) is \(C^{k}\) in \(x\), i.e.\ \(\psi\) is \(C^{k+1}\) in \(x\). Since moreover \(\pp_{t}\psi = X_{t}(\psi)\), every derivative in \(t\) is a polynomial expression in derivatives of \(X\) and of \(\psi\) in \(x\); hence all mixed partial derivatives of \(\psi\) in \((t,x)\) of every order exist and are continuous, and \(\psi\) is smooth.
Diffeomorphism. Fix \(t\in[0,1]\) and let \(\eta_{s}\) be the flow built by the same construction for the field \(Y_{s}(y) = -X_{t-s}(y)\), which also vanishes at \(p\); let \(B_{2}\) be its ball. That ball is the same for every \(t\): the construction used only \(r\), \(M\) and \(L\), and those depend on \(X\) and the box \(Q\) alone, not on \(t\), while \(\abs{Y}\) and \(\norm{D_{x}Y}\) have the same maxima over \(Q\) as \(\abs{X}\) and \(\norm{D_{x}X}\). Shrinking \(B\) once and for all so that \(\ee^{L}\delta\) is smaller than the radius of \(B_{2}\), Equation (A.197) puts \(\psi_{t}(B)\) inside \(B_{2}\) for every \(t\). For \(x\in B\) the curve \(s\mapsto\psi_{t-s}(x)\) satisfies \(\dd/\dd s = -X_{t-s}(\psi_{t-s}(x)) = Y_{s}(\cdot)\) and starts at \(\psi_{t}(x)\), so uniqueness (Theorem 9.8) gives \(\eta_{t}(\psi_{t}(x)) = \psi_{0}(x) = x\). Put \(W = \set{y\in B_{2}\mid \eta_{t}(y)\in B}\), open because \(\eta_{t}\) is continuous. Then \(\psi_{t}(B)\subseteq W\); and for \(y\in W\) the same uniqueness argument run backwards gives \(\psi_{t}(\eta_{t}(y)) = y\), so \(W\subseteq\psi_{t}(B)\). Hence \(\psi_{t}:B\longrightarrow W\) is a bijection onto an open set, with inverse \(\eta_{t}\), and both are smooth.
∎Let \(\varphi:I\times U\longrightarrow V\) be smooth, \(I\subseteq\R\) an interval, \(U,V\subseteq\R^{n}\) open, and let \(X_{t}\) be a smooth time-dependent vector field on \(V\) with
Let \(\alpha_{t}\) be a smooth time-dependent \(k\)-form on \(V\). Then
where the Lie derivative of a \(k\)-form has the components
the index \(l\) standing in the \(a\)-th slot. Rests on Definition A.70, Proposition 13.127 and Proposition 7.73.
Derives Lemma A.75. By Equation (A.186),
Differentiate in \(t\) by the product rule. The first factor contributes, by the chain rule Equation (7.52) and Equation (A.201),
For the remaining \(k\) factors, the mixed partial derivatives of \(\varphi\) may be exchanged (Proposition 7.73), so
again by Equation (7.52). The \(a\)-th such term is therefore
which, after renaming the summation indices \(j_{a}\leftrightarrow l\), is the pullback of the \(a\)-th drag term of Equation (A.203). Collecting the \(k+1\) contributions gives Equation (A.202). Finally, taking \(\alpha_{t}=\alpha\) independent of \(t\), \(\varphi_{0}=\id\) and \(X_{t}=X\), the identity at \(t=0\) reads \(\dd\left(\varphi_{t}^{\ast}\alpha\right)/\dd t\big|_{0} =\mathcal{L}_{X}\alpha\) with the right-hand side Equation (A.203): this is the definition Equation (13.267) of the Lie derivative, so Equation (A.203) is indeed the \(k\)-form case of Proposition 13.127 — one transport term and one drag term per index — and no separate verification is needed.
∎The local converse of the Poincaré lemma
Lemma 13.107 states one half of the correspondence: an exact form is closed. The half needed here is the converse, which is false in general (it fails on a punctured plane) and true on a star-shaped set. Differentiable Manifolds, Tensors, and Curvature announces the dependence on topology but does not prove the positive statement; it is proved here.
An open set \(B\subseteq\R^{n}\) is star-shaped about \(x_{0}\in B\) if \(x_{0}+t\left(x-x_{0}\right)\in B\) for every \(x\in B\) and every \(t\in[0,1]\). Every open ball is star-shaped about each of its points. Rests on Definition 6.26.
Let \(B\subseteq\R^{n}\) be open and star-shaped about \(x_{0}\), and let \(\alpha\) be a smooth closed \(k\)-form on \(B\) with \(k\ge1\). Then \(\alpha=\dd\sigma\), where \(\sigma\) is the \((k-1)\)-form with components
In particular \(\sigma(x_{0})=0\), and for \(k=2\) the one-form \(\sigma\) vanishes at \(x_{0}\). Rests on Definition A.76, Lemma A.75, Proposition 13.128 and Lemma A.73.
Derives Lemma A.77. Translating, take \(x_{0}=0\). Put \(\varphi_{t}(x)=tx\) for \(t\in(0,1]\), which maps \(B\) into \(B\) by Definition A.76, and
which is Equation (A.201). Both \(\varphi\) and \(X\) are smooth on \((0,1]\times B\). By Lemma A.75 with \(\alpha_{t}=\alpha\) fixed, and by Cartan's magic formula Equation (13.276) together with \(\dd\alpha=0\),
the last step by Equation (A.188). The components of the form being differentiated are
since \(\left(i_{X_{t}}\alpha\right)_{i_{2}\ldots i_{k}}(y) = t^{-1}y^{j}\alpha_{j\,i_{2}\ldots i_{k}}(y)\) by the interior-product convention of Proposition 13.128, each of the \(k-1\) pullback factors \(\pp_{i}\varphi^{j}_{t}=t\,\delta^{j}_{\ i}\) contributes a \(t\), and \(y=tx\) supplies \(y^{j}/t = x^{j}\). The right-hand side of Equation (A.206) extends continuously to \(t=0\) because \(k\ge1\), and so do its partial derivatives in \(x\), so \(\sigma\) of Equation (A.204) is a well-defined smooth \((k-1)\)-form and, by Lemma A.73, \(\dd\) may be taken under the integral sign. Integrating Equation (A.205) from \(\varepsilon\) to \(1\),
Now \(\varphi_{1}=\id\), so the first term is \(\alpha\); and \(\left(\varphi_{\varepsilon}^{\ast}\alpha\right)_{i_{1}\ldots i_{k}}(x) = \varepsilon^{k}\alpha_{i_{1}\ldots i_{k}}(\varepsilon x)\) tends to \(0\) as \(\varepsilon\to0\), again because \(k\ge1\) and \(\alpha\) is continuous at the origin. Letting \(\varepsilon\to0\) in Equation (A.207), with the interchange of \(\dd\) and the integral justified by Lemma A.73, gives \(\alpha=\dd\sigma\) with \(\sigma\) as in Equation (A.204). Setting \(x=x_{0}\) there makes the factor \((x-x_{0})^{j}\) vanish, so \(\sigma(x_{0})=0\).
∎Moser's homotopy argument
Proof of Theorem A.68. Derives Theorem A.68. Step 1: the pointwise normal form. Let \(n=\dim M\). The bilinear form \(\omega_{P}\) on the \(n\)-dimensional real vector space \(T_{M}(P)\) is antisymmetric and non-degenerate (Definition A.67), so Proposition 5.115 applies: \(n=2m\) is even and there is a basis \(\set{e_{1},\ldots,e_{m},f_{1},\ldots,f_{m}}\) of \(T_{M}(P)\) in which Equation (5.143) holds, that is, in which the matrix of \(\omega_{P}\) is Equation (5.144).
Step 2: a chart adapted at the point. Choose any chart \(x=(x^{1},\ldots,x^{2m})\) around \(P\) with \(x(P)=0\) and with the coordinate vectors \(\pp_{i}\) at \(P\) equal to that basis — possible because the coordinate frame of a chart may be composed with any constant invertible linear map (Definition 13.2), and the change of basis carrying the coordinate frame to \(\set{e_{i},f_{i}}\) is invertible. Let \(B_{0}\) be an open ball in the image of the chart, centred at the origin. Define on \(B_{0}\) the constant-coefficient two-form
all other components vanishing. Its matrix is exactly Equation (5.144), so by Step 1
and \(\dd\omega_{0}=0\) because its components are constants (Definition 13.103).
Step 3: the interpolation. Put \(\tau=\omega-\omega_{0}\) and
Each \(\omega_{t}\) is closed, since \(\omega\) and \(\omega_{0}\) are; and by Equation (A.209), \(\left.\omega_{t}\right|_{P}=\left.\omega\right|_{P}\) for every \(t\), which is non-degenerate. The function
is continuous on \([0,1]\times B_{0}\) — a polynomial in the components, which are continuous — and \(D(t,0)\ne0\) for every \(t\). For each \(t\in[0,1]\) continuity therefore supplies an open interval \(I_{t}\ni t\) and a ball \(B^{(t)}\ni 0\) on which \(D\) does not vanish; the intervals \(I_{t}\) cover the compact \([0,1]\) (Theorem 6.11), finitely many of them suffice, and the intersection \(B_{1}\) of the corresponding finitely many balls is an open ball about the origin on which
Step 4: the primitive. \(B_{1}\) is star-shaped about the origin (Definition A.76) and \(\tau\) is a closed two-form on it, so Lemma A.77 gives a smooth one-form \(\sigma\) on \(B_{1}\) with
The vanishing of \(\sigma\) at the origin — which is Equation (A.209) feeding through the explicit homotopy formula — is what will pin the point \(P\) and is the reason the argument is local rather than global.
Step 5: the vector field. Let \(W^{ij}(t,x)\) be the inverse of the matrix \(\left(\omega_{t}\right)_{ij}(x)\), which exists on \([0,1]\times B_{1}\) by Equation (A.211) and whose entries are smooth there, being rational functions of the components with non-vanishing denominator \(D\). Define the time-dependent vector field
the equivalence because \(\left(i_{X}\omega_{t}\right)_{i}=X^{j}\left(\omega_{t}\right)_{ji}\), and contracting \(X^{j}(\omega_{t})_{ji}=-\sigma_{i}\) with \(W^{ik}\) returns Equation (A.213). It is smooth in \((t,x)\), and by Equation (A.212)
Step 6: the flow. By Equation (A.214), Theorem A.74 applies: there is a ball \(B_{2}\ni0\) and a smooth family \(\psi_{t}:B_{2}\longrightarrow B_{1}\), \(t\in[0,1]\), of diffeomorphisms onto open sets, with \(\psi_{0}=\id\), \(\pp_{t}\psi_{t}=X_{t}\circ\psi_{t}\) and \(\psi_{t}(0)=0\).
Step 7: the homotopy equation. By Lemma A.75, Cartan's magic formula Equation (13.276), \(\dd\omega_{t}=0\), \(\pp_{t}\omega_{t}=\tau\) and Equation (A.213),
where the last line used Equation (A.212). Hence \(\psi_{t}^{\ast}\omega_{t}\) is independent of \(t\) on \(B_{2}\), and at \(t=0\) it equals \(\psi_{0}^{\ast}\omega_{0}=\omega_{0}\). Taking \(t=1\),
Step 8: reading off the chart. Let \(W=\psi_{1}(B_{2})\), an open neighbourhood of \(0\), and define on it
a chart of \(M\) around \(P\) with \(u(P)=0\), since \(\psi_{1}\) is a diffeomorphism and \(\psi_{1}(0)=0\). Applying \(\left(\psi_{1}^{-1}\right)^{\ast}\) to Equation (A.216) and using Equation (A.189), \(\omega = \left(\psi_{1}^{-1}\right)^{\ast}\omega_{0}\) on \(W\). By Equation (A.190), \(\left(\psi_{1}^{-1}\right)^{\ast}\dd x^{a} = \dd\left(x^{a}\circ\psi_{1}^{-1}\right) = \dd u^{a}\), and by Equation (A.187) the pullback distributes over the wedge, so Equation (A.208) becomes
which is Equation (A.184) with \(q^{i}=u^{i}\) and \(p_{i}=u^{m+i}\).
∎Reading the result
Both hypotheses of Definition A.67 are load-bearing and they enter at different places. Non-degeneracy is used twice: at one point in Step 1, to get the linear normal form from Proposition 5.115, and on a neighbourhood in Step 5, to solve Equation (A.213) for \(X_{t}\) — without it the homotopy equation cannot be solved even though it is a linear algebraic equation. Closedness is used twice as well: in Step 4, where \(\dd\tau=0\) is precisely the hypothesis of Lemma A.77, and in Step 7, where \(i_{X_{t}}\dd\omega_{t}\) is dropped from Cartan's formula. A non-degenerate two-form that is not closed has no canonical chart, and the obstruction is exactly the three-form \(\dd\omega\), which is a genuine local invariant.
Remark 5.116 records that on the phase space of a mechanical system the value \(\omega(u,v)\) carries the SI unit \(\mathrm{J}\,\mathrm{s}\), that of action and of \(\hbar\). Nothing in the proof above disturbs this: the chart of Step 2 is obtained by a constant linear change of frame, which distributes the unit between the two halves of the canonical pair, and every later step is an identity between two-forms and so is unit-homogeneous. In the canonical chart of Equation (A.184) the product \(q^{i}p_{i}\) carries \(\mathrm{J}\,\mathrm{s}\) while the individual factors do not have to: for a point particle of mass \(m\) moving in three-dimensional space, \(q^{i}\) is a Cartesian position in \(\mathrm{m}\) and \(p_{i}\) the conjugate momentum in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), so each summand \(\dd q^{i}\wedge\dd p_{i}\) carries \(\mathrm{kg}\,\mathrm{m}^{2}/\mathrm{s}=\mathrm{J}\,\mathrm{s}\), as it must.
Take \(M = T^{\ast}\R^{3}\), the phase space of a single particle moving in three-dimensional space: \(m=3\), \(\dim M = 6\), global coordinates \((q^{1},q^{2},q^{3},p_{1},p_{2},p_{3})\) and \(\omega=\sum_{i=1}^{3}\dd q^{i}\wedge\dd p_{i}\). Here the canonical chart is global and Theorem A.68 says nothing new. Its force appears when the same phase space is written in another coordinate system — spherical coordinates and their conjugate momenta, or the action–angle variables of an integrable system: the theorem guarantees that the transformed \(\omega\) can always be brought back to Equation (A.184) near any point, so that no computation in one such system can disagree with a computation in another about a local question. Rests on Theorem A.68 and Remark A.79.
The theorem is Darboux's; the proof given here is not his, but the homotopy argument that interpolates between two forms and integrates a time-dependent vector field, introduced by Moser for volume forms [Moser:1965] and adapted to the symplectic case shortly afterwards. Darboux's own memoir is not in this treatise's bibliography, so that half of the attribution is made in words; nothing in The pullback of a $k$-form, Flows of time-dependent vector fields, The local converse of the Poincaré lemma and Moser's homotopy argument rests on either, every step above having been carried out from the material of Linear Algebra and Representation Theory, Real Analysis, Ordinary Differential Equations and Sturm–Liouville Theory and Differentiable Manifolds, Tensors, and Curvature.
Darboux's Theorem: Local Canonical Coordinates discharges the derivation owed at Remark 5.116 of Linear Algebra and Representation Theory, where Proposition 5.115 establishes the pointwise half of the statement — a non-degenerate antisymmetric form on a single vector space has no invariant but its dimension — and Darboux's theorem is named as the geometric counterpart that does not follow from it. The gap between the two is exactly Steps 3 to 7 above: the linear normal form can be achieved at every point separately by a basis that varies from point to point, and it is closedness of \(\omega\), through the Poincaré lemma and the homotopy equation, that lets those bases be chosen to fit together into a chart. The three tools built along the way — the pullback of a \(k\)-form (Definition A.70), the smoothness of a flow (Theorem A.74) and the converse Poincaré lemma (Lemma A.77) — belong with the differential-form machinery of Section 13.6, Section 13.6.5 and Section 13.9 and are used again in The General Stokes Theorem for Differential Forms.
The Korn–Lichtenstein Theorem and the Elliptic Canonical Form
This appendix discharges the obligation recorded in Remark 10.14 of Partial Differential Equations: the elliptic branch of Proposition 10.13, the reduction of a second-order operator in two independent variables to the canonical form \(u_{\xi\xi}+u_{\eta\eta}\), under coefficients that are merely \(C^{1}\) and therefore only locally Hölder. The chapter's argument constructs the required complex first integral by complexifying the characteristic equation, which needs the coefficients to be real analytic; nothing in real ordinary differential equation theory replaces that step, because what is really being asked for is not a first integral of an ordinary equation at all but a solution of a first-order elliptic system in two variables — a Beltrami equation.
Four things are done below. First the reduction is made exact: the canonical form exists at a point if and only if a Beltrami equation with a coefficient of modulus bounded away from \(1\) has a local solution with non-vanishing Jacobian, and each of the equivalent formulations is derived from the previous one with nothing left implicit (From the canonical form to a Beltrami equation). Then the problem is normalized so that the Beltrami coefficient is as small as one pleases, in the Hölder norm and not merely in the supremum norm (Normalizing the Beltrami coefficient); this is where the Hölder hypothesis is spent, and it is what turns an existence problem into a convergent series. Then the singular-integral machinery is set out and the one deep estimate is stated precisely as a quoted theorem, with a remark saying exactly what is assumed and why this treatise does not prove it (The Cauchy and Beurling transforms). Finally the solution is constructed as an explicitly convergent Neumann series and the Jacobian is bounded below by hand (The construction and the Jacobian), which completes the proof.
Throughout, \(x\) and \(y\) are the two independent variables of the partial differential equation — not two spatial dimensions — and \(z = x+\ii y\), \(\bar z = x-\ii y\), with
so that \(\pp_{x} = \pp_{z}+\pp_{\bar z}\) and \(\pp_{y} = \ii\bigl(\pp_{z}-\pp_{\bar z}\bigr)\), and \(\nabla^{2} = 4\pp_{z}\pp_{\bar z}\). We write \(D_{\rho} = \set{z \in \C \mid \abs{z} < \rho}\), and for \(0 < \alpha < 1\)
\(C^{\alpha}\) denoting the functions with \([f]_{\alpha} < \infty\) and \(C^{k,\alpha}\) those whose derivatives up to order \(k\) lie in \(C^{\alpha}\).
Statement
Let \(L u = a\,u_{xx} + 2b\,u_{xy} + c\,u_{yy} + \cdots\) be elliptic on a neighbourhood of \(P_{0}\), that is \(\Delta = b^{2}-ac < 0\) there, with \(a,b,c\) of class \(C^{\alpha}\) for some \(\alpha \in (0,1)\) — in particular whenever they are \(C^{1}\). Then there are a neighbourhood \(V\) of \(P_{0}\) and a map \((\xi,\eta) : V \longrightarrow \R^{2}\) of class \(C^{1,\alpha}\) whose Jacobian \(J_{0} = \xi_{x}\eta_{y}-\xi_{y}\eta_{x}\) is bounded away from \(0\) on \(V\), such that the transformed principal coefficients of Equation (10.12) satisfy
so that in the coordinates \((\xi,\eta)\) the principal part of \(L\) is \(\delta J_{0}\bigl(\pp_{\xi}^{2}+\pp_{\eta}^{2}\bigr)\), a positive multiple of the Laplacian. If in addition \(a,b,c \in C^{1,\alpha}\), the map is of class \(C^{2,\alpha}\) and the reduction is classical: the transformed operator is a genuine second-order operator with continuous coefficients, of the third form in Equation (10.10). Rests on Proposition 10.13, Definition 10.6 and Equation (10.6).
Two remarks fix what Theorem A.83 does and does not claim, and the second is the reason the theorem is stated in two halves.
The congruence Equation (10.6) that carries \(A\) to \(\tilde A\) involves the Jacobian \(J\) and nothing else, so it is meaningful for any \(C^{1}\) change of coordinates and is a pointwise statement of linear algebra. The extra term of Equation (10.7), by contrast, carries the second derivatives \(\pp_{i}\pp_{j}y^{a}\) of the change; a \(C^{1,\alpha}\) change does not have them, so for coefficients that are merely \(C^{\alpha}\) the transformed equation is not literally an equation with continuous first-order coefficients — the first-order remainder exists only as a distribution. Nothing is wrong with the reduction; what is wrong is the naive reading of what “the equation becomes \(u_{\xi\xi}+u_{\eta\eta}+\cdots\)” asserts. This is why the theorem separates the two statements, and why the second half buys the classical form with one extra derivative on the coefficients. The chapter's sentence that the conclusion of Proposition 10.13 “does hold at \(C^{1}\)” is correct in the first sense and should be read in it.
There is no loss in assuming \(a > 0\). Ellipticity gives \(ac > b^{2} \ge 0\), so \(a\) and \(c\) are nonzero and of the same sign; replacing \(L\) by \(-L\), which changes no solution of \(Lu=f\) except by the sign of \(f\), makes both positive. The matrix \(A\) of Equation (10.11) is then positive definite, and \(\delta = \sqrt{ac-b^{2}} = \sqrt{\det A} > 0\) is continuous and bounded away from \(0\) near \(P_{0}\).
From the canonical form to a Beltrami equation
Let \(\xi,\eta\) be real \(C^{1}\) functions near \(P_{0}\) and set \(\zeta = \xi + \ii\eta\). Then, with \(\tilde a,\tilde b,\tilde c\) the transformed principal coefficients of Equation (10.12),
Hence \(\tilde a = \tilde c\) and \(\tilde b = 0\) hold at a point if and only if \(\zeta\) satisfies the characteristic equation Equation (10.8) there, with the complex \(\nabla\zeta\) in place of \(\nabla\phi\). Rests on Equations (10.8) and (10.12).
Derives Proposition A.86. Expand, using \(\zeta_{x} = \xi_{x}+\ii\eta_{x}\) and \(\zeta_{y} = \xi_{y}+\ii\eta_{y}\):
The real parts sum to \(\bigl(a\xi_{x}^{2}+2b\xi_{x}\xi_{y}+c\xi_{y}^{2}\bigr) -\bigl(a\eta_{x}^{2}+2b\eta_{x}\eta_{y}+c\eta_{y}^{2}\bigr) = \tilde a - \tilde c\) by Equation (10.12), and the imaginary parts sum to twice the polarisation \(a\xi_{x}\eta_{x}+b(\xi_{x}\eta_{y}+\eta_{x}\xi_{y})+c\xi_{y}\eta_{y} = \tilde b\).
∎Assume \(a>0\) and \(\Delta<0\), and write \(\delta = \sqrt{ac-b^{2}}>0\). A \(C^{1}\) pair \((\xi,\eta)\) satisfies the equation of Proposition A.86 with the branch \(a\zeta_{x}+(b+\ii\delta)\zeta_{y}=0\) if and only if
For such a pair,
so \(J_{0} > 0\) exactly where \(\nabla\xi \neq 0\), and Equation (A.220) holds there. Rests on Proposition A.86 and Equation (10.11).
Derives Proposition A.87. The quadratic \(a t^{2}+2bt+c\) has the two roots \(t_{\pm} = \bigl(-b\pm\ii\delta\bigr)/a\), so \(a\zeta_{x}^{2}+2b\zeta_{x}\zeta_{y}+c\zeta_{y}^{2} = a\bigl(\zeta_{x}-t_{+}\zeta_{y}\bigr) \bigl(\zeta_{x}-t_{-}\zeta_{y}\bigr)\) and the equation of Proposition A.86 holds if and only if one of the two factors vanishes. Take the factor \(\zeta_{x}-t_{-}\zeta_{y} = 0\), i.e.\ \(a\zeta_{x} + \bigl(b+\ii\delta\bigr)\zeta_{y} = 0\) (the other branch is the complex conjugate statement and produces \(\bar\zeta\)). Separating real and imaginary parts of \(a(\xi_{x}+\ii\eta_{x}) + (b+\ii\delta)(\xi_{y}+\ii\eta_{y}) = 0\),
The first is the second half of Equation (A.222). Substituting it into the second and using \(\delta^{2} = ac-b^{2}\), so that \(b^{2}+\delta^{2} = ac\),
and dividing by \(a \neq 0\) gives the first half. The steps are reversible, which is the “only if”.
For Equation (A.223), substitute Equation (A.222):
which is the displayed quadratic form. That form is \(\vect{v}\transpose A\,\vect{v}/\delta\) with \(\vect{v} = (\xi_{x},\xi_{y})\) and \(A\) the positive definite matrix of Equation (10.11) (Remark A.85), so it is positive whenever \(\vect{v} \neq 0\) and zero otherwise. Finally Proposition A.86 gives \(\tilde a = \tilde c\) and \(\tilde b = 0\), and Equation (A.223) identifies \(\tilde a = \delta J_{0}\), which is Equation (A.220).
∎Eliminating \(\eta\) from Equation (A.222) by the equality of mixed partials \(\pp_{y}\eta_{x} = \pp_{x}\eta_{y}\) gives, for \(\xi\) alone, the second-order equation in divergence form
whose principal part is again \(L\)'s. That is the exact sense in which the elliptic branch of Proposition 10.13 is circular if attacked head-on: putting an elliptic operator in canonical form is equivalent to solving an elliptic equation with the same principal part. The hyperbolic and parabolic branches escape this because there the two characteristic directions are real, so the corresponding first integrals come from Theorem 9.8 applied to a real ordinary differential equation, which needs no more than continuity and a Lipschitz condition. The complex route below breaks the circle by solving Equation (A.225) not by elliptic theory but by inverting a constant-coefficient operator, \(\pp_{\bar z}\), and treating everything else as a perturbation.
With \(a>0\), \(\Delta<0\) and \(\delta = \sqrt{ac-b^{2}}\), the branch \(a\zeta_{x}+(b+\ii\delta)\zeta_{y}=0\) of Proposition A.87 is exactly the Beltrami equation
and its coefficient satisfies
with \(\mu = 0\) exactly where \(A\) is a positive multiple of the identity. Moreover the Jacobian of \(z \longmapsto \zeta\) is
Rests on Proposition A.87 and Equation (A.218).
Derives Proposition A.89. Substitute \(\zeta_{x} = \pp_{z}\zeta+\pp_{\bar z}\zeta\) and \(\zeta_{y} = \ii\bigl(\pp_{z}\zeta-\pp_{\bar z}\zeta\bigr)\) from Equation (A.218) into \(a\zeta_{x}+(b+\ii\delta)\zeta_{y}=0\) and use \((b+\ii\delta)\ii = -\delta+\ii b\):
The coefficient of \(\pp_{\bar z}\zeta\) has modulus squared \((a+\delta)^{2}+b^{2} > 0\) and never vanishes, so the equation may be solved for \(\pp_{\bar z}\zeta\), giving Equation (A.226). Its modulus is Equation (A.227), and \((a-\delta)^{2} < (a+\delta)^{2}\) because \(a\) and \(\delta\) are both positive; equality of the numerator with \(0\) requires \(b = 0\) and \(a = \delta = \sqrt{ac}\), i.e. \(a = c\) and \(b=0\).
For Equation (A.228), write \(A_{1} = \zeta_{x}\) and \(B_{1} = \ii\zeta_{y}\), so that \(\pp_{z}\zeta = \tfrac12(A_{1}-B_{1})\) and \(\pp_{\bar z}\zeta = \tfrac12(A_{1}+B_{1})\). Then
and \(\bar\zeta_{x}\zeta_{y} = (\xi_{x}-\ii\eta_{x})(\xi_{y}+\ii\eta_{y})\) has imaginary part \(\xi_{x}\eta_{y}-\eta_{x}\xi_{y} = J_{0}\). Inserting Equation (A.226) gives the second equality.
∎Proposition A.89 reduces Theorem A.83 to a single question: does Equation (A.226) have a local \(C^{1,\alpha}\) solution with \(\pp_{z}\zeta \neq 0\)? Note that Equation (A.228) has already disposed of the Jacobian: because \(\abs{\mu}\) is bounded away from \(1\), a solution whose \(\pp_{z}\zeta\) does not vanish automatically has \(J_{0}\) bounded away from \(0\), and no separate argument is needed.
Normalizing the Beltrami coefficient
There is an invertible real linear change of the independent variables after which \(A(P_{0}) = \identity\), \(\delta(P_{0}) = 1\) and \(\mu(P_{0}) = 0\), the transformed coefficients being again of class \(C^{\alpha}\) and again elliptic. Rests on Proposition A.89 and Theorem 10.8.
Derives Lemma A.90. \(A(P_{0})\) is symmetric and positive definite (Remark A.85), so it factors as \(A(P_{0}) = S\,S\transpose\) with \(S\) real and invertible — take \(S = O\Lambda^{1/2}\) with \(A(P_{0}) = O\Lambda O\transpose\) its orthogonal diagonalization, the eigenvalues in \(\Lambda\) being positive. Change variables by \(\hat{\vect{x}} = S^{-1}\vect{x}\), whose Jacobian is the constant matrix \(J = S^{-1}\); by Equation (10.6) the new principal matrix is
Congruence with a constant matrix preserves Hölder continuity and, by Theorem 10.8, ellipticity. At \(P_{0}\) one then has \(\hat a = \hat c = 1\), \(\hat b = 0\), \(\hat\delta = 1\), and Equation (A.226) gives \(\mu(P_{0}) = -(1-1+0)/(1+1-0) = 0\). Since a composition of two changes of coordinates with non-vanishing Jacobians is again one, proving Theorem A.83 for the normalized problem proves it in general.
∎Assume the normalization of Lemma A.90 and identify \(P_{0}\) with \(0 \in \C\). Fix once and for all a cutoff \(\chi \in C^{\infty}(\C)\) with \(0 \le \chi \le 1\), \(\chi = 1\) on \(\overline{D_{1}}\) and \(\chi = 0\) off \(D_{2}\). For \(r>0\) put
Then \(\nu_{r} \in C^{\alpha}(\C)\) vanishes off \(\overline{D_{2}}\), coincides with \(z \longmapsto \mu(rz)\) on \(\overline{D_{1}}\), satisfies \(\norm{\nu_{r}}_{\infty} \le \sup\abs{\mu} < 1\), and
with \(K\) depending only on \(\chi\) and \(\alpha\). If moreover \(a,b,c \in C^{1,\alpha}\), the same holds with the \(C^{1,\alpha}\) norm in place of the \(C^{\alpha}\) norm and \(r\) in place of \(r^{\alpha}\). Rests on Lemma A.90 and Equation (A.219).
Derives Lemma A.91. \(\mu\) is a rational function of \(a,b,\delta\) with non-vanishing denominator, hence of class \(C^{\alpha}\) on a neighbourhood of \(0\); shrink that neighbourhood to a disc \(D_{\rho}\) and take \(r < \rho/2\) so that \(rz \in D_{\rho}\) for \(z \in D_{2}\). Write \(\mu^{(r)}(z) = \mu(rz)\). Since \(\mu(0)=0\) by Lemma A.90,
and directly from Equation (A.219), \([\mu^{(r)}]_{\alpha,D_{2}} = r^{\alpha}[\mu]_{\alpha}\). For the product with the cutoff, use \(\abs{\chi f(z)-\chi f(w)} \le \abs{\chi(z)}\abs{f(z)-f(w)} + \abs{f(w)}\abs{\chi(z)-\chi(w)}\), so that
which with Equation (A.232) gives \([\nu_{r}]_{\alpha} \le r^{\alpha}[\mu]_{\alpha} + (2r)^{\alpha}[\mu]_{\alpha}[\chi]_{\alpha}\). Adding \(\norm{\nu_{r}}_{\infty} \le (2r)^{\alpha}[\mu]_{\alpha}\) gives Equation (A.231) with \(K = 1+2^{\alpha}\bigl(1+[\chi]_{\alpha}\bigr)\). The bound \(\norm{\nu_{r}}_{\infty} \le \sup\abs{\mu} < 1\) is Equation (A.227) and \(0\le\chi\le1\). Under the stronger hypothesis \(a,b,c \in C^{1,\alpha}\) one has \(\mu \in C^{1,\alpha}\) and \(\pp\mu^{(r)} = r\,(\pp\mu)(rz)\), so every seminorm entering the \(C^{1,\alpha}\) norm of \(\mu^{(r)}\) carries at least one factor \(r\), and the same product estimate applies.
∎Lemma A.91 is the whole use made of the Hölder continuity of the coefficients, and it is worth seeing why mere continuity would not do. Continuity of \(\mu\) at \(P_{0}\) makes \(\norm{\nu_{r}}_{\infty}\) small, which is enough for an \(L^{p}\) theory but not for a Hölder one; what the exponent \(\alpha\) buys is the factor \(r^{\alpha}\) in Equation (A.231), which makes the seminorm small as well and thereby turns the Neumann series of The construction and the Jacobian into a convergent series in the very space where the solution's regularity is measured. A zooming argument of this kind is available only because the Beltrami equation is invariant under \(z \longmapsto rz\): if \(\hat\zeta\) solves \(\pp_{\bar z}\hat\zeta = \mu^{(r)}\,\pp_{z}\hat\zeta\) on \(D_{1}\), then \(\zeta(w) = \hat\zeta(w/r)\) solves Equation (A.226) on \(D_{r}\), both derivatives picking up the same factor \(1/r\), and the Jacobian is multiplied by the positive number \(r^{-2}\).
The Cauchy and Beurling transforms
For \(f\) continuous and vanishing outside a compact set put
\(\dd A\) being the area element of the plane and the principal value meaning the limit of the integral over \(\abs{z-w} > \varepsilon\) as \(\varepsilon \to 0^{+}\). \(P\) is the Cauchy transform and \(T\) the Beurling transform. Rests on Equation (A.218).
The integral defining \(Pf\) converges absolutely, the kernel \(\abs{z-w}^{-1}\) being integrable in the plane; the one defining \(Tf\) does not, and the principal value is not a notational convenience but the entire difficulty. The next lemma explains why \(P\) is the right object, and is the one part of the analytic package this treatise can prove outright, because it is the two-dimensional twin of Proposition 10.67.
In the sense of distributions on \(\C \cong \R^{2}\),
Consequently \(Pf = \bigl(\pi z\bigr)^{-1} * f\) satisfies \(\pp_{\bar z}Pf = f\) distributionally. Rests on Proposition 10.67 and Equation (A.218).
Derives Lemma A.94. Let \(G(z) = (2\pi)^{-1}\ln\abs{z}\), which is harmonic on \(\C \setminus \set{0}\) because \(\nabla^{2}\ln r = r^{-1}\pp_{r}(r\,\pp_{r}\ln r) = r^{-1}\pp_{r}(1)=0\) in plane polar coordinates. Let \(\varphi\) be a test function. Since \(\ln\abs{z}\) is integrable near the origin, \(\int G\nabla^{2}\varphi\,\dd A = \lim_{\varepsilon\to0^{+}}\int_{r>\varepsilon}G\nabla^{2}\varphi\,\dd A\). On \(r>\varepsilon\) apply Green's second identity Equation (10.41) with \(p=1\), \(u=\varphi\), \(v=G\); the term \(\varphi\nabla^{2}G\) vanishes, and the only boundary is the circle \(r=\varepsilon\), whose outward normal for the region \(r>\varepsilon\) points at the origin, so \(\pp_{\nu} = -\pp_{r}\) there. Hence
using \(\pp_{r}G = (2\pi r)^{-1}\). The circle has length \(2\pi\varepsilon\), so the first term tends to \(-\varphi(0)\) by continuity of \(\varphi\) and the second is bounded by \(\varepsilon\abs{\ln\varepsilon}\sup\abs{\nabla\varphi} \to 0\). Thus \(\int G\nabla^{2}\varphi\,\dd A = \varphi(0)\), which is the first half of Equation (A.235). For the second, note \(\nabla^{2} = 4\pp_{z}\pp_{\bar z}\) and \(\pp_{z}\ln\abs{z} = \pp_{z}\tfrac12\ln\bigl(z\bar z\bigr) = 1/(2z)\), so \(\delta = 4\pp_{\bar z}\pp_{z}G = 4\pp_{\bar z}\bigl(1/(4\pi z)\bigr) = \pp_{\bar z}\bigl(1/(\pi z)\bigr)\). Finally Equation (A.234) is the convolution of \(f\) with \(1/(\pi z)\), and differentiating a convolution moves the derivative onto either factor.
∎Let \(0 < \alpha < 1\) and let \(f \in C^{\alpha}(\C)\) vanish outside \(\overline{D_{2}}\). Then:
-
the principal value defining \(Tf\) in Equation (A.234) exists at every point of \(\C\);
-
\(Pf\) is continuously differentiable on \(\C\), with
\begin{equation}\tag{A.236} \pp_{\bar z}Pf = f\ec \qquad \pp_{z}Pf = Tf \end{equation}holding classically;
-
\(Tf \in C^{\alpha}(\C)\) and
\begin{equation}\tag{A.237} \norm{Tf}_{\infty} + [Tf]_{\alpha} \le C_{\alpha}\,[f]_{\alpha}\ec \end{equation}with \(C_{\alpha}\) depending only on \(\alpha\);
-
if in addition \(f \in C^{k,\alpha}\) for an integer \(k \ge 1\), then \(Tf \in C^{k,\alpha}\) and \(\norm{Tf}_{C^{k,\alpha}} \le C_{k,\alpha}\norm{f}_{C^{k,\alpha}}\).
Theorem A.95 is the only input of this section that is not proved, and it is worth being exact about its three parts. Part (ii) is half proved above: \(\pp_{\bar z}Pf = f\) is Lemma A.94, and the Hölder hypothesis upgrades it from a distributional to a classical identity. Parts (i) and the second half of (ii) — that the singular integral converges and that it is the \(z\)-derivative of \(Pf\) — are the statement that differentiating the Newtonian-type potential \(Pf\) a second time is legitimate in the principal-value sense; the same question arises for the Newtonian potential of Proposition 10.67, where the chapter needed only one derivative and could avoid it. Part (iii) is the substantial content: it says the Beurling transform, a singular integral operator whose kernel \((z-w)^{-2}\) is exactly at the critical homogeneity for the plane, is bounded on Hölder spaces. That is the Korn–Lichtenstein inequality, proved for such kernels by Korn in 1914 and Lichtenstein in 1916 and placed in its general setting by the Calderón–Zygmund theory of singular integrals [Calderon:1952]; the standard modern account of the Beltrami equation built on it is Vekua's. The two original memoirs and Vekua's monograph are not entries of this bibliography, so those attributions are made in words; what the treatise does contain and does cite for the classical, real-analytic route is Courant and Hilbert [Courant:1962].
Why the estimate is not proved here is not a matter of length. Its proof requires the Calderón–Zygmund decomposition of a function relative to a cube, or an equivalent covering argument, and both rest on Lebesgue measure and on the maximal function; Real Analysis develops the Riemann integral only, and the Lebesgue theory these arguments belong to is not built anywhere in this book. Everything else in this section — the reduction of From the canonical form to a Beltrami equation, the normalization of Normalizing the Beltrami coefficient, the convergent series and the Jacobian bound of The construction and the Jacobian — is elementary once Equation (A.237) is granted, and is carried out in full.
Two features of Equation (A.237) are used below and are worth isolating. First, the right-hand side involves only the seminorm \([f]_{\alpha}\), which is legitimate because \(f\) vanishes off \(\overline{D_{2}}\): every point of \(\overline{D_{2}}\) lies within distance \(4\) of a point where \(f\) vanishes, so \(\norm{f}_{\infty} \le 4^{\alpha}[f]_{\alpha}\), and \([\,\cdot\,]_{\alpha}\) is therefore a genuine norm on the space of such \(f\), complete because a \([\,\cdot\,]_{\alpha}\)-Cauchy sequence converges uniformly and its limit inherits the seminorm bound. Second, the left-hand side controls \(\norm{Tf}_{\infty}\) as well as \([Tf]_{\alpha}\), and the supremum bound is what will be needed to keep \(\pp_{z}\zeta\) away from zero.
The construction and the Jacobian
Let \(\nu \in C^{\alpha}(\C)\) vanish off \(\overline{D_{2}}\) and set
with \(C_{\alpha}\) the constant of Equation (A.237). If \(\varepsilon_{\nu} < 1\) then the series
converges in the space \(X\) of \(C^{\alpha}\) functions vanishing off \(\overline{D_{2}}\), and its sum is the unique element of \(X\) with
Rests on Theorem A.95 and Equation (A.219).
Derives Proposition A.98. Each \(h_{n}\) lies in \(X\): it is a product one of whose factors is \(\nu\), which vanishes off \(\overline{D_{2}}\), and it is Hölder because \(Th_{n-1}\) is, by Equation (A.237). By the product estimate Equation (A.233) applied to \(\nu\,Tg\) and then Equation (A.237),
for every \(g \in X\). Hence \([h_{n}]_{\alpha} \le \varepsilon_{\nu}^{\,n}[\nu]_{\alpha}\) by induction, and the series Equation (A.239) is dominated term by term by a geometric series of ratio \(\varepsilon_{\nu}<1\). Since \([\,\cdot\,]_{\alpha}\) is a complete norm on \(X\) (Remark A.97), the series converges there, with the bound stated in Equation (A.240); the same domination and Lemma 9.6 give uniform convergence, so the sum may be rearranged and \(T\) applied term by term, \(T\) being bounded. Then
which is Equation (A.240). Uniqueness: if \(h,h'\in X\) both satisfy it, then \([h-h']_{\alpha} = [\nu T(h-h')]_{\alpha} \le \varepsilon_{\nu}[h-h']_{\alpha}\) by Equation (A.241), and \(\varepsilon_{\nu}<1\) forces \([h-h']_{\alpha}=0\), hence \(h=h'\) since both vanish off a compact set.
∎In the situation of Proposition A.98, suppose in addition
Then \(\zeta(z) = z + Ph(z)\) is of class \(C^{1,\alpha}\) on \(\C\), solves \(\pp_{\bar z}\zeta = \nu\,\pp_{z}\zeta\) there, and satisfies
Rests on Proposition A.98, Theorem A.95 and Equation (A.228).
Derives Proposition A.99. By part (ii) of Theorem A.95 applied to \(h \in X\),
both continuous and Hölder by part (iii), so \(\zeta \in C^{1,\alpha}\). The Beltrami equation \(\pp_{\bar z}\zeta = \nu\,\pp_{z}\zeta\) reads, after Equation (A.244), precisely \(h = \nu(1+Th)\), which is Equation (A.240). For the Jacobian, Equation (A.237) and Equation (A.240) give
by Equation (A.242), so \(\abs{\pp_{z}\zeta} = \abs{1+Th} \ge 1 - \tfrac12 = \tfrac12\) everywhere. Equation (A.228) then yields the second half of Equation (A.243), and \(\norm{\nu}_{\infty}<1\) makes it strictly positive.
∎Proof of Theorem A.83. Derives Theorem A.83. By Remark A.85 assume \(a>0\), and by Lemma A.90 assume further that \(A(P_{0}) = \identity\) and \(\mu(P_{0})=0\), \(P_{0}\) being the origin of \(\C\); both reductions are linear changes of the independent variables with constant non-vanishing Jacobian, and composing them with the map constructed below gives the map of the theorem.
Let \(\nu_{r}\) be the cutoff coefficient of Lemma A.91. By Equation (A.231) both \(\varepsilon_{\nu_{r}}\) of Equation (A.238) and the left-hand side of Equation (A.242) are bounded by \(C_{\alpha}K r^{\alpha}[\mu]_{\alpha}\) times a constant, so choosing \(r\) small enough makes \(\varepsilon_{\nu_{r}} \le \tfrac12\) and Equation (A.242) hold. Fix such an \(r\). Propositions A.98 and A.99 then produce \(\hat\zeta \in C^{1,\alpha}(\C)\) with \(\pp_{\bar z}\hat\zeta = \nu_{r}\pp_{z}\hat\zeta\) and Jacobian bounded below by a positive constant.
On \(\overline{D_{1}}\) one has \(\nu_{r}(z) = \mu(rz)\), so \(\zeta(w) = \hat\zeta(w/r)\) satisfies Equation (A.226) on \(D_{r}\) by the scaling identity of Remark A.92, with Jacobian multiplied by \(r^{-2}>0\) and hence still bounded away from \(0\). Set \(V = D_{r}\) and \((\xi,\eta) = (\Re\zeta,\Im\zeta)\), which is \(C^{1,\alpha}\). By Proposition A.89 the pair satisfies the branch \(a\zeta_{x}+(b+\ii\delta)\zeta_{y}=0\) of the complex characteristic equation, so by Propositions A.86 and A.87 the transformed coefficients obey \(\tilde b = 0\) and \(\tilde a = \tilde c = \delta J_{0} > 0\), which is Equation (A.220). The transformed principal part is therefore \(\tilde a\,u_{\xi\xi}+\tilde c\,u_{\eta\eta} = \delta J_{0}\bigl(u_{\xi\xi}+u_{\eta\eta}\bigr)\), as claimed.
For the last sentence of the theorem, suppose \(a,b,c\in C^{1,\alpha}\). Then \(\mu \in C^{1,\alpha}\), and by the last clause of Lemma A.91 the quantity \(\norm{\nu_{r}}_{C^{1,\alpha}}\) is \(O(r)\). Running Proposition A.98 in the space of \(C^{1,\alpha}\) functions vanishing off \(\overline{D_{2}}\) — legitimate because part (iv) of Theorem A.95 bounds \(T\) there, and because the product estimate Equation (A.233) has an evident \(C^{1,\alpha}\) counterpart obtained by applying it to \(f\) and to each \(\pp f\) — gives \(h \in C^{1,\alpha}\), hence \(\zeta \in C^{2,\alpha}\) by Equation (A.244). The change of coordinates then has continuous second derivatives, the extra term of Equation (10.7) is an honest continuous first-order coefficient, and the transformed equation is the third canonical form of Equation (10.10) in the classical sense.
∎For real-analytic coefficients the chapter's own construction is available and is much cheaper: complexify \(y\), integrate the characteristic ordinary differential equation Equation (10.13) along the complex branch, and take real and imaginary parts. That argument produces an analytic \(\zeta\) and needs nothing beyond existence for an analytic ordinary differential equation — which is the Cauchy–Kovalevskaya theorem Theorem 10.22 in its simplest instance. What it cannot do is survive the loss of analyticity, because a complexified non-analytic coefficient has no meaning at all. The route taken here never complexifies the coefficients: it complexifies only the unknown, which turns one real second-order equation into one complex first-order equation, and then solves that by inverting the constant-coefficient operator \(\pp_{\bar z}\) and treating the variable-coefficient part as a small perturbation. The price is Theorem A.95, and the gain is that \(C^{1}\) coefficients suffice — which matters because the coefficients of a physical equation are material properties, measured and interpolated, and are never analytic for any reason a physicist could give.
The Korn–Lichtenstein Theorem and the Elliptic Canonical Form discharges the obligation stated in Remark 10.14 and completes the elliptic branch of Proposition 10.13 in Partial Differential Equations. Its consequence for the chapter is that the trichotomy of Equation (10.10) is a statement about all elliptic operators with \(C^{1}\) coefficients and not only about the real-analytic ones, so that the invariance of type proved in Theorem 10.8 is matched by an equally general normal form. One honest limitation survives and is stated in Remark A.84: at \(C^{\alpha}\) coefficients the normal form is a statement about the principal part alone, and the classical form with continuous lower-order coefficients costs one more derivative. The single quoted input is Theorem A.95, delimited in Remark A.96.
The Cauchy–Kovalevskaya Theorem
This appendix proves Theorem 10.22 of Partial Differential Equations: a Cauchy problem in normal form, with real-analytic coefficients and real-analytic data prescribed on a non-characteristic surface, has a real-analytic solution near the point in question, and it is the only analytic one. The method is Kovalevskaya's [Kowalevsky:1875] and is the model for every convergence proof of this type: write the solution as a formal power series, observe that the equation and the data determine its coefficients by a recursion whose coefficients are non-negative, replace the data by larger data for which the same recursion can be summed in closed form, and conclude that the original series converges because it is dominated term by term by one that does.
The argument is arranged in four steps. The formal series and its recursion come first (Formal series, the recursion, and the reduction), together with the reduction of a normal problem of order \(k\) to a first-order quasilinear system with vanishing data — carried out at the level of formal series, which is all that is needed and avoids the compatibility questions a reduction of the differential problem would raise. The notion of a majorant and the standard majorizing function then follow (Majorants), the explicit solution of the majorant problem (The majorant problem has a closed-form solution), and the assembly (Proof of the theorem). A closing remark exhibits Lewy's equation, which shows that the analyticity of the data cannot be weakened even to smoothness [Lewy:1957].
Notation. \(t \in \R\) and \(x = (x_{1},\dots,x_{n}) \in \R^{n}\); \(\beta \in \N^{n}\) and \(\gamma \in \N^{m}\) are multi-indices in the sense of Equation (10.2), with \(\N\) containing \(0\); \(x^{\beta} = x_{1}^{\beta_{1}}\cdots x_{n}^{\beta_{n}}\); \(e_{k}\) is the \(k\)-th coordinate multi-index; and a multi-index in the \(1+n\) variables \((t,x)\) is written \(\alpha = (\alpha_{0},\alpha')\) with \(\pp^{\alpha} = \pp_{t}^{\alpha_{0}}\pp_{x}^{\alpha'}\). A function is real analytic near a point if on some polydisc about it the function is the sum of an absolutely convergent power series; that is the only property of analyticity used below.
Formal series, the recursion, and the reduction
A normal Cauchy problem of order \(k\) for the scalar unknown \(w\) is
with \(G\) and the \(\varphi_{p}\) real analytic near the relevant origins. Rests on Definition 10.1 and Proposition 10.11.
For the unknown \(u = (u_{1},\dots,u_{m})\),
with \(a_{ij}^{k}\) and \(b_{i}\) real analytic on a neighbourhood of \((x,u) = (0,0)\) in \(\R^{n+m}\). Rests on Definitions 10.2 and A.102.
Write the coefficients of Equation (A.246) as convergent series \(a_{ij}^{k} = \sum_{\beta,\gamma}a_{ij,\beta\gamma}^{k}\,x^{\beta} u^{\gamma}\) and \(b_{i} = \sum_{\beta,\gamma}b_{i,\beta\gamma}\,x^{\beta}u^{\gamma}\), and seek a formal solution \(u_{i} = \sum_{\beta,l}u_{i,\beta,l}\,x^{\beta}t^{l}\). Then the coefficients \(u_{i,\beta,l}\) are uniquely determined; they vanish for \(l=0\); and for every \(l \ge 1\)
where \(R_{i,\beta,l}\) is a polynomial with non-negative rational coefficients, the same polynomial for every problem of the shape Equation (A.246) with the same \(m\) and \(n\). Rests on Definition A.103 and Equation (10.2).
Derives Lemma A.104. The data give \(u_{i,\beta,0} = 0\) for every \(i,\beta\). Substitute the series into Equation (A.246). On the left,
On the right, \(\pp_{k}u_{j} = \sum_{\beta,l}(\beta_{k}+1)u_{j,\beta+e_{k},l}\,x^{\beta}t^{l}\), whose coefficients are non-negative integers times coefficients of \(u\); and \(a_{ij}^{k}(x,u)\), being a power series in \(x\) and in \(u\) composed with the series for \(u\), has coefficients that are polynomials with non-negative integer coefficients in the \(a_{ij,\beta'\gamma}^{k}\) and the \(u_{j',\beta',l'}\), since composition and multiplication of formal series involve only sums of products. The same holds for \(b_{i}\). Multiplying the two series and collecting the coefficient of \(x^{\beta}t^{l}\), the right-hand side is such a polynomial in which every \(u\)-coefficient carries a time index \(l' \le l\): a factor \(t^{l'}\) from one series can only be accompanied by factors \(t^{l''}\) with \(l'+l''\le l\), and no negative powers occur.
Comparing with Equation (A.248),
with \(Q\) a polynomial with non-negative integer coefficients. Induction on \(l\) therefore determines every coefficient uniquely from the vanishing ones at \(l=0\), and substituting the previously obtained expressions into Equation (A.249) expresses \(u_{i,\beta,l}\) as a polynomial in the coefficients of \(a\) and \(b\) alone. Non-negativity survives every substitution — the only new factors introduced are the positive rationals \(1/(l+1)\) — which is Equation (A.247). That \(R\) does not depend on the particular \(a,b\) is clear from the construction: only \(m\), \(n\) and the indices entered.
∎Every normal Cauchy problem Equation (A.245) has a unique formal power series solution \(\hat w\) about the origin. Moreover there is a system Equation (A.246), with \(m\) and with coefficients \(a_{ij}^{k},b_{i}\) built analytically from \(G\) and the \(\varphi_{p}\), whose unique formal solution \(\hat u\) has one distinguished component equal to
Rests on Lemma A.104 and Definition A.102.
Derives Lemma A.105. Homogenization. Put \(v = w - \sum_{p<k}\varphi_{p}(x)t^{p}/p!\). Since \(\pp_{t}^{p}\bigl(\varphi_{q}t^{q}/q!\bigr)(0,x) = \delta_{pq}\varphi_{q}(x)\) for \(p,q<k\), the new unknown satisfies \(\pp_{t}^{p}v(0,x) = 0\) for \(p<k\), and Equation (A.245) becomes an equation of the same shape with a new right-hand side that is \(G\) evaluated on the shifted arguments — again analytic near the origin, because it is the composition of \(G\) with polynomials in \(t\) whose coefficients are derivatives of the analytic \(\varphi_{p}\).
The system. Take as unknowns the quantities \(u_{\alpha}\) indexed by the multi-indices \(\alpha\) in \((t,x)\) with \(\abs{\alpha}\le k-1\), intended to be \(\pp^{\alpha}v\), together with one extra unknown \(u_{\ast}\) intended to be \(t\). Their equations are
-
\(\pp_{t}u_{\ast} = 1\);
-
\(\pp_{t}u_{\alpha} = u_{\alpha+e_{t}}\) whenever \(\abs{\alpha}\le k-2\), so that \(\abs{\alpha+e_{t}}\le k-1\);
-
for \(\abs{\alpha} = k-1\) with \(\alpha_{0} < k-1\) — so that \(\alpha\) carries at least one \(x\)-derivative, say \(\alpha \ge e_{x_{j}}\) — \(\pp_{t}u_{\alpha} = \pp_{x_{j}}u_{\alpha+e_{t}-e_{x_{j}}}\), the index on the right having length \(k-1\);
-
for \(\alpha = (k-1)e_{t}\), \(\pp_{t}u_{\alpha} = G\bigl(u_{\ast},x,\{\cdot\}\bigr)\), where each argument \(\pp^{\gamma}v\) with \(\abs{\gamma}\le k-1\) is the unknown \(u_{\gamma}\) and each argument with \(\abs{\gamma}=k\), \(\gamma_{0}<k\), carries an \(x\)-derivative and is written \(\pp_{x_{j}}u_{\gamma-e_{x_{j}}}\).
Every right-hand side is analytic in \((x,u)\) near the origin — the \(t\)-slot of \(G\) has been replaced by the unknown \(u_{\ast}\) — and depends linearly on the first \(x\)-derivatives of the unknowns, which is the shape of Equation (A.246). All data vanish: \(u_{\ast}(0,x)=0\), and \(\pp^{\alpha}v(0,x) = \pp_{x}^{\alpha'}\bigl[\pp_{t}^{\alpha_{0}}v(0,x)\bigr] = 0\) because \(\alpha_{0}\le k-1\).
The two formal solutions agree. The recursion of Equation (A.245) determines a unique formal \(\hat v\): rewriting the equation as \(\pp_{t}^{k}v = \widetilde G\) and matching coefficients of \(x^{\beta}t^{l}\) expresses \(l\)-th and higher \(t\)-coefficients in terms of lower ones exactly as in Lemma A.104, the data supplying the \(k\) lowest. Now form the formal derivatives \(\pp^{\alpha}\hat v\) and \(\hat u_{\ast} = t\). Formal differentiation is a ring homomorphism compatible with formal composition, so applying \(\pp^{\alpha}\) to the formal identity \(\pp_{t}^{k}\hat v = \widetilde G(\cdots)\) and to the trivial identities \(\pp_{t}\pp^{\alpha}\hat v = \pp^{\alpha+e_{t}}\hat v\) shows that the family \(\bigl(\pp^{\alpha}\hat v,\ t\bigr)\) satisfies every equation (i)–(iv) formally, with the correct vanishing data. By the uniqueness half of Lemma A.104 it is \(\hat u\), and in particular \(\hat u_{(0,0)} = \hat v\), which is Equation (A.250).
∎The reduction has been made at the level of formal series on purpose. A reduction of the differential problems would require showing that a solution of the system has components that really are the derivatives of its first one, which is an extra argument; here nothing of the kind is needed, because all the reduction is asked to do is to transport convergence from the system to the original problem, and convergence is a property of a formal series.
Majorants
For formal power series \(f = \sum_{\sigma}f_{\sigma}w^{\sigma}\) and \(F = \sum_{\sigma}F_{\sigma}w^{\sigma}\) in the same variables, write \(f \ll F\) — “\(F\) majorizes \(f\)” — if \(\abs{f_{\sigma}} \le F_{\sigma}\) for every multi-index \(\sigma\). In particular every coefficient of \(F\) is then non-negative. Rests on Equation (10.2).
Let \(g = \sum_{\sigma}g_{\sigma}w^{\sigma}\) be real analytic near the origin of \(\R^{N}\), and let \(r>0\) be such that \(C = \sum_{\sigma}\abs{g_{\sigma}}\,r^{\abs{\sigma}} < \infty\). Then \(\abs{g_{\sigma}} \le C\,r^{-\abs{\sigma}}\) and
Rests on Definition A.107.
Derives Lemma A.108. The bound \(\abs{g_{\sigma}}r^{\abs{\sigma}} \le C\) is immediate, every term of the defining sum being non-negative. For the majorant, expand the geometric series and then the powers by the multinomial theorem:
where \(\sigma! = \sigma_{1}!\cdots\sigma_{N}!\). The multinomial coefficient \(\abs{\sigma}!/\sigma!\) is a positive integer, hence at least \(1\), so the coefficient of \(w^{\sigma}\) in Equation (A.252) is at least \(C r^{-\abs{\sigma}} \ge \abs{g_{\sigma}}\).
∎Let Equation (A.246) have coefficients \(a,b\) and let \(A_{ij}^{k}, B_{i}\) be power series with \(a_{ij}^{k} \ll A_{ij}^{k}\) and \(b_{i} \ll B_{i}\). Let \(U_{i,\beta,l}\) be the coefficients of the formal solution of the majorant problem — the system Equation (A.246) with \(A,B\) in place of \(a,b\) and the same vanishing data. Then
Rests on Lemma A.104 and Definition A.107.
Derives Lemma A.109. By Lemma A.104 both sets of coefficients are given by the same polynomials \(R_{i,\beta,l}\), evaluated at the coefficients of \(a,b\) and of \(A,B\) respectively. A polynomial with non-negative coefficients satisfies \(\abs{R(\vect{s})} \le R(\abs{\vect{s}})\) by the triangle inequality applied monomial by monomial, and is non-decreasing in each argument on the non-negative orthant; hence
the middle inequality using \(\abs{a} \le A\) and \(\abs{b} \le B\) coefficientwise.
∎The majorant problem has a closed-form solution
Fix \(C,r>0\) and take, in the notation of Lemma A.108 with the \(N = n+m\) variables \((x,u)\),
Put \(K = m(n+1)\). Then the functions \(U_{i} = V\), \(i=1,\dots,m\), with
solve the majorant problem, are real analytic on the polydisc
and all their Taylor coefficients at the origin are non-negative. Rests on Lemma A.108 and Definition A.103.
Derives Proposition A.110. Write \(\sigma = r-s\) and \(R = \sqrt{\sigma^{2}-2KCrt}\), so \(V = (\sigma - R)/K\). Since \(\pp_{t}R = -KCr/R\) and \(\pp_{\sigma}R = \sigma/R\),
With \(U_{i}=V\) for every \(i\) one has \(U_{1}+\cdots+U_{m} = mV\) and \(\sum_{j,k}\pp_{k}U_{j} = mn\,\pp_{s}V\), so the majorant problem reads
Substituting Equation (A.257), the requirement Equation (A.258) is
that is, after cross-multiplying by \(R\) and by the denominator,
Because \(\sigma\) and \(R\) are functionally independent, Equation (A.259) holds identically if and only if the coefficients match separately: \(1 - m/K = mn/K\) and \(m/K = 1 - mn/K\). Both give \(K = m + mn = m(n+1)\), which is the choice made. At \(t = 0\), \(R = \abs{\sigma} = \sigma\) for \(s<r\), so \(V(0,x) = 0\): the data hold.
Analyticity and positivity. Factor
where \(1-\sqrt{1-y} = \sum_{p\ge1}c_{p}y^{p}\) is the binomial series, whose coefficients \(c_{p} = \dfrac{(2p-2)!}{2^{2p-1}\,p!\,(p-1)!}\) are strictly positive, and which converges absolutely for \(\abs{y}<1\). Each factor \(\sigma^{1-2p} = r^{1-2p}\bigl(1-s/r\bigr)^{-(2p-1)}\) expands, for \(\abs{s}<r\), as \(r^{1-2p}\sum_{d\ge0}\binom{2p-2+d}{d}(s/r)^{d}\), again with positive coefficients, and \(s^{d} = (x_{1}+\cdots+x_{n})^{d}\) expands with non-negative integer coefficients in the \(x^{\beta}\). Hence every Taylor coefficient of \(V\) is non-negative, and the multiple series converges absolutely wherever \(\abs{s} < r\) and \(2KCr\abs{t} < \bigl(r-\abs{s}\bigr)^{2}\). On the polydisc Equation (A.256) one has \(\abs{s} \le \sum_{k}\abs{x_{k}} < r/4\), so \(\bigl(r-\abs{s}\bigr)^{2} > 9r^{2}/16\) and the second condition follows from \(\abs{t} < 9r/(32KC)\).
∎The choice Equation (A.254) is the whole trick, and its two features are independent. Making all the coefficient functions the same single function of \(s\) and of \(u_{1}+\cdots+u_{m}\) collapses the \(m\)-component system to one scalar equation in two variables; and making that function the geometric kernel \(Cr/(r-s-\sum u_{i})\) makes the resulting scalar equation quasilinear of first order with the closed-form solution Equation (A.255). Notice that Equation (A.258) is exactly an equation of the type treated by Proposition 10.15: written as \(\bigl(\sigma - mV\bigr)\pp_{t}V - Crmn\,\pp_{s}V = Cr\) it is quasilinear of first order in the two variables \((t,s)\), with characteristic system \(\dot t = \sigma - mV\), \(\dot s = -Crmn\), \(\dot V = Cr\). It is because that system can be integrated — or, as done above, because the resulting closed form can be guessed and then verified — that the majorant problem is solvable at all; the square root in Equation (A.255) is the mark of the quadratic relation between \(V\) and \(t\) that the last two characteristic equations impose. That the exponent works out to \(K = m(n+1)\) and not to something depending on the data is what makes the radius of convergence in Equation (A.256) explicit.
Proof of the theorem
The problem Equation (A.246) has a real-analytic solution on a polydisc about the origin of \(\R^{1+n}\), and it is the only analytic solution there. Rests on Lemma A.109, Proposition A.110 and Lemma A.108.
Derives Theorem A.112. By hypothesis the \(a_{ij}^{k}\) and \(b_{i}\) are analytic near \((x,u) = (0,0)\); choose \(r>0\) so small that all of them are represented by absolutely convergent series on the polydisc of radius \(r\) in the \(n+m\) variables, and let \(C\) be the largest of the finitely many sums \(\sum_{\beta\gamma}\abs{a_{ij,\beta\gamma}^{k}}r^{\abs{\beta}+\abs{\gamma}}\) and \(\sum_{\beta\gamma}\abs{b_{i,\beta\gamma}}r^{\abs{\beta}+\abs{\gamma}}\). By Lemma A.108 every coefficient is then majorized by the function Equation (A.254).
Let \(U_{i,\beta,l}\) be the formal solution of the majorant problem. By Proposition A.110 the function \(V\) of Equation (A.255) solves that problem and is analytic on the polydisc Equation (A.256); its Taylor coefficients therefore satisfy the recursion of Lemma A.104 for the majorant data, and by the uniqueness half of that lemma they are the \(U_{i,\beta,l}\). Combining with Lemma A.109,
on Equation (A.256), where \(\abs{x}^{\beta}\) means \(\abs{x_{1}}^{\beta_{1}}\cdots\); the middle equality holds because all the \(U\) are non-negative, so the series is its own absolute value. Hence the formal series for each \(u_{i}\) converges absolutely on that polydisc and defines a function there.
The sum solves the problem. On any closed subpolydisc of ratio \(\theta<1\) the differentiated series are dominated term by term by \(\theta^{-1}\sup_{q\ge0}\bigl[(q+1)\theta^{q}\bigr]\) times the original one, hence converge uniformly by Lemma 9.6, so term-by-term differentiation in \(t\) and in each \(x_{k}\) is legitimate, one variable at a time, by Theorem 7.51. The coefficients satisfy Equation (A.249) by construction, so the analytic function \(\pp_{t}u_{i} - \sum_{j,k}a_{ij}^{k}(x,u)\pp_{k}u_{j} - b_{i}(x,u)\) has every Taylor coefficient equal to zero and therefore vanishes identically; and \(u_{i}(0,x)=0\) because \(u_{i,\beta,0}=0\).
Uniqueness. An analytic solution on any polydisc about the origin has a Taylor series which, by the computation of Lemma A.104, must satisfy the same recursion with the same initial values; the coefficients are therefore those just constructed, and two analytic functions with the same Taylor series at a point agree near it.
∎Proof of Theorem 10.22. Derives Theorem 10.22. Let the problem be in the normal form Equation (A.245), which is what normal Cauchy problem means, with the initial surface already the slice \(\set{t=0}\). By Lemma A.105 its unique formal solution \(\hat w\) is recovered by Equation (A.250) from the distinguished component of the formal solution \(\hat u\) of an associated system Equation (A.246). By Theorem A.112 that formal solution converges absolutely on a polydisc about the origin; hence so does \(\hat u_{(0,0)}\), and hence so does \(\hat w\), the finitely many extra terms \(\varphi_{p}(x)t^{p}/p!\) being analytic. Its sum \(w\) is analytic, and its Taylor coefficients satisfy the recursion of the normal problem, so — by the argument already used in Theorem A.112 — the analytic function \(\pp_{t}^{k}w - G(\cdots)\) has all Taylor coefficients zero and vanishes, while \(\pp_{t}^{p}w(0,x)=\varphi_{p}(x)\) for \(p<k\) by construction. Any other analytic solution has the same Taylor coefficients, by the same recursion, and therefore coincides with \(w\) near the origin.
∎The proof above is complete for a problem presented in the normal form Equation (A.245), which is the form Theorem 10.22 names. Bringing a general non-characteristic Cauchy problem to that form uses two facts that this treatise does not prove:
-
Flattening the initial surface. If the analytic hypersurface \(S\) is \(\set{\phi = 0}\) with \(\nabla\phi\neq0\), the map \(x \longmapsto \bigl(\phi(x),\psi_{2}(x),\dots\bigr)\) is an analytic change of coordinates near the point, by the analytic inverse function theorem.
-
Solving for the highest normal derivative. For a fully nonlinear equation \(F\bigl(x,\set{\pp^{\alpha}u}\bigr)=0\), non-characteristic means \(\pp F/\pp\bigl(\pp_{t}^{k}u\bigr) \neq 0\) at the point (Proposition 10.11 is the linear instance of the same computation), and the analytic implicit function theorem then supplies an analytic \(G\) with \(\pp_{t}^{k}u = G(\cdots)\).
Both are the analytic versions of theorems whose smooth versions belong to Real Analysis; both are themselves usually proved by the majorant method of this section, so nothing circular is being borrowed, but neither is derived here and neither is stated in the treatise. For a linear or quasilinear equation — which is every equation this book actually solves — step (ii) is not needed at all: being non-characteristic then means that the coefficient of \(\pp_{t}^{k}u\) is nonzero, by Equation (10.8) and the computation in the proof of Proposition 10.11, and one divides by it. Nothing else in this section is quoted: the recursion, the reduction, the majorant, the closed-form solution and the convergence estimate are all carried out above.
Every hypothesis of Theorem 10.22 was used, and the analyticity of the data — not merely of the coefficients — is the one whose necessity is least obvious and most complete. Consider on \(\R^{3}\), with coordinates \((x_{1},x_{2},x_{3})\), the first-order operator
Its coefficients are polynomials, hence analytic, and it is elementary to check that it is non-characteristic in the \(x_{3}\)-direction at points where \(x_{1}+\ii x_{2} \neq 0\). Lewy proved that there exist \(f \in C^{\infty}(\R^{3})\) — real-valued, and depending on \(x_{3}\) alone — for which \(L u = f\) has no solution of class \(C^{1}\) on any neighbourhood of any point whatever [Lewy:1957]. The failure is therefore not of uniqueness or of regularity but of existence, it is local and universal, and it is produced by an operator whose coefficients satisfy every hypothesis of Cauchy–Kovalevskaya. What is lost when \(f\) ceases to be analytic is exactly what the proof above consumed: the formal series still exists in the analytic case because the recursion Equation (A.249) always runs, and it is the majorant that turns it into a function. With a merely smooth right-hand side there is no series to majorize, and Equation (A.262) shows that nothing takes its place.
The reader should also keep in view the second limitation, which the chapter states in Remark 10.23 and which is independent of this one: even where the theorem applies it establishes existence and not well-posedness. Analytic data cannot be varied independently in disjoint regions, so continuous dependence in the sense of Definition 10.24 is not even being asked for, and the Cauchy problem Cauchy–Kovalevskaya solves for the Laplace equation is the ill-posed one of Remark 10.26.
The Cauchy–Kovalevskaya Theorem discharges the proof obligation of Theorem 10.22 in Partial Differential Equations. The result is used in the treatise wherever an analytic construction is needed and nothing weaker will do: it is the existence statement behind the classical, real-analytic branch of Proposition 10.13 — the complex first integral constructed there by complexifying Equation (10.13) is supplied by this theorem, and The Korn–Lichtenstein Theorem and the Elliptic Canonical Form is what replaces it when the coefficients are only \(C^{1}\) — and it is the reason a formal power-series ansatz may be trusted at all when the data of a problem are analytic. Its limitations are the substance of Remark A.114 and Remark 10.23, and they are the reason the rest of Partial Differential Equations proceeds by energy estimates, transforms and Green's functions rather than by series.
The Stäckel Conditions and the Eleven Separable Systems
This appendix proves Theorem 10.34 of Partial Differential Equations: the Helmholtz equation Equation (10.25) separates in an orthogonal coordinate system of three-dimensional Euclidean space if and only if the reciprocal squared scale factors are the first column of the inverse of a Stäckel matrix — a matrix each of whose rows depends on one coordinate alone — and the volume element satisfies the accompanying Robertson condition. Both conditions are derived here, not postulated: they are read off by substituting the product ansatz into the Laplacian and demanding that the three separation constants enter the separated ordinary differential equations linearly. The three systems the treatise uses (Example 10.31, Example 10.32, Example 10.33) are then verified against the criterion, and so is the generic member of the family, the ellipsoidal system, from which all the others descend by confluence. The section closes with the enumeration of the eleven systems and with an exact statement of the one input that is quoted rather than proved.
Everything is set in the observed three-dimensional space, signature \(3+0\), with the flat Euclidean metric; the classification is a statement about \(\R^{3}\) and has no higher-dimensional counterpart in this treatise. Latin indices \(i,j,l\) run over \(1,2,3\) and are not summed unless a summation sign is written, because the sums that appear below are not all over the same range.
Orthogonal coordinates and the Laplacian
Let \(\vect{x} = \vect{x}(q^{1},q^{2},q^{3})\) be a \(C^{2}\) diffeomorphism of an open set of \(\R^{3}\) onto another, with
The \(h_{i}\) are the scale factors, the line element is \(\dd\ell^{2} = \sum_{i}h_{i}^{2}\,(\dd q^{i})^{2}\), and the volume element is \(\dd^{3}x = g\,\dd q^{1}\dd q^{2}\dd q^{3}\) with \(g = h_{1}h_{2}h_{3}\). Rests on Definition 7.66 and Equation (10.25).
For \(\psi \in C^{2}\),
Rests on Definition A.116 and Theorem 7.101.
Derives Lemma A.117. Let \(\hat{\vect{e}}_{i} = h_{i}^{-1}\pp\vect{x}/\pp q^{i}\); by Equation (A.263) these form an orthonormal triad at each point. The gradient is characterized by \(\dd\psi = \nabla\psi\cdot\dd\vect{x}\); writing \(\dd\vect{x} = \sum_{i}h_{i}\,\dd q^{i}\,\hat{\vect{e}}_{i}\) and \(\dd\psi = \sum_{i}\pp_{i}\psi\,\dd q^{i}\) and matching coefficients of the independent increments \(\dd q^{i}\),
For the divergence of \(\vect{V} = \sum_{i}V_{i}\hat{\vect{e}}_{i}\) apply the divergence theorem Equation (7.123) to the coordinate box \([q^{1},q^{1}+\dd q^{1}]\times\cdots\). Its face at fixed \(q^{1}\) has outward normal \(-\hat{\vect{e}}_{1}\) and area \(h_{2}h_{3}\,\dd q^{2}\dd q^{3} = (g/h_{1})\,\dd q^{2}\dd q^{3}\), and the opposite face has the same area evaluated at \(q^{1}+\dd q^{1}\); the net flux through the pair is therefore \(\pp_{1}\bigl((g/h_{1})V_{1}\bigr)\,\dd q^{1}\dd q^{2}\dd q^{3}\) to first order. Summing over the three pairs and dividing by the volume \(g\,\dd q^{1}\dd q^{2}\dd q^{3}\),
Composing Equation (A.266) with Equation (A.265), so that \(V_{i} = h_{i}^{-1}\pp_{i}\psi\), gives Equation (A.264).
∎What separation means, and what it forces
The Helmholtz equation Equation (10.25) separates simply in the orthogonal system \((q^{i})\) if there exist nowhere-vanishing functions \(f_{i}(q^{i})\) and functions \(\phi_{ij}(q^{i})\) (\(i,j = 1,2,3\)), each row index \(i\) carrying dependence on \(q^{i}\) alone, with
such that for every triple of constants \((\lambda_{1},\lambda_{2},\lambda_{3})\) with \(\lambda_{1} = k^{2}\), any solutions \(X_{i}\) of the three ordinary differential equations
compose into a solution \(\psi = X_{1}X_{2}X_{3}\) of Equation (10.25). The matrix \(\Phi = (\phi_{ij})\) is the Stäckel matrix and \(S\) its Stäckel determinant. Rests on Definition A.116 and Lemma 10.28.
Three remarks fix the content of Definition A.118 before it is used. First, the form Equation (A.268) is not a restriction but the general second-order linear equation in \(q^{i}\) with no first-derivative term beyond the one a weight can absorb: any equation \(X'' + p X' + \sigma X = 0\) is brought to it by \(f = \exp\int p\). Second, the requirement that the constants \(\lambda_{j}\) enter linearly is the definition of separation in Stäckel's sense, and it is the requirement that carries the physics: each \(\lambda_{j}\) is an eigenvalue of a commuting operator, and the separated problem is a Sturm–Liouville problem in each variable (Ordinary Differential Equations and Sturm–Liouville Theory) exactly because Equation (A.268) is linear in them. Third, \(\lambda_{1} = k^{2}\) singles out the first column of \(\Phi\); the remaining two constants are free, and \(S\neq0\) is what makes them genuinely independent.
Let \(\Phi = (\phi_{ij})\) be a \(3\times3\) matrix with \(S = \det\Phi \neq 0\), and let \(M_{ij}\) be the cofactor of \(\phi_{ij}\), that is \((-1)^{i+j}\) times the determinant of the \(2\times2\) matrix obtained by deleting row \(i\) and column \(j\). Then
and, if the \(i\)-th row of \(\Phi\) depends on \(q^{i}\) alone, then \(M_{i1}\) does not depend on \(q^{i}\) at all. Rests on Equation (A.267).
Derives Lemma A.119. For \(j = 1\) the sum in Equation (A.269) is the Laplace expansion of \(\det\Phi\) along the first column. For \(j \neq 1\) it is the Laplace expansion along the first column of the matrix obtained from \(\Phi\) by overwriting its first column with its \(j\)-th, since the cofactors \(M_{i1}\) do not involve the first column at all; that matrix has two equal columns and hence vanishing determinant. Dividing Equation (A.269) by \(S\) identifies \(M_{i1}/S\) as the \((1,i)\) entry of the inverse. The last assertion is immediate: the minor belonging to \(M_{i1}\) is built from the two rows other than the \(i\)-th, which by hypothesis carry no dependence on \(q^{i}\).
∎The Helmholtz equation separates simply in the orthogonal system \((q^{i})\), in the sense of Definition A.118, if and only if there is a Stäckel matrix \(\Phi = (\phi_{ij}(q^{i}))\) with \(S = \det\Phi \neq 0\) such that
and the volume element satisfies the Robertson condition
for some functions \(f_{i}\) of one variable each — the same \(f_{i}\) that weight the separated equations Equation (A.268). Rests on Definition A.118, Lemma A.117 and Lemma A.119.
Sufficiency. Derives Theorem A.120. Assume Equation (A.270) and Equation (A.271), and let the \(X_{i}\) solve Equation (A.268). Put \(\psi = X_{1}X_{2}X_{3}\). By Equation (A.270) and \(g = S f_{1}f_{2}f_{3}\),
and by the last clause of Lemma A.119 both \(M_{i1}\) and the factors \(f_{l}\) with \(l \neq i\) are independent of \(q^{i}\). Hence
because everything except \(f_{i}X_{i}'\) passes through \(\pp_{i}\) untouched. Insert Equation (A.273) into Equation (A.264) and divide by \(\psi = X_{1}X_{2}X_{3}\), using \(g = Sf_{1}f_{2}f_{3}\) once more:
The separated equations Equation (A.268) say that the last factor equals \(-\sum_{j}\lambda_{j}\phi_{ij}\), so by the cofactor identity Equation (A.269)
that is \(\nabla^{2}\psi + k^{2}\psi = 0\). Note where each hypothesis acted: Equation (A.271) made Equation (A.273) a one-variable derivative, and Equation (A.270) collapsed the sum in Equation (A.275) onto the single constant \(\lambda_{1}\), killing the two free constants exactly as separation requires.
∎Necessity. Derives Theorem A.120. Assume the equation separates simply and substitute \(\psi = X_{1}X_{2}X_{3}\) into Equation (10.25) through Equation (A.264), dividing by \(\psi\):
obtained by expanding \(g^{-1}\pp_{i}\bigl((g/h_{i}^{2})\pp_{i}\psi\bigr)\) and cancelling the \(l \neq i\) factors of \(\psi\).
Step 1: the weights. The \(i\)-th bracket in Equation (A.276) must be expressible through \(q^{i}\) and \(X_{i}\) alone — that is what makes the \(i\)-th separated equation an ordinary differential equation in \(q^{i}\) — and \(X_{i}\) is an arbitrary solution of one, so \(X_{i}''/X_{i}\) and \(X_{i}'/X_{i}\) are functionally independent quantities that cannot cancel each other's coefficients. Hence \(\pp_{i}\ln(g/h_{i}^{2})\) is a function of \(q^{i}\) alone. Call it \((\ln f_{i})'(q^{i})\), which defines \(f_{i}>0\) up to a constant; the bracket is then \(\bigl(f_{i}X_{i}'\bigr)'/(f_{i}X_{i})\), and Equation (A.276) reads
each \(E_{i}\) a function of \(q^{i}\) alone.
Step 2: linearity in the constants. By hypothesis the separated equations carry the three constants linearly, so \(E_{i} = -\sum_{j}\lambda_{j}\phi_{ij}(q^{i})\) with \(\lambda_{1}=k^{2}\) and \(\lambda_{2},\lambda_{3}\) free. Substituting into Equation (A.277),
and since this must hold identically in the three independent constants \(\lambda_{j}\), every bracket vanishes: \(\sum_{i}\phi_{ij}/h_{i}^{2} = \delta_{j1}\). In matrix language the row vector \(v\) with \(v_{i} = 1/h_{i}^{2}\) satisfies \(v\,\Phi = e_{1}\transpose\), whence \(\Phi\) is invertible — so \(S\neq0\), which is the independence of the constants — and \(v_{i} = \bigl(\Phi^{-1}\bigr)_{1i} = M_{i1}/S\) by Lemma A.119. That is Equation (A.270).
Step 3: the Robertson condition. Step 1 gave \(\pp_{i}\ln\bigl(g/h_{i}^{2}\bigr) = (\ln f_{i})'(q^{i})\). Insert Equation (A.270): \(g/h_{i}^{2} = (g/S)\,M_{i1}\), and \(M_{i1}\) is independent of \(q^{i}\) by Lemma A.119, so its logarithmic \(q^{i}\)-derivative vanishes and
Write \(\Lambda = \ln\abs{g/S}\). Then \(\pp_{1}\Lambda\) depends on \(q^{1}\) alone, so \(\Lambda = \ln\abs{f_{1}(q^{1})} + \Lambda_{1}(q^{2},q^{3})\); applying Equation (A.279) for \(i=2\) to this expression gives \(\pp_{2}\Lambda_{1} = (\ln f_{2})'(q^{2})\), hence \(\Lambda_{1} = \ln\abs{f_{2}(q^{2})} + \Lambda_{2}(q^{3})\), and one more step gives \(\Lambda_{2} = \ln\abs{f_{3}(q^{3})}\) up to an additive constant, which is absorbed into any one \(f_{i}\). Exponentiating, \(\abs{g/S} = \abs{f_{1}f_{2}f_{3}}\); the sign is constant on a connected domain and is absorbed likewise, giving Equation (A.271).
∎Equation (A.270) is a condition on the metric alone and is the substantive one: it says the three functions \(1/h_{i}^{2}\), which in general are functions of all three coordinates, are the first column of the inverse of a matrix assembled out of nine functions of one coordinate each. That is an enormous restriction, and it is what makes the list of admissible systems finite. Equation (A.271) is a condition on the volume element relative to \(S\), and is the one that distinguishes the Helmholtz operator from the Hamilton–Jacobi equation of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy: the latter involves no second derivatives of \(\psi\) and so imposes Equation (A.270) alone. Every system in the list below satisfies both, but the logical order matters — Stäckel's condition governs separation of the classical problem, Robertson's is the extra price of the wave operator, which is why the two subjects are almost but not exactly the same, as Remark 10.35 records.
The three systems the treatise uses
In each case we exhibit \(\Phi\), \(S\), the cofactors \(M_{i1}\) and the weights \(f_{i}\), and check Equation (A.270) and Equation (A.271) outright. The separation constants are named to match Section 10.2.1.
\(h_{1}=h_{2}=h_{3}=1\), \(g=1\). With \(\lambda_{1}=k^{2}\), \(\lambda_{2}=\beta^{2}\), \(\lambda_{3}=\gamma^{2}\) the separated equations of Example 10.31 are \(X''+(k^{2}-\beta^{2}-\gamma^{2})X=0\), \(Y''+\beta^{2}Y=0\), \(Z''+\gamma^{2}Z=0\), so \(f_{i}=1\) and
The cofactors of the first column are \(M_{11} = +\det\begin{pmatrix}1&0\\0&1\end{pmatrix} = 1\), \(M_{21} = -\det\begin{pmatrix}-1&-1\\0&1\end{pmatrix} = 1\) and \(M_{31} = +\det\begin{pmatrix}-1&-1\\1&0\end{pmatrix} = 1\), all equal to \(1/h_{i}^{2}\); and \(g/S = 1 = f_{1}f_{2}f_{3}\). Rests on Theorem A.120 and Example 10.31.
\((q^{1},q^{2},q^{3}) = (\rho,\varphi,z)\) with \(h_{1}=1\), \(h_{2}=\rho\), \(h_{3}=1\) and \(g=\rho\). Set \(\lambda_{1}=k^{2}\), \(\lambda_{2}=m^{2}\), \(\lambda_{3}=\gamma^{2}\); the separated equations of Example 10.32 are Equation (10.32), \(\Psi''+m^{2}\Psi=0\) and \(Z''+\gamma^{2}Z=0\), so \(f_{1}=\rho\), \(f_{2}=f_{3}=1\) and
Then \(M_{11} = 1 = 1/h_{1}^{2}\), \(M_{21} = -\det\begin{pmatrix}-\rho^{-2}&-1\\0&1\end{pmatrix} = \rho^{-2} = 1/h_{2}^{2}\) and \(M_{31} = +\det\begin{pmatrix}-\rho^{-2}&-1\\1&0\end{pmatrix} = 1 = 1/h_{3}^{2}\); and \(g/S = \rho = f_{1}f_{2}f_{3}\). Rests on Theorem A.120 and Example 10.32.
\((q^{1},q^{2},q^{3}) = (r,\theta,\varphi)\) with \(h_{1}=1\), \(h_{2}=r\), \(h_{3}=r\sin\theta\) and \(g = r^{2}\sin\theta\). Set \(\lambda_{1}=k^{2}\), \(\lambda_{2}=l(l+1)\), \(\lambda_{3}=m^{2}\); the separated equations of Example 10.33 are Equation (10.34), the associated Legendre equation and \(\Psi''+m^{2}\Psi = 0\), so \(f_{1}=r^{2}\), \(f_{2}=\sin\theta\), \(f_{3}=1\) and
The first-column cofactors are \(M_{11} = 1 = 1/h_{1}^{2}\), \(M_{21} = -\det\begin{pmatrix} -r^{-2}&0\\0&1\end{pmatrix} = r^{-2} = 1/h_{2}^{2}\) and \(M_{31} = +\det\begin{pmatrix} -r^{-2}&0\\1&-(\sin\theta)^{-2}\end{pmatrix} = r^{-2}\bigl(\sin\theta\bigr)^{-2} = 1/h_{3}^{2}\); and \(g/S = r^{2}\sin\theta = f_{1}f_{2}f_{3}\). The three checks reproduce, in one line each, the structure the chapter obtained by hand: the centrifugal term \(-l(l+1)/r^{2}\) of Equation (10.34) is the entry \(\phi_{12}\), and the term \(-m^{2}/\sin^{2}\theta\) of the Legendre equation is \(\phi_{23}\). Rests on Theorem A.120 and Example 10.33.
The generic system: confocal quadrics
The three systems just checked are degenerate members of one family. The generic member is the ellipsoidal system, and it is worth carrying out in full because everything else follows from it by letting parameters coincide or run to infinity.
Fix \(a_{1}^{2} > a_{2}^{2} > a_{3}^{2} > 0\) and write \(x_{1},x_{2},x_{3}\) for Cartesian coordinates and
The confocal family is
Through a generic point pass exactly three members of the family, with parameters \(\xi_{1} > \xi_{2} > \xi_{3}\) interlacing the poles,
an ellipsoid, a hyperboloid of one sheet and a hyperboloid of two sheets respectively. These are the ellipsoidal coordinates. Rests on Definition A.116.
With the notation of Definition A.125,
the system is orthogonal, and
Rests on Definition A.125 and Equation (A.263).
Derives Lemma A.126. Multiply Equation (A.284), in the form \(1 - \sum_{p}x_{p}^{2}/(\theta+a_{p}^{2}) = 0\), by \(P(\theta)\). The result,
is a monic cubic in \(\theta\) whose roots are by definition \(\xi_{1},\xi_{2},\xi_{3}\), which identifies it with the right-hand side. Evaluating Equation (A.288) at \(\theta = -a_{p}^{2}\) kills \(P\) and every term of the sum except the \(p\)-th, giving \(-x_{p}^{2}\prod_{q\neq p}(a_{q}^{2}-a_{p}^{2}) = -\prod_{i}(\xi_{i}+a_{p}^{2})\), which is Equation (A.286) after the two sign changes in the denominator.
Differentiating Equation (A.286) with respect to \(\xi_{i}\) removes one factor from the numerator: \(2x_{p}\,\pp_{i}x_{p} = x_{p}^{2}/(\xi_{i}+a_{p}^{2})\), hence
Divide Equation (A.288) by \(P(\theta)\) and rearrange:
For \(i \neq j\) the inner product of the two coordinate tangent vectors is, by Equation (A.289) and the partial fraction \(\bigl[(\xi_{i}+a)(\xi_{j}+a)\bigr]^{-1} = (\xi_{j}-\xi_{i})^{-1} \bigl[(\xi_{i}+a)^{-1}-(\xi_{j}+a)^{-1}\bigr]\),
because \(N(\xi_{i}) = N(\xi_{j}) = 0\) makes \(G(\xi_{i}) = G(\xi_{j}) = 1\) by Equation (A.290): the system is orthogonal. For the scale factors, Equation (A.289) gives \(h_{i}^{2} = \frac14\sum_{p}x_{p}^{2} \bigl(\xi_{i}+a_{p}^{2}\bigr)^{-2} = -\tfrac14 G'(\xi_{i})\), and differentiating Equation (A.290),
using \(N(\xi_{i})=0\). That is Equation (A.287). Positivity follows from Equation (A.285): \(P(\xi_{1})>0\) and \(\prod_{l\neq1}(\xi_{1}-\xi_{l})>0\); \(P(\xi_{2})<0\) and \(\prod_{l\neq2}(\xi_{2}-\xi_{l})<0\); \(P(\xi_{3})>0\) and \(\prod_{l\neq3}(\xi_{3}-\xi_{l})>0\).
∎For distinct \(\xi_{1},\xi_{2},\xi_{3}\) and an integer \(0 \le p \le 2\),
Rests on Equation (A.287).
Derives Lemma A.127. The polynomial \(L(\theta) = \sum_{i}\xi_{i}^{\,p}\prod_{l\neq i} \frac{\theta-\xi_{l}}{\xi_{i}-\xi_{l}}\) has degree at most \(2\) and agrees with \(\theta^{p}\) at the three distinct points \(\xi_{1},\xi_{2},\xi_{3}\). Since \(\theta^{p}\) also has degree at most \(2\), their difference is a polynomial of degree at most \(2\) with three roots, hence identically zero: \(L(\theta) = \theta^{p}\). Comparing the coefficients of \(\theta^{2}\) on the two sides gives Equation (A.293), the left-hand coefficient being the sum displayed and the right-hand one \(\delta_{p2}\).
∎The ellipsoidal system of Definition A.125 satisfies Equation (A.270) and Equation (A.271) with
and, with \(\epsilon_{1}=\epsilon_{3}=+1\) and \(\epsilon_{2}=-1\) the signs of \(P\) on the three intervals of Equation (A.285),
Only the product of the three signs is fixed, by Equation (A.271); the sign of an individual \(f_{i}\) is immaterial because \(f_{i}\) occurs twice in Equation (A.268). The separated equations Equation (A.268) become the Lamé wave equations
with \(P\) evaluated at \(\xi_{i}\). Rests on Lemma A.126, Lemma A.127 and Theorem A.120.
Derives Proposition A.128. Each row of \(\Phi\) in Equation (A.294) depends on \(\xi_{i}\) alone, as a Stäckel matrix must. By Equation (A.287), \(1/h_{i}^{2} = 4P(\xi_{i})/\prod_{l\neq i}(\xi_{i}-\xi_{l})\), so
by Lemma A.127. By Step 2 of the proof of Theorem A.120 this identity is equivalent to Equation (A.270), and it also shows \(S \neq 0\).
For the determinant, factor \(1/\bigl(4P(\xi_{i})\bigr)\) out of the \(i\)-th row: \(S = \bigl[64\prod_{i}P(\xi_{i})\bigr]^{-1}\det\bigl(\xi_{i}^{3-j}\bigr)\). Reversing the column order — one transposition, of columns \(1\) and \(3\) — turns \(\det(\xi_{i}^{3-j})\) into \(-\det(\xi_{i}^{j-1})\), and the latter is the Vandermonde determinant \(\prod_{i<l}(\xi_{l}-\xi_{i}) = -W\); hence \(\det(\xi_{i}^{3-j}) = W\) and \(S\) is as stated.
For Robertson, take the product of Equation (A.287) over \(i\). The numerator is \(\prod_{i}\prod_{l\neq i}(\xi_{i}-\xi_{l}) = \prod_{i<l}\bigl[-(\xi_{i}-\xi_{l})^{2}\bigr] = -W^{2}\), so \(g^{2} = -W^{2}\bigl[64\prod_{i}P(\xi_{i})\bigr]^{-1}\); the sign is consistent because \(\prod_{i}P(\xi_{i}) < 0\) by Equation (A.285), and therefore
the second because \(\prod_{i}\epsilon_{i} = -1\). Dividing,
the last step by Equation (A.295), whose three signs multiply to \(-1\): a product of three functions of one variable each, which is exactly Equation (A.271). Substituting \(f_{i}\) and \(\sum_{j}\lambda_{j}\phi_{ij} = \bigl(k^{2}\xi_{i}^{2}+\lambda_{2}\xi_{i}+\lambda_{3}\bigr)/ \bigl(4P(\xi_{i})\bigr)\) into Equation (A.268) gives Equation (A.296), the constant factor \(2\) in \(f_{i}\) cancelling between the two occurrences.
∎For \(k = 0\), Equation (A.296) is Lamé's equation and its polynomial solutions are the ellipsoidal harmonics, the analogue for a triaxial body of the spherical harmonics Equation (9.163). The chapter's three examples are visible in Equation (A.296) as limits: when two of the \(a_{p}^{2}\) coincide one of the three coordinates becomes an angle and the corresponding Lamé equation degenerates to the associated Legendre equation Equation (9.157); when all three coincide the remaining two degenerate to the Legendre equation and to \(\Psi''+m^{2}\Psi=0\), and Example A.124 is recovered.
The enumeration
What remains is to count the systems. The count has two halves, and only one of them is proved here.
Let \((q^{i})\) be an orthogonal coordinate system of \(\R^{3}\) in which the Helmholtz equation separates simply. Then, up to a Euclidean motion and a relabelling and reparametrization of the coordinates, the coordinate surfaces are the members of a confocal family of quadrics Equation (A.284) or of one of its degenerate limits.
Theorem A.130 is the only statement in this section that is not proved. It is the substantial half of the classification: the assertion that a metric of flat space admitting a Stäckel matrix must have confocal quadrics for its coordinate surfaces. Its proof is a long analysis of the second-order system that Equation (A.270) imposes on the nine functions \(\phi_{ij}\) together with the vanishing of the Riemann tensor of the Euclidean metric written in the coordinates \(q^{i}\) — a computation belonging to classical differential geometry rather than to the theory of partial differential equations, and one this treatise does not carry out anywhere. The result and the resulting tables of separated equations, system by system, are those of Morse and Feshbach [Morse:1953]; the differential-geometric analysis behind it is Eisenhart's, whose paper is not among the sources catalogued in this treatise's bibliography, so the attribution here is uncited and the reader who wishes to check it must go to the secondary account just named. Everything else in this section — the two conditions, their necessity and sufficiency, the ellipsoidal verification and the degeneration analysis below — is derived in full.
Granting Theorem A.130, the enumeration becomes a finite case analysis over the ways the confocal family Equation (A.284) can degenerate, and that we carry out. The family is fixed by the cubic \(P\) of Equation (A.283), i.e. by the three parameters \(a_{1}^{2} > a_{2}^{2} > a_{3}^{2}\); a common shift of all three is a reparametrization \(\theta \to \theta+c\) and changes nothing, so only the two differences \(a_{1}^{2}-a_{2}^{2}\) and \(a_{2}^{2}-a_{3}^{2}\) matter, and a common rescaling fixes the unit of length. The degenerations are then exhausted by three independent binary choices:
-
Coincidence. Either both differences are nonzero (three distinct foci), or exactly one vanishes (an axis of rotational symmetry appears, and there are two inequivalent ways for this to happen — the focal set degenerates to a segment when \(a_{2}^{2}=a_{3}^{2}\) and to a disc when \(a_{1}^{2}=a_{2}^{2}\)), or both vanish (full rotational symmetry: concentric spheres).
-
A focus at infinity. One may let one root of \(P\) recede, \(a_{1}^{2}\to\infty\), while rescaling; the ellipsoids and one family of hyperboloids open up into paraboloids. This can be done to the generic family and to the rotationally symmetric one, but not to the spherical one, where there is nothing left to send away.
-
Translation. One may instead let the whole configuration become independent of one Cartesian direction, so that the quadrics become cylinders over a two-dimensional confocal family of conics. The two-dimensional families are themselves classified by the same first two choices, one dimension down: confocal conics with two distinct foci, confocal parabolas, concentric circles, or the degenerate family of parallel lines.
Reading the choices off gives the list, in which the first seven are the genuinely three-dimensional families and the last four the cylindrical ones.
| System | Coordinate surfaces | Degeneration |
|---|---|---|
| Ellipsoidal | three confocal quadrics | none: $a_{1}^{2}>a_{2}^{2}>a_{3}^{2}$ |
| Paraboloidal | elliptic and hyperbolic paraboloids | one focus to infinity |
| Prolate spheroidal | prolate spheroids, two-sheeted hyperboloids, half-planes | $a_{2}^{2}=a_{3}^{2}$ (focal segment) |
| Oblate spheroidal | oblate spheroids, one-sheeted hyperboloids, half-planes | $a_{1}^{2}=a_{2}^{2}$ (focal disc) |
| Parabolic | two families of paraboloids of revolution, half-planes | rotational, one focus to infinity |
| Spherical | spheres, cones of revolution, half-planes | $a_{1}^{2}=a_{2}^{2}=a_{3}^{2}$ |
| Conical | spheres and two families of elliptic cones | spheres with the focal cone retained |
| Elliptic cylindrical | confocal elliptic and hyperbolic cylinders, planes | translational, two foci |
| Parabolic cylindrical | confocal parabolic cylinders, planes | translational, one focus at infinity |
| Circular cylindrical | circular cylinders, half-planes, planes | translational, coincident foci |
| Cartesian | three families of parallel planes | translational, all foci at infinity |
The degeneration scheme just described yields exactly the eleven systems of Table A.1, and no others. Rests on Proposition A.128 and Theorem A.130.
Derives Proposition A.132. Consider first the non-translational families. The coincidence pattern of the three parameters has four cases: all distinct; \(a_{2}^{2}=a_{3}^{2}\); \(a_{1}^{2}=a_{2}^{2}\); all equal. In the first three cases one may additionally send a focus to infinity or not, which doubles them — but the two rotationally symmetric cases have the same paraboloidal limit, because sending the distinct root to infinity destroys the distinction between a focal segment and a focal disc, leaving the single parabolic system. That gives \(1+1\) (ellipsoidal, paraboloidal) plus \(2+1\) (prolate, oblate, parabolic), i.e. five. The fully symmetric case admits no focus at infinity, but it does admit two inequivalent coordinate systems on the sphere: the one whose second and third coordinates are the polar angle and the azimuth — spherical — and the one in which the sphere is cut by two families of elliptic cones sharing a vertex, which is the conical system and is the residue of the ellipsoidal system when the quadrics have collapsed but their asymptotic cones have not. That gives seven.
The translational families are cylinders over a two-dimensional confocal family of conics, and the same analysis one dimension down applies to the plane cross-section: two distinct foci give confocal ellipses and hyperbolae (elliptic cylindrical); sending one focus to infinity gives confocal parabolae (parabolic cylindrical); coincident foci give concentric circles and their radii (circular cylindrical); and sending both to infinity gives two families of parallel lines (Cartesian). No further case arises, because a family of conics in the plane is fixed by the two foci alone, and their configuration — distinct, coincident, one at infinity, both at infinity — has exactly these four types. That gives four, and \(7+4 = 11\).
∎Proof of Theorem 10.34. Derives Theorem 10.34. The criterion asserted by the theorem is Theorem A.120: separation holds if and only if the reciprocal squared scale factors are expressible through a matrix each of whose rows depends on one coordinate alone — explicitly, if and only if \(1/h_{i}^{2} = M_{i1}/S\) in Equation (A.270), together with the Robertson condition Equation (A.271) on the volume element, which the chapter's statement subsumes under “the Stäckel conditions”. The list of eleven systems is Table A.1: that each of them separates is verified by exhibiting its Stäckel matrix, done in Examples A.122, A.123 and A.124 for the three the treatise uses and in Proposition A.128 for the generic one, from which the remaining entries follow by the confluences of Proposition A.132; and that there are no others is Proposition A.132 together with the quoted Theorem A.130.
∎The Stäckel Conditions and the Eleven Separable Systems discharges the proof obligation of Theorem 10.34 in Partial Differential Equations. Two consequences are used elsewhere in the treatise. The first is negative and is the practical one recorded in Remark 10.35: a boundary that is not a coordinate surface of one of the eleven systems admits no separation, which is why the boundary chooses the coordinates and not the operator. The second is that Equation (A.270) alone — without Robertson's Equation (A.271) — is the condition for the Hamilton–Jacobi equation of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy to separate, so the separation constants \(\lambda_{2},\lambda_{3}\) of Equation (A.268) are the quantum counterparts of the classical integrals of motion of an integrable system; the spherical case (Example A.124), where they are \(l(l+1)\) and \(m^{2}\), is the one worked out physically in The Hydrogen Atom.
The Mean-Value Property Characterizes Harmonic Functions
This appendix proves the continuous form of the converse half of Theorem 10.69 (Partial Differential Equations): a function that is merely continuous on an open set and equals its own spherical average over every ball whose closure the set contains is automatically infinitely differentiable there, and harmonic. The chapter proves the converse only for \(u \in C^{2}\), because the argument it uses decides the sign of a Laplacian whose existence is exactly what is in question (Remark 10.70). The gap is closed here by constructing the smoothing device the chapter does not have — a radial mollifier — and showing that convolution against it returns the function unchanged. Once that is done the chapter's own sign argument applies verbatim.
Nothing is quoted. Everything below is built from the Riemann integral of a continuous function over a compact region, the mean value theorem of one-variable calculus, and the polar decomposition of a volume integral in \(\R^{3}\) already used in Lemma 10.58 and in the proof of Theorem 10.69; in particular the Lebesgue theory that Real Analysis deliberately does not develop is nowhere needed, because every integrand that appears is continuous with compact support.
Throughout, \(\Omega \subset \R^{3}\) is open and nonempty, \(B(\vect{x},r) = \set{\vect{y} \in \R^{3} \mid \abs{\vect{y}-\vect{x}} < r}\), and \(M_{u}(\vect{x},r)\) is the spherical mean Equation (10.54). The space is the observed three-dimensional one; nothing in the argument depends on that choice beyond the numerical factor \(4\pi\).
Statement
A function \(u \in C^{0}(\Omega)\) has the mean-value property on \(\Omega\) if and only if
for every \(\vect{x} \in \Omega\) and every \(r>0\) with \(\overline{B(\vect{x},r)} \subset \Omega\). Rests on Equation (10.54) and Definition 6.2.
Let \(u \in C^{0}(\Omega)\) have the mean-value property of Definition A.134. Then \(u \in C^{\infty}(\Omega)\) and \(\nabla^{2}u = 0\) on \(\Omega\); that is, \(u\) is harmonic in the sense of Definition 10.66. Consequently the two conditions harmonic and continuous with the mean-value property define the same class of functions. Rests on Definition A.134, Theorem 10.69 and Definition 10.66.
The proof occupies the rest of the section: an elementary lemma on differentiating an integral with respect to a parameter (Differentiation under the integral sign), the construction of the mollifier and the smoothness of the convolution (A radial mollifier), the reproduction identity \(u * \rho_{\varepsilon} = u\) (The mollification reproduces the function), and the assembly (Proof of the theorem).
Differentiation under the integral sign
The one analytic tool the argument needs is stated and proved here, in the only form it will be used: the integrand is continuous, the domain of integration is a fixed compact set, and the parameter runs over an open set.
Let \(U \subset \R^{3}\) be open, let \(Q \subset \R^{3}\) be a compact box, and let \(F : U \times Q \longrightarrow \R\) be such that \(F\) and the partial derivative \(\pp F/\pp x^{i}\) (taken in the first argument) are continuous on \(U \times Q\) for one fixed \(i\). Then
has a continuous partial derivative \(\pp_{i}g\) on \(U\), and
Rests on Definition 7.66 and Theorem 6.11.
Derives Lemma A.136. Fix \(\vect{x} \in U\) and choose \(r>0\) with \(\overline{B(\vect{x},r)} \subset U\). The set \(\overline{B(\vect{x},r)}\times Q\) is compact, so \(\pp F/\pp x^{i}\) is uniformly continuous on it: given \(\varepsilon>0\) there is \(\delta \in (0,r)\) such that
for every \(\vect{y} \in Q\), the bound being uniform in \(\vect{y}\) — which is the whole content of uniform continuity on the product and the reason the compactness of \(Q\) is needed.
Let \(\vect{e}_{i}\) be the \(i\)-th coordinate vector and let \(0 < \abs{h} < \delta\). For each fixed \(\vect{y}\) the mean value theorem applied to the function \(s \longmapsto F(\vect{x}+s\vect{e}_{i},\vect{y})\) on the interval between \(0\) and \(h\) supplies \(\theta = \theta(h,\vect{y}) \in (0,1)\) with
Both sides of Equation (A.304) are continuous in \(\vect{y}\) — the left side manifestly, and hence the right side as well, whatever the measurability of \(\theta\) — so both are Riemann integrable over \(Q\). Subtracting Equation (A.302) and using Equation (A.303) with \(\abs{\theta h} < \delta\),
with \(\abs{Q}\) the volume of the box. Since \(\varepsilon\) was arbitrary the difference quotient converges, which is Equation (A.302). Continuity of \(\pp_{i}g\) follows from the same estimate with \(\vect{x}'' = \vect{x}\) and \(\vect{x}'\) a nearby point, the integral of a quantity bounded by \(\varepsilon\) being bounded by \(\varepsilon\abs{Q}\).
∎Lemma A.136 iterates: if \(F\) and all its partial derivatives in \(\vect{x}\) of every order are continuous on \(U\times Q\), then \(g \in C^{\infty}(U)\) and every derivative may be taken under the integral sign, because the conclusion of the lemma is again a function of the same form with \(\pp F/\pp x^{i}\) in place of \(F\).
A radial mollifier
The device the chapter lacks is a smooth, radial, non-negative function supported in a ball and of unit integral. Its existence is not obvious: a function that is analytic cannot vanish on an open set without vanishing identically, so the construction must use a function that is smooth and not analytic. There is exactly one standard source of such a function.
Define \(f : \R \longrightarrow \R\) by
Then \(f \in C^{\infty}(\R)\), \(f^{(n)}(0) = 0\) for every \(n\), and \(f(t)>0\) exactly for \(t>0\). Rests on Definition 7.66 and Theorem 7.38.
Derives Lemma A.138. Derivatives for \(t>0\). We claim that for every \(n \ge 0\) there is a polynomial \(p_{n}\) with
For \(n=0\) take \(p_{0} = 1\). Differentiating Equation (A.307) and using \(\dd(1/t)/\dd t = -1/t^{2}\),
so Equation (A.307) holds at \(n+1\) with \(p_{n+1}(s) = s^{2}\bigl(p_{n}(s) - p_{n}'(s)\bigr)\).
Decay at the origin. For \(s>0\) and any integer \(m \ge 0\) the exponential series gives \(\ee^{s} \ge s^{m+1}/(m+1)!\), hence \(s^{m}\ee^{-s} \le (m+1)!/s \longrightarrow 0\) as \(s \to +\infty\). Consequently, for any polynomial \(p\),
Derivatives at the origin. We show by induction that \(f^{(n)}\) exists everywhere, is continuous, and \(f^{(n)}(0)=0\). For \(n=0\) this is Equation (A.309) with \(p=1\). Assume it at \(n\). For \(t<0\) the difference quotient of \(f^{(n)}\) at \(0\) vanishes identically; for \(t>0\) it is
which tends to \(0\) by Equation (A.309) applied to the polynomial \(s\,p_{n}(s)\). Hence \(f^{(n+1)}(0)\) exists and is \(0\). It is continuous at \(0\) because \(f^{(n+1)}(t) = p_{n+1}(1/t)\ee^{-1/t} \to 0\) as \(t\to0^{+}\) by Equation (A.309) again, and vanishes for \(t<0\). Positivity for \(t>0\) is immediate.
∎With \(f\) of Equation (A.306) set
and, for \(\varepsilon > 0\), \(\rho_{\varepsilon}(\vect{x}) = \varepsilon^{-3}\rho(\vect{x}/\varepsilon)\). Rests on Lemma A.138.
\(\rho_{\varepsilon} \in C^{\infty}(\R^{3})\); it is non-negative, vanishes outside \(\overline{B(\vect{0},\varepsilon)}\), depends on \(\vect{x}\) only through \(\abs{\vect{x}}\), and \(\int_{\R^{3}}\rho_{\varepsilon}\,\dd^{3}x = 1\). Rests on Definition A.139 and Lemma A.138.
Derives Lemma A.140. \(\abs{\vect{x}}^{2} = x^{2}+y^{2}+z^{2}\) is a polynomial, hence \(C^{\infty}\), and the composition of a \(C^{\infty}\) function of one variable with a \(C^{\infty}\) function of three is \(C^{\infty}\) by the chain rule (Proposition 7.72); this is where Lemma A.138 is used, and it is the only place where smoothness that is not analyticity is needed. The argument \(1-\abs{\vect{x}}^{2}\) is positive exactly on the open unit ball and non-positive outside it, so \(\rho > 0\) there and \(\rho = 0\) elsewhere; in particular \(\rho\) vanishes on a whole neighbourhood of every point of \(\abs{\vect{x}} = 1\) together with all its derivatives. The normalizing integral in Equation (A.311) is the integral of a continuous non-negative function that is positive on an open set, hence finite and strictly positive, so \(c_{0}\) is well defined. Radial dependence is manifest. Finally the substitution \(\vect{x} = \varepsilon\vect{z}\), whose Jacobian is \(\varepsilon^{3}\), gives \(\int\rho_{\varepsilon}(\vect{x})\dd^{3}x = \int\rho(\vect{z})\dd^{3}z = 1\), and rescales the support to \(\abs{\vect{x}} \le \varepsilon\).
∎For \(u \in C^{0}(\Omega)\) and \(\varepsilon>0\) put
an open subset of \(\Omega\) whose union over \(\varepsilon>0\) is \(\Omega\), and define on it
The two integrals in Equation (A.313) agree by the substitution \(\vect{y} = \vect{x}-\vect{z}\), of unit Jacobian; both integrands are continuous on a compact set, so both are ordinary Riemann integrals. That \(\Omega_{\varepsilon}\) is open, and that every point of \(\Omega\) lies in some \(\Omega_{\varepsilon}\), is immediate from \(\Omega\) being open.
\(u_{\varepsilon} \in C^{\infty}(\Omega_{\varepsilon})\), and every partial derivative may be taken under the integral sign in the second form of Equation (A.313). Rests on Lemmas A.136 and A.140.
Derives Proposition A.142. Fix \(\vect{x}_{0} \in \Omega_{\varepsilon}\) and \(\sigma>0\) so small that \(\overline{B(\vect{x}_{0},\sigma+\varepsilon)} \subset \Omega\); let \(U = B(\vect{x}_{0},\sigma)\) and let \(Q\) be a closed box containing \(\overline{B(\vect{x}_{0},\sigma+\varepsilon)}\). Extend \(u\) from \(\overline{B(\vect{x}_{0},\sigma+\varepsilon)}\) to \(Q\) by any continuous function — the extension is irrelevant, because for \(\vect{x} \in U\) the factor \(\rho_{\varepsilon}(\vect{x}-\vect{y})\) already vanishes for \(\vect{y}\) outside \(\overline{B(\vect{x}_{0},\sigma+\varepsilon)}\), so
for every \(\vect{x} \in U\). In Equation (A.314) the whole dependence on the parameter \(\vect{x}\) sits in the factor \(\rho_{\varepsilon}\), which by Lemma A.140 is \(C^{\infty}\); hence \(F\) and all its \(\vect{x}\)-derivatives of every order,
are continuous on \(U\times Q\), being products of continuous functions. Lemma A.136 and Remark A.137 therefore apply and give \(u_{\varepsilon} \in C^{\infty}(U)\) with \(\pp^{\alpha}u_{\varepsilon}(\vect{x}) = \int_{Q}\bigl(\pp^{\alpha}\rho_{\varepsilon}\bigr) (\vect{x}-\vect{y})\,u(\vect{y})\,\dd^{3}y\). Since \(\vect{x}_{0} \in \Omega_{\varepsilon}\) was arbitrary, the conclusion holds on all of \(\Omega_{\varepsilon}\).
∎The mollification reproduces the function
This is the step at which the mean-value hypothesis enters, and it is the reason the mollifier was required to be radial.
Let \(u \in C^{0}(\Omega)\) have the mean-value property. Then
Rests on Definition A.134, Lemma A.140 and Definition A.141.
Derives Proposition A.143. Write \(\rho_{\varepsilon}(\vect{z}) = \varepsilon^{-3}\tilde\rho\bigl(\abs{\vect{z}}/\varepsilon\bigr)\), where \(\tilde\rho(s) = c_{0}f(1-s^{2})\) is the radial profile supplied by Lemma A.140. Decompose the first integral of Equation (A.313) into spheres, \(\dd^{3}z = s^{2}\,\dd s\,\dd\Omega\) — the same decomposition already used in Lemma 10.58 to pass between a solid and a spherical integral, and legitimate here because the integrand is continuous on the compact ball:
The inner integral is unchanged under \(\vect{n} \longmapsto -\vect{n}\), which is a symmetry of the unit sphere and of its solid-angle element, so it equals \(\int_{S^{2}}u(\vect{x}+s\vect{n})\,\dd\Omega = 4\pi\,M_{u}(\vect{x},s)\). For \(\vect{x} \in \Omega_{\varepsilon}\) and \(0 < s \le \varepsilon\) the closed ball \(\overline{B(\vect{x},s)}\) lies in \(\Omega\), so the mean-value property Equation (A.300) gives \(M_{u}(\vect{x},s) = u(\vect{x})\) — a constant, which comes out of the \(s\)-integral:
the middle step being Equation (A.317) read backwards with \(u \equiv 1\), and the last being the normalization of Lemma A.140.
∎Both hypotheses on \(\rho\) are used and neither can be dropped. Radial dependence is what lets the weight \(\rho_{\varepsilon}\) be pulled outside the angular integral in Equation (A.317), leaving exactly a spherical mean for the hypothesis to act on; a non-radial bump would leave an angular weight and the mean-value property would say nothing about the result. Unit mass is what makes the surviving factor equal to \(1\) rather than to some other constant. The identity Equation (A.316) is thus not an approximation statement — it is not \(u_{\varepsilon} \to u\), which holds for every continuous \(u\) — but an exact equality, valid for each \(\varepsilon\) separately, and that exactness is the whole content of the argument.
Proof of the theorem
Proof of Theorem A.135. Derives Theorem A.135. Smoothness. Let \(\vect{x}_{0} \in \Omega\). Since \(\Omega\) is open there is \(\varepsilon>0\) with \(\overline{B(\vect{x}_{0},2\varepsilon)} \subset \Omega\), and then \(B(\vect{x}_{0},\varepsilon) \subset \Omega_{\varepsilon}\). On \(\Omega_{\varepsilon}\) we have \(u = u_{\varepsilon}\) by Proposition A.143, and \(u_{\varepsilon} \in C^{\infty}(\Omega_{\varepsilon})\) by Proposition A.142. Hence \(u\) is \(C^{\infty}\) on a neighbourhood of \(\vect{x}_{0}\). Smoothness being a local property and \(\vect{x}_{0}\) arbitrary, \(u \in C^{\infty}(\Omega)\). Note that the smoothing parameter \(\varepsilon\) is allowed to depend on the point, as it must: no single \(\varepsilon\) works near \(\pp\Omega\).
Harmonicity. Now that \(u \in C^{2}(\Omega)\) is established, the argument of Theorem 10.69 applies without change, and we repeat it for completeness. Suppose \(\nabla^{2}u(\vect{x}_{0}) > 0\) for some \(\vect{x}_{0} \in \Omega\). By continuity of \(\nabla^{2}u\) there is \(r>0\) with \(\overline{B(\vect{x}_{0},r)} \subset \Omega\) and \(\nabla^{2}u > 0\) throughout that ball. Darboux's identity Equation (10.55) then gives, for \(0 < s \le r\),
so \(s \longmapsto M_{u}(\vect{x}_{0},s)\) is strictly increasing on \((0,r]\). But the mean-value property makes it constant, equal to \(u(\vect{x}_{0})\), on that whole interval — a contradiction. The case \(\nabla^{2}u(\vect{x}_{0}) < 0\) is the same argument with the inequalities reversed. Hence \(\nabla^{2}u \equiv 0\) on \(\Omega\), which is Definition 10.66.
The two classes coincide. A harmonic function is \(C^{2}\) by definition and has the mean-value property by the direct half of Theorem 10.69; conversely a continuous function with the mean-value property is harmonic by what has just been proved.
∎Every harmonic function on \(\Omega\) is of class \(C^{\infty}(\Omega)\), although its definition (Definition 10.66) demands only \(C^{2}\). Rests on Theorems 10.69 and A.135.
Derives Corollary A.145. A harmonic \(u\) is continuous and has the mean-value property (Theorem 10.69); Theorem A.135 then returns \(u \in C^{\infty}(\Omega)\).
∎If \(u_{n}\) are harmonic on \(\Omega\) and \(u_{n} \to u\) uniformly on every compact subset of \(\Omega\), then \(u\) is harmonic. Rests on Theorem A.135 and Definition A.134.
Derives Corollary A.146. \(u\) is continuous, being a locally uniform limit of continuous functions. Fix \(\vect{x}\) and \(r\) with \(\overline{B(\vect{x},r)}\subset\Omega\); the sphere \(\abs{\vect{y}-\vect{x}} = r\) is compact, so \(u_{n} \to u\) uniformly on it and the averages converge: \(M_{u}(\vect{x},r) = \lim_{n} M_{u_{n}}(\vect{x},r) = \lim_{n} u_{n}(\vect{x}) = u(\vect{x})\). Thus \(u\) is continuous with the mean-value property, and Theorem A.135 applies. The corresponding statement with \(C^{2}\) convergence would be trivial; the point is that mere uniform convergence suffices, and it does so only because the characterization proved here does not mention derivatives at all.
∎Nothing. The three inputs are the mean value theorem of one-variable calculus, uniform continuity of a continuous function on a compact set (Theorem 6.11), and the polar decomposition \(\dd^{3}z = s^{2}\,\dd s\,\dd\Omega\) of a volume integral, all of them already in force in Real Analysis and used by the chapter itself in Lemma 10.58. Every integrand appearing above is continuous on a compact region, so the Riemann integral of Real Analysis suffices throughout and no appeal to Lebesgue's theory — which this treatise does not build — is made or needed. In particular Lemma A.136 is proved, not cited: the usual statement of differentiation under the integral sign is a dominated-convergence argument, and the compactness of the domain of integration replaces the domination here.
The gain over Theorem 10.69 is not cosmetic. The \(C^{2}\) converse cannot be applied to a function one has only constructed as a limit, an average or a supremum, because such a function is typically known to be continuous and nothing more; the continuous converse can, and Corollary A.146 is the first instance. It is also the cleanest statement of the rigidity of the Laplace equation: the second-order differential condition \(\nabla^{2}u=0\) is equivalent to a condition — Equation (A.300) — that mentions no derivative whatever, and that is why solutions of it are automatically smooth while solutions of, say, the wave equation are not.
The Mean-Value Property Characterizes Harmonic Functions discharges the proof obligation left open in Remark 10.70 and completes Theorem 10.69 of Partial Differential Equations. The chapter uses only the \(C^{2}\) form in what follows — the strong maximum principle Theorem 10.77 is proved from it — so nothing there depended on the present section; what the section adds is the equivalence that makes the mean-value property a definition of harmonicity rather than a consequence of one, and with it the stability under uniform limits recorded in Corollary A.146. The mollifier constructed in A radial mollifier is the same device that underlies the smooth test functions of Fourier Analysis and Integral Transforms and the approximation arguments of Section 10.6.2.
Kruzhkov's Doubling of Variables and the Entropy Solution
This appendix proves Theorem 10.103 of Partial Differential Equations: a scalar conservation law Equation (10.81) with bounded measurable datum has at most one entropy solution in the sense of Definition 10.101, two entropy solutions contract in the mean, and at least one exists. The uniqueness half is Kruzhkov's argument by doubling of variables, and it is the whole content of the section; the existence half is the vanishing-viscosity limit already announced in Remark 10.100, whose architecture is set out in Existence by vanishing viscosity with its two compactness inputs named.
Three features of the argument are worth flagging before it begins, because each is what makes it work. The first is the choice of entropy: the family \(\eta(u) = \abs{u-k}\), indexed by a real constant \(k\), is not \(C^{2}\) and so is not admitted by Definition 10.101 directly; it is reached by approximation in From the admissible entropies to Kruzhkov's family, and it is the family for which the entropy flux of two solutions can be compared. The second is that the comparison is carried out in four variables — \(u\) is evaluated at \((t,x)\) and \(v\) at \((s,y)\) — with each solution's own entropy inequality tested against the same function and the constant \(k\) taken to be the other solution. The third is the cancellation that makes the whole thing finite: with a test function of the form \(\chi\bigl(\tfrac{t+s}{2},\tfrac{x+y}{2}\bigr)\) times a mollifier in the differences, the two entropy inequalities add so that the mollifier is differentiated in neither of them.
Throughout, \(F \in C^{1}(\R)\), \(u\) and \(v\) are entropy solutions of Equation (10.81) with bounded measurable data \(u_{0}, v_{0}\),
and \(\varphi\) denotes a non-negative test function of the kind appearing in Equation (10.92). Convexity of \(F\), assumed in Theorem 10.103, is nowhere used below: Kruzhkov's theorem needs only that \(F\) be \(C^{1}\), which is a strictly stronger result and is recorded in Remark A.165.
Statement
For \(k \in \R\) put
Let \(u,v\) be entropy solutions in the sense of Definition 10.101 obeying Equation (A.320). Then for almost every \(0 < \tau_{1} < \tau_{2}\) and every \(R>0\)
and, using Lemma A.154, the same inequality holds with \(\tau_{1}=0\) and \(u,v\) replaced there by \(u_{0},v_{0}\). If in addition \(u_{0}-v_{0}\) is integrable on \(\R\), then for almost every \(t>0\)
An entropy solution with a given bounded measurable datum is unique up to a null set, and the solution map \(u_{0}\longmapsto u(t,\cdot)\) is a contraction for the \(L^{1}\) distance — which is condition (iii) of Definition 10.24 with the mean as the norm. Rests on Theorem A.151 and Definition 10.24.
Derives Corollary A.152. Let \(u_{0} = v_{0}\) almost everywhere. Then the right-hand side of Equation (A.322) with \(\tau_{1}=0\) vanishes for every \(R\), so \(u(t,\cdot) = v(t,\cdot)\) almost everywhere on every bounded interval and hence on \(\R\), for almost every \(t\). Continuous dependence is Equation (A.323) itself, the map being \(1\)-Lipschitz from \(L^{1}\) to \(L^{1}\).
∎The inputs this treatise does not build
Unlike the rest of Partial Differential Equations, the subject of this section is not formulated in terms of the Riemann integral and cannot be. A weak solution in the sense of Definition 10.95 is a bounded measurable function, its defining identity is an integral over a region of the plane against an arbitrary test function, and the conclusions are statements valid almost everywhere; every one of these words belongs to Lebesgue's theory, which Real Analysis does not develop — it builds the Riemann integral only, as Probability and Statistics also records where it needs dominated convergence and declines to use it. The following facts are therefore assumed, and they are assumed for the definitions as much as for the proofs:
-
The Lebesgue integral of a bounded measurable function over a measurable subset of \(\R^{d}\), its linearity and monotonicity, and the fact that a function vanishing almost everywhere has zero integral.
-
Fubini's theorem for a bounded measurable function on a product, used in Proposition A.159 to integrate one solution's entropy inequality over the other solution's variables, and again to change variables from \((t,x,s,y)\) to sum and difference coordinates.
-
Dominated convergence, used at each passage to a limit.
-
Continuity of translation in \(L^{1}\): for \(w\) bounded and measurable and \(K\) compact, \(\int_{K}\abs{w(\cdot+h)-w} \longrightarrow 0\) as \(h \to 0\). Equivalently, almost every point is a Lebesgue point. This is the one substantial analytic fact of Proposition A.162, and it is exactly what replaces the continuity that a classical solution would have had.
-
Lebesgue's differentiation theorem in one variable, used in The limit and the cone estimate to pass from an average over a short time interval to a value at almost every instant.
-
Kruzhkov's initial-layer lemma (Lemma A.154): an entropy solution attains its datum in the local mean. This is the only theorem of the subject that is quoted rather than proved, and it is isolated as such below.
-
For the existence half only (Existence by vanishing viscosity): classical solvability of the viscous problem, and a compactness theorem for families of uniformly bounded functions of uniformly bounded variation.
Everything else — the passage from the \(C^{2}\) entropies of Definition 10.101 to the Kruzhkov family, the doubling, the cancellation, the limit, the cone estimate, and the derivation of the entropy inequality from the viscous equation — is carried out in full. The primary source is [Kruzhkov:1970], where the entropy class, the doubling of variables and the \(L^{1}\) contraction first appear together.
Let \(u\) be an entropy solution in the sense of Definition 10.101 with datum \(u_{0}\). Then for every compact \(K \subset \R\)
the limit being taken through the instants at which \(u(t,\cdot)\) is defined.
Equation (A.324) is stronger than what the definition plainly gives, and the gap is worth naming. Testing Equation (10.85) with \(\varphi(x,t)=\rho(x)\alpha(t)\) shows that \(t \longmapsto \int u(t,x)\rho(x)\,\dd x\) agrees almost everywhere with an absolutely continuous function taking the value \(\int u_{0}\rho\) at \(t = 0\), so the datum is attained weakly; that much is a two-line computation from the definition. Weak attainment is not enough for Theorem A.151, because the functional \(w \longmapsto \int\abs{w}\) is only lower semicontinuous under weak convergence and the inequality needed runs the other way. Kruzhkov obtains Equation (A.324) from the entropy inequality itself applied with constant \(k\) and a test function concentrating at \(t=0\), which yields an upper bound on \(\limsup_{t\to0}\int\abs{u(t,x)-k}\rho\) by \(\int\abs{u_{0}-k}\rho\); a covering argument over a countable dense set of values \(k\), together with approximation of \(u_{0}\) by step functions, upgrades that to the strong statement. The argument is measure-theoretic throughout and is not reproduced here.
From the admissible entropies to Kruzhkov's family
Let \(u\) be an entropy solution in the sense of Definition 10.101. Then for every \(k \in \R\) and every non-negative test function \(\varphi\),
with \(q_{k}\) of Equation (A.321). Rests on Definition 10.101, Definition A.150 and Equation (10.92).
Derives Proposition A.156. For \(\delta>0\) set
Then \(\eta_{\delta}\) is \(C^{\infty}\), with
so it is convex; \(q_{\delta}\) is \(C^{1}\) with \(q_{\delta}' = \eta_{\delta}'F'\) by the fundamental theorem of calculus. Thus \((\eta_{\delta},q_{\delta})\) is an admissible entropy pair in the sense of Definition 10.101, and Equation (10.92) holds for it.
Uniform convergence of the entropy. From \(\abs{w-k} \le \sqrt{(w-k)^{2}+\delta^{2}} \le \abs{w-k}+\delta\),
Uniform convergence of the flux on the range. Let \(\abs{w} \le M_{0}\) and abbreviate \(r_{\delta}(\lambda) = \sgn(\lambda-k) - \eta_{\delta}'(\lambda)\), so that \(\abs{r_{\delta}} \le 2\) and, for \(\abs{\lambda-k}\ge\varepsilon\),
Since \(\sgn(\lambda-k)\) is constant and equal to \(\sgn(w-k)\) for \(\lambda\) strictly between \(k\) and \(w\), \(\int_{k}^{w}\sgn(\lambda-k)F'(\lambda)\,\dd\lambda = \sgn(w-k)\bigl(F(w)-F(k)\bigr) = q_{k}(w)\), whence
the interval of integration having been split at \(\abs{\lambda-k}=\varepsilon\) and its length bounded by \(2M_{0}\) — here \(\abs{k} \le M_{0}\) may be assumed, since for \(\abs{k}>M_{0}\) the sign of \(u-k\) is constant and Equation (A.325) reduces to Equation (10.85) with a sign. Choosing first \(\varepsilon\) and then \(\delta\), the right-hand side of Equation (A.330) is made arbitrarily small uniformly in \(w\): \(q_{\delta} \longrightarrow q_{k}\) uniformly on \([-M_{0},M_{0}]\).
Passage to the limit. Apply Equation (10.92) to \((\eta_{\delta},q_{\delta})\) and subtract the corresponding expression built from \((\eta_{k},q_{k})\). Each of the three differences is bounded in modulus by the uniform bounds Equation (A.328) and Equation (A.330) times the integral of \(\abs{\pp_{t}\varphi}\), \(\abs{\pp_{x}\varphi}\) or \(\abs{\varphi(\cdot,0)}\) over the compact support of \(\varphi\), a fixed finite number. Letting \(\delta\to0\) therefore gives Equation (A.325). No convergence theorem is needed at this step: the convergence is uniform on a set of finite measure.
∎For all \(a,b \in [-M_{0},M_{0}]\), with \(\Phi(a,b) = \sgn(a-b)\bigl(F(a)-F(b)\bigr)\),
so that \(\Phi\) is symmetric, continuous, and Lipschitz:
Rests on Definition A.150 and Equation (A.320).
Derives Lemma A.157. If \(a>b\) then \(\sgn(a-b)=1\), \(\max=a\), \(\min=b\) and both sides of Equation (A.331) equal \(F(a)-F(b)\); if \(a<b\) then \(\sgn(a-b)=-1\), \(\max=b\), \(\min=a\) and both sides equal \(F(b)-F(a)\); if \(a=b\) both vanish. Symmetry and continuity are then manifest, \(\max\) and \(\min\) being continuous. For the bounds, the mean value theorem gives \(\abs{F(\max)-F(\min)} \le M\abs{\max-\min} = M\abs{a-b}\), which is the first; and since \(a \longmapsto \max(a,b)\) and \(a\longmapsto\min(a,b)\) are \(1\)-Lipschitz in each argument, so is \(F\circ\max\) and \(F\circ\min\) up to the factor \(M\), giving the second.
∎Equation (A.331) is the reason Kruzhkov's family is the right one. What the doubling will produce is the pair \(\bigl(\abs{u-v},\,\Phi(u,v)\bigr)\) built from two solutions, and for this to be usable it must be a genuine entropy pair in each variable separately with the other frozen — which is exactly what Definition A.150 provides, the frozen solution playing the role of the constant \(k\). No single convex \(\eta\) of Definition 10.101 has that property, because a general \(\eta\) knows nothing about the second solution. The price is that \(\abs{u-k}\) is not \(C^{2}\), and Proposition A.156 is what pays it.
Doubling the variables
Let \(\psi = \psi(t,x,s,y) \ge 0\) be smooth with compact support in \(\set{t>0}\times\R\times\set{s>0}\times\R\). Then
the integral being over \((t,x,s,y)\) with \(t,s>0\). Rests on Proposition A.156 and Lemma A.157.
Derives Proposition A.159. Fix \((s,y)\) with \(s>0\) and apply Equation (A.325) to the solution \(u\), with the constant \(k = v(s,y)\) — legitimate for almost every \((s,y)\), since \(\abs{v(s,y)}\le M_{0}\) is then a genuine real number — and with the non-negative test function \((t,x)\longmapsto\psi(t,x,s,y)\), which is smooth with compact support in \(\set{t>0}\times\R\). The initial term is absent because \(\psi\) vanishes for \(t\) near \(0\). Recalling \(q_{v(s,y)}\bigl(u(t,x)\bigr) = \Phi\bigl(u(t,x),v(s,y)\bigr)\) from Equation (A.321),
The left-hand side of Equation (A.334) is a bounded measurable function of \((s,y)\) — bounded because the integrand is bounded by \(2M_{0}\abs{\pp_{t}\psi}+2MM_{0}\abs{\pp_{x}\psi}\) using Equation (A.332) — so it may be integrated over \(\set{s>0}\times\R\), and Fubini's theorem rearranges the result into a single integral over the four variables. This gives Equation (A.333) with only the \(\pp_{t}\) and \(\pp_{x}\) terms.
Now exchange the roles. Fix \((t,x)\) with \(t>0\) and apply Equation (A.325) to \(v\), with \(k = u(t,x)\) and the test function \((s,y)\longmapsto\psi(t,x,s,y)\); by the symmetry of \(\Phi\) recorded in Lemma A.157, \(q_{u(t,x)}\bigl(v(s,y)\bigr) = \Phi\bigl(v(s,y),u(t,x)\bigr) = \Phi\bigl(u(t,x),v(s,y)\bigr)\), and \(\abs{v-u} = \abs{u-v}\). Integrating over \((t,x)\) gives Equation (A.333) with the \(\pp_{s}\) and \(\pp_{y}\) terms. Adding the two inequalities gives Equation (A.333).
∎Let \(\chi \ge 0\) be smooth with compact support in \((0,\infty)\times\R\), let \(\theta \in C^{\infty}(\R)\) be even and non-negative with support in \((-1,1)\) and \(\int\theta = 1\), and put \(\theta_{\delta}(\sigma) = \delta^{-1}\theta(\sigma/\delta)\) and
For \(\delta\) small enough \(\psi_{\delta}\) is admissible in Proposition A.159, and in the sum-and-difference coordinates
one has
Rests on Proposition A.159 and Equation (A.335).
Derives Proposition A.160. By the chain rule applied to Equation (A.336), \(\pp_{t} = \tfrac12\pp_{\tau}+\tfrac12\pp_{\sigma}\) and \(\pp_{s} = \tfrac12\pp_{\tau}-\tfrac12\pp_{\sigma}\), so \(\pp_{t}+\pp_{s} = \pp_{\tau}\); likewise \(\pp_{x}+\pp_{y} = \pp_{\zeta}\). In Equation (A.335) the factors \(\theta_{\delta}\) depend only on \(\sigma\) and \(\omega\) and are therefore annihilated by \(\pp_{\tau}\) and \(\pp_{\zeta}\), while \(\chi\) depends only on \((\tau,\zeta)\); that is Equation (A.337). Admissibility: the support of \(\psi_{\delta}\) lies within distance \(2\delta\) of the diagonal \(t=s\), \(x=y\) over the support of \(\chi\), so if \(2\delta\) is smaller than the distance from the support of \(\chi\) to \(\set{\tau = 0}\) then \(t,s>0\) throughout.
∎This is the step the method is named for and the only one that could fail. Had the mollifier been differentiated, each derivative would have produced a factor \(\delta^{-1}\), and the limit \(\delta\to0\) would have been meaningless. It survives because the two entropy inequalities are added, not subtracted, and because the sum of the two time derivatives — one for each copy of the equation — is a derivative along the diagonal, in whose direction the mollifier is constant. In effect the argument tests the difference of two solutions against a function that is smooth along the diagonal and singular across it, and only the diagonal direction is ever differentiated.
The limit and the cone estimate
For every non-negative \(\chi\) smooth with compact support in \((0,\infty)\times\R\),
that is, \(\pp_{t}\abs{u-v} + \pp_{x}\Phi(u,v) \le 0\) in the sense of distributions on \((0,\infty)\times\R\). Rests on Proposition A.160 and Lemma A.157.
Derives Proposition A.162. Insert Equation (A.337) into Equation (A.333) and change variables to Equation (A.336). The map \((t,x,s,y)\longmapsto(\tau,\zeta,\sigma,\omega)\) is linear with \(\dd t\,\dd s = 2\,\dd\tau\,\dd\sigma\) and \(\dd x\,\dd y = 2\,\dd\zeta\,\dd\omega\), so
after dividing by the positive constant \(4\), where
Let \(K\) be a compact neighbourhood of the support of \(\chi\) in \((0,\infty)\times\R\). By continuity of translation in \(L^{1}\) (item 4 of Remark A.153) both \((\tau,\zeta)\longmapsto u(\tau+\sigma,\zeta+\omega)\) and \((\tau,\zeta)\longmapsto v(\tau-\sigma,\zeta-\omega)\) converge to \(u\) and \(v\) in \(L^{1}(K)\) as \((\sigma,\omega)\to(0,0)\); hence, the maps \((a,b)\longmapsto\abs{a-b}\) and \((a,b)\longmapsto\Phi(a,b)\) being Lipschitz by Equation (A.332),
Denote by \(\Xi(\sigma,\omega)\) the inner bracket of Equation (A.339) integrated against \(\chi\)'s derivatives, i.e. the quantity \(\iint\bigl(G_{\sigma\omega}\pp_{\tau}\chi + H_{\sigma\omega}\pp_{\zeta}\chi\bigr)\dd\zeta\,\dd\tau\), and by \(\Xi_{0}\) the left-hand side of Equation (A.338). Then \(\abs{\Xi(\sigma,\omega)-\Xi_{0}} \le \norm{\nabla\chi}_{\infty}\int_{K} \bigl(\abs{G_{\sigma\omega}-\abs{u-v}} + \abs{H_{\sigma\omega}-\Phi(u,v)}\bigr)\), which by Equation (A.341) tends to \(0\) as \((\sigma,\omega)\to0\). But Equation (A.339) says exactly \(\iint\Xi(\sigma,\omega)\theta_{\delta}(\sigma)\theta_{\delta}(\omega) \ge 0\), an average of \(\Xi\) over the square \(\abs{\sigma},\abs{\omega}<\delta\) with a non-negative weight of total mass \(1\); as \(\delta\to0\) that average converges to \(\Xi_{0}\). Hence \(\Xi_{0}\ge0\), which is Equation (A.338).
∎Proof of Theorem A.151. Derives Theorem A.151. Write \(w = \abs{u-v}\) and \(\Psi = \Phi(u,v)\), so that Equation (A.338) holds and \(\abs{\Psi} \le M\,w\) by Equation (A.332).
The cone cutoff. Fix \(R>0\) and \(T>0\). For \(\epsilon>0\) let \(\Lambda_{\epsilon} \in C^{\infty}(\R)\) be non-decreasing with \(\Lambda_{\epsilon} = 0\) on \((-\infty,0]\), \(\Lambda_{\epsilon}=1\) on \([\epsilon,\infty)\) and \(0 \le \Lambda_{\epsilon}' \le 2/\epsilon\); and let \(n_{\epsilon} \in C^{\infty}(\R)\) be even with \(\abs{n_{\epsilon}'} \le 1\) and \(\abs{n_{\epsilon}(x)-\abs{x}} \le \epsilon\) — for instance \(n_{\epsilon} = \sqrt{x^{2}+\epsilon^{2}}\). Put
a smooth function with values in \([0,1]\), equal to \(1\) well inside the backward cone \(\set{n_{\epsilon}(x) \le R+M(T-t)}\) and to \(0\) outside it. Let \(\alpha \ge 0\) be smooth with compact support in \((0,T]\) and take \(\chi = \alpha\beta\), which is admissible in Equation (A.338) after multiplication by a fixed spatial cutoff equal to \(1\) on the (bounded) support of \(\beta\) — the cone is bounded for \(t \ge 0\) — so that \(\chi\) has compact support.
The lateral term has a sign. Since \(\pp_{t}\chi = \alpha'\beta + \alpha\,\pp_{t}\beta\) and \(\pp_{x}\chi = \alpha\,\pp_{x}\beta\), with
the contribution of the second group to Equation (A.338) is
because \(\alpha \ge 0\), \(\Lambda_{\epsilon}'\ge0\), and \(M w + \Psi n_{\epsilon}' \ge Mw - \abs{\Psi} \ge 0\) by \(\abs{n_{\epsilon}'}\le1\) and \(\abs{\Psi}\le Mw\). This is the point of choosing the slope of the cone to be \(M\): the flux can never carry mass out through the lateral boundary faster than the boundary retreats. Subtracting Equation (A.344) from Equation (A.338) leaves
From an average to a value. Let \(0<\tau_{1}<\tau_{2}\le T\) and let \(\alpha = \alpha_{h}\) rise smoothly from \(0\) to \(1\) on \([\tau_{1},\tau_{1}+h]\) and fall back to \(0\) on \([\tau_{2},\tau_{2}+h]\), with \(\alpha_{h}'\) bounded by \(2/h\) and of one sign on each interval. Then Equation (A.345) reads, in the limit of the standard smooth ramps,
up to an error tending to \(0\) with the width of the ramps. \(\Theta\) is bounded and measurable, so by Lebesgue's differentiation theorem (item 5 of Remark A.153) the two averages converge, as \(h\to0\), to \(\Theta(\tau_{1})\) and \(\Theta(\tau_{2})\) for almost every \(\tau_{1},\tau_{2}\). Hence \(\Theta(\tau_{2}) \le \Theta(\tau_{1})\) for almost all \(\tau_{1}<\tau_{2}\).
Removing the smoothing. As \(\epsilon\to0\), \(\beta(t,x)\) tends to \(1\) at every \(x\) with \(\abs{x} < R+M(T-t)\) and to \(0\) at every \(x\) with \(\abs{x} > R+M(T-t)\), and \(0\le\beta\le1\) throughout; by dominated convergence, applied with the integrable majorant \(2M_{0}\,\indic{\abs{x}\le R+MT}\),
Choosing \(T = \tau_{2}\) turns Equation (A.347) into Equation (A.322).
Down to \(t=0\). By Lemma A.154 applied to \(u\) and to \(v\) on the compact set \(\abs{x}\le R+M\tau_{2}\), \(w(\tau_{1},\cdot)\longrightarrow\abs{u_{0}-v_{0}}\) in the mean over that set as \(\tau_{1}\to0^{+}\) — the triangle inequality gives \(\bigl|\,\abs{u(\tau_{1})-v(\tau_{1})} - \abs{u_{0}-v_{0}}\,\bigr| \le \abs{u(\tau_{1})-u_{0}} + \abs{v(\tau_{1})-v_{0}}\) — so the right-hand side of Equation (A.322) converges to \(\int_{\abs{x}\le R+M\tau_{2}}\abs{u_{0}-v_{0}}\) along a sequence of admissible \(\tau_{1}\), which proves the version with \(\tau_{1}=0\).
The global statement. If \(u_{0}-v_{0}\) is integrable, let \(R\to\infty\) in that version: the left-hand side increases to \(\int_{\R}w(t,x)\,\dd x\) by monotone passage on an increasing family of sets, and the right-hand side increases to \(\int_{\R}\abs{u_{0}-v_{0}}\), giving Equation (A.323).
∎Existence by vanishing viscosity
The uniqueness half is now complete. Existence is the limit \(\varepsilon \to 0^{+}\) of the viscous problem, and the one step that is genuinely about entropy — and the only one that would be invisible in a purely functional-analytic account — is carried out here in full.
Let \(u^{\varepsilon}\) be a bounded classical solution of
with \(\abs{u^{\varepsilon}} \le M_{0}\), and let \((\eta,q)\) be an entropy pair in the sense of Definition 10.101. Then for every non-negative test function \(\varphi\)
with \(C_{\varphi} = \max_{\abs{w}\le M_{0}}\abs{\eta(w)}\cdot \iint\abs{\pp_{x}^{2}\varphi}\), a constant independent of \(\varepsilon\). Rests on Definition 10.101, Equation (10.81) and Theorem 10.80.
Derives Proposition A.163. Multiply Equation (A.348) by \(\eta'(u^{\varepsilon})\). On the left, the chain rule and \(q' = \eta'F'\) give \(\eta'(u^{\varepsilon})\pp_{t}u^{\varepsilon} = \pp_{t}\eta(u^{\varepsilon})\) and \(\eta'(u^{\varepsilon})\pp_{x}F(u^{\varepsilon}) = \eta'(u^{\varepsilon})F'(u^{\varepsilon})\pp_{x}u^{\varepsilon} = \pp_{x}q(u^{\varepsilon})\). On the right,
the identity being the chain rule applied twice and the inequality being \(\eta'' \ge 0\), which is the convexity demanded by Definition 10.101 and the only place it is used. Hence
pointwise. Multiply Equation (A.351) by \(\varphi \ge 0\), integrate over \(\set{t>0}\times\R\) and integrate by parts, all boundary terms at spatial infinity vanishing with \(\varphi\) and the one at \(t=0\) producing \(-\int\eta(u_{0}^{\varepsilon})\varphi(x,0)\,\dd x\):
and the right-hand side is bounded in modulus by \(\varepsilon C_{\varphi}\), which is Equation (A.349). The two integrations by parts on the viscous term are the reason the estimate is \(O(\varepsilon)\) and not \(O(1)\): all the differentiation has been moved onto the test function, and no derivative of \(u^{\varepsilon}\) survives.
∎Proposition A.163 is the whole of the entropy argument; what it does not supply is a limit to take it to. Three further steps are needed, of which the first is elementary in this treatise and the other two are the quoted items (7) of Remark A.153.
-
The supremum bound. \(\abs{u^{\varepsilon}} \le \norm{u_{0}^{\varepsilon}}_{\infty} \le \norm{u_{0}}_{\infty}\), by the parabolic maximum principle Theorem 10.80 applied to Equation (A.348) on large space-time rectangles: the equation is of the form treated there with the extra first-order term \(F'(u^{\varepsilon})\pp_{x}u^{\varepsilon}\), which vanishes at an interior extremum and so does not disturb the argument. This is why \(M_{0}\) and hence \(M\) in Equation (A.320) can be fixed independently of \(\varepsilon\).
-
Classical solvability of Equation (A.348) for smooth bounded data \(u_{0}^{\varepsilon} = u_{0}\) mollified, on \(\set{t>0}\times\R\). This is parabolic theory and is quoted.
-
Compactness. The spatial derivative \(p = \pp_{x}u^{\varepsilon}\) satisfies the linear parabolic equation \(\pp_{t}p + \pp_{x}\bigl(F'(u^{\varepsilon})p\bigr) = \varepsilon\pp_{x}^{2}p\), from which the total variation \(\int\abs{p(t,x)}\,\dd x\) is non-increasing in \(t\); together with (i) this gives a family uniformly bounded and of uniformly bounded variation, and a compactness theorem for such families produces a sequence \(\varepsilon_{j}\to0\) with \(u^{\varepsilon_{j}}\to u\) in \(L^{1}\) on compact sets and almost everywhere. The compactness theorem is quoted.
Granting these, pass to the limit in Equation (A.349): \(\eta\) and \(q\) are continuous and \(u^{\varepsilon_{j}}\to u\) almost everywhere with values in a fixed bounded interval, so dominated convergence carries both integrals, the data converge because \(u_{0}^{\varepsilon}\to u_{0}\) in \(L^{1}\) on compact sets, and the right-hand side tends to \(0\). The limit is exactly Equation (10.92). Taking \(\eta(w)=w\) and \(q = F\) — a legitimate, if degenerate, member of Definition 10.101, for which Equation (A.350) is an equality — gives the same statement with equality, which is Equation (10.85): the limit is a weak solution.
Proof of Theorem 10.103. Derives Theorem 10.103. Existence is Remark A.164: the vanishing-viscosity limit is a weak solution in the sense of Definition 10.95 satisfying Equation (10.92) for every entropy pair, hence an entropy solution in the sense of Definition 10.101. Uniqueness and continuous dependence in the mean are Corollary A.152, which rests on Theorem A.151. The last clause of the chapter's statement — that where the solution is piecewise \(C^{1}\) across a \(C^{1}\) curve its jumps obey Equation (10.91) — is Proposition 10.102, proved in the chapter, and is not re-derived here.
∎Theorem 10.103 assumes \(F\) convex, and the chapter has good reason to: convexity is what makes Proposition 10.102 equate the entropy inequality with the geometric condition Equation (10.91), and it is what makes the shock picture of Section 10.6.3 the whole story. But the proof given above never used it. Only \(F \in C^{1}\) entered, through the bound \(M\) of Equation (A.320) and the Lipschitz estimate Equation (A.332). What is lost without convexity is not uniqueness but the description of the admissible discontinuities: for a non-convex flux the entropy inequality still selects one solution, but the admissible jumps are described by a chord condition — the chord joining the two states must lie on one side of the graph of \(F\) between them — rather than by the simple inequality \(u_{-} \ge u_{+}\). The function \(G\) of Equation (10.94) is exactly that chord difference, and the chapter's proof of Proposition 10.102 is where convexity is spent: it is used to pass from \(G(u_{-}) = G(u_{+}) = 0\) to \(G \le 0\) between the two states, and that implication is what fails when \(F\) is not convex.
The chapter already observes, in the paragraph following Definition 10.101, that Equation (10.91) presupposes one-sided limits that a bounded measurable weak solution need not have. The proof above shows what is gained by not presupposing them: at no point did the argument speak of a discontinuity, a curve, or a one-sided limit. It compared two solutions through their entropy inequalities alone, and the only regularity used — continuity of translation in the mean — is a property every bounded measurable function has. Example 10.98 is settled by the same stroke without any examination of the two candidate solutions: whichever of them satisfies Equation (10.92) is the only one that does.
Kruzhkov's Doubling of Variables and the Entropy Solution discharges the proof obligation of Theorem 10.103 in Partial Differential Equations, and with it closes the pattern that Remark 10.104 describes: a nonlinear equation destroys its own classical solution in finite time (Proposition 10.94), weak solutions restore existence at the cost of uniqueness (Example 10.98), and the entropy inequality Equation (10.92) restores uniqueness, giving back all three of Hadamard's conditions (Definition 10.24) with the mean as the norm. The physical reading is Remark 10.100, and the gas-dynamical form of the admissible jump is Fluid Dynamics. Six inputs are assumed and are listed in Remark A.153; five of them are the Lebesgue theory that this treatise deliberately does not build, and the sixth is Lemma A.154.
Lévy's Continuity Theorem
This appendix proves Theorem 11.45 of Probability and Statistics: if the characteristic functions of a sequence of random variables converge pointwise to a limit that is continuous at the origin, then that limit is itself a characteristic function and the variables converge in distribution. It is the bridge the central limit theorem crosses — Theorem 11.46 produces \(\varphi_{Z_{n}}(t)\rightarrow\ee^{-t^{2}/2}\) and needs to conclude \(Z_{n}\rightarrow\mathcal{N}(0,1)\) — and it is used again in The Lindeberg–Feller Central Limit Theorem and The Berry–Esseen Inequality.
The route has four stages, each of which is a theorem in its own right: Helly's selection theorem, which extracts a limiting monotone function from any sequence of distribution functions (Helly's selection theorem); the Fourier inversion formula, which recovers the distribution from the transform and thereby shows the transform determines it (The inversion formula); the tightness estimate, which is the single place continuity of the limit at the origin is used (Tightness from continuity at the origin); and the assembly (Proof of the theorem). Everything rests on one classical integral, Dirichlet's, which is evaluated first because the inversion formula is built out of it.
Throughout, \(F_{X}(x)=\Pr(X\le x)\) is the distribution function of Definition 11.21 — non-decreasing, right-continuous, with limits \(0\) and \(1\) at \(\mp\infty\), and not assumed continuous — and \(\varphi_{X}(t)=\avg{\ee^{\ii tX}}\) the characteristic function of Definition 11.42. Convergence in distribution means \(F_{X_{n}}(x)\rightarrow F_{X}(x)\) at every continuity point \(x\) of \(F_{X}\). All the quantities here are pure numbers: \(t\) carries the reciprocal of the SI unit of \(X\), so that \(tX\), and with it every exponent below, is dimensionless.
Statement
Let \(X_{1},X_{2},\dots\) be real random variables with characteristic functions \(\varphi_{n}\). Suppose \(\varphi_{n}(t)\rightarrow\varphi(t)\) for every \(t\in\R\), and suppose the limit function \(\varphi\) is continuous at \(t=0\). Then \(\varphi\) is the characteristic function of a random variable \(X\), and \(X_{n}\rightarrow X\) in distribution. Rests on Definitions 11.21 and 11.42.
The converse is elementary and is recorded at the end (Corollary A.181); it is the direction the theorem is not usually needed in.
The two facts imported from Lebesgue theory
Real Analysis builds the Riemann–Darboux integral (Definition 7.39) and no more. Two theorems of the Lebesgue theory are used below and are not proved anywhere in this treatise; they are stated here in exactly the form in which they are applied, so that the reader can see precisely how much is being assumed.
Let \(Z,Z_{1},Z_{2},\dots\) be random variables with \(Z_{m}\rightarrow Z\) pointwise (or in probability), and suppose there is a random variable \(D\) with \(\avg{D}<\infty\) and \(\abs{Z_{m}}\le D\) for every \(m\). Then \(\avg{Z_{m}}\rightarrow\avg{Z}\). The same statement holds for a family indexed by a continuous parameter. In particular, if \(\abs{Z_{m}}\le c\) for a constant \(c\), then \(\avg{Z_{m}}\rightarrow \avg{Z}\) [Billingsley:1995]. Rests on Definition 11.4.
Let \(h(t,\omega)\) be jointly measurable, let \(\abs{h}\le c\) for a constant \(c\), and let \(T<\infty\). Then
both sides being finite [Billingsley:1995]. Rests on Definition 11.4.
Theorems A.169 and A.170 are the only statements in this appendix that are assumed rather than derived. Both belong to measure theory: dominated convergence is the theorem that makes the Lebesgue integral complete under limits — the property the Riemann integral lacks, which is why it cannot be proved from Real Analysis — and Fubini's theorem is the corresponding statement for a product measure. This treatise develops neither, by the editorial choice recorded at Lemma 11.44, where the same gap was closed by an explicit truncation instead. No such elementary substitute is available here: the interchange Equation (A.352) is applied to the expectation of an oscillatory integral whose integrand depends on the random variable through \(\ee^{\ii tX}\), and there is nothing to truncate. Each use below is flagged by name at the point where it occurs; there are four in all (two of each). Everything else — Helly's theorem, the Helly–Bray lemma, the Dirichlet integral, the inversion formula, the tightness estimate — is derived here from the real analysis of Real Analysis.
The Dirichlet integral
Write \(S(Y)=\int_{0}^{Y}\left(\sin s\right)/s\,\dd s\) for \(Y>0\), the integrand being extended by the value \(1\) at \(s=0\), where it is continuous. Then
Consequently, for every \(u\in\R\) and \(T>0\), \(\int_{0}^{T}\left(\sin(tu)\right)/t\,\dd t=S(Tu)\) for \(u>0\), \(=-S(T\abs{u})\) for \(u<0\) and \(=0\) for \(u=0\); so the integral is bounded by \(C_{0}\) uniformly in \(T\) and \(u\) and tends to \(\tfrac{\pi}{2}\sgn(u)\) as \(T\rightarrow\infty\). Rests on Theorem 7.43 and Equation (7.28).
Derives Lemma A.172. The uniform bound. On \([0,1]\) the integrand is bounded by \(1\), so \(\abs{\int_{0}^{1}}\le1\). For \(Y\ge1\) integrate by parts (Equation (7.28)) with \(\sin s\,\dd s=-\dd\cos s\):
whose first term is at most \(1/Y+\cos1\le2\) in modulus and whose second is at most \(\int_{1}^{\infty}s^{-2}\dd s=1\). Hence \(\abs{S(Y)}\le1+2+1=4\) for \(Y\ge1\), and \(\abs{S(Y)}\le1\) for \(Y\le1\). The same estimate applied on \([Y,Y']\) with \(1\le Y<Y'\) gives
so \(S(Y)\) satisfies the Cauchy criterion (Theorem 7.8) and the limit \(D=\lim_{Y\rightarrow\infty}S(Y)\) exists.
The value. For \(\lambda>0\) put \(I(\lambda)=\int_{0}^{\infty}\ee^{-\lambda s}\left(\sin s\right)/s\,\dd s\), which converges absolutely since the integrand is bounded by \(\ee^{-\lambda s}\). Its \(\lambda\)-derivative is \(-\ee^{-\lambda s}\sin s\) in modulus at most \(\ee^{-\lambda_{0}s}\) for \(\lambda\ge\lambda_{0}>0\), an integrable bound independent of \(\lambda\), so the difference quotients converge uniformly on \([\lambda_{0},\infty)\) and differentiation under the integral sign is legitimate there:
Also \(\abs{I(\lambda)}\le\int_{0}^{\infty}\ee^{-\lambda s}\dd s=1/\lambda\rightarrow0\). Integrating Equation (A.356) from \(\lambda\) to \(\infty\) therefore gives \(I(\lambda)=\tfrac{\pi}{2}-\arctan\lambda\).
The two agree in the limit. Fix \(Y\ge1\) and estimate \(I(\lambda)-D\) in three pieces. On \([0,Y]\), using \(\abs{\ee^{-\lambda s}-1}\le\lambda s\) and \(\abs{\sin s}\le1\),
The tail of \(D\) is at most \(2/Y\) by Equation (A.355). For the tail of \(I\), integrate by parts with \(u=\ee^{-\lambda s}/s\):
the middle step bounding \(1/s\le1/Y\) in the \(\lambda\)-term and \(\ee^{-\lambda s}\le1\) in the other. Hence \(\abs{I(\lambda)-D}\le\lambda Y+5/Y\) for every \(Y\ge1\) and \(\lambda>0\). Given \(\varepsilon>0\) choose \(Y=10/\varepsilon\) and then \(\lambda<\varepsilon/(2Y)\): the right side is below \(\varepsilon\). So \(D=\lim_{\lambda\downarrow0}I(\lambda)=\tfrac{\pi}{2}-\arctan0 =\tfrac{\pi}{2}\).
The final statement is the substitution \(s=tu\) (Equation (7.27)), under which \(\left(\sin(tu)\right)/t\,\dd t=\left(\sin s\right)/s\,\dd s\) for \(u>0\); oddness of \(\sin\) in \(u\) supplies the case \(u<0\).
∎Helly's selection theorem
Let \(F_{1},F_{2},\dots\) be distribution functions. Then there are a subsequence \(F_{n_{1}},F_{n_{2}},\dots\) and a non-decreasing, right-continuous \(F:\R\rightarrow[0,1]\) such that \(F_{n_{k}}(x)\rightarrow F(x)\) at every continuity point \(x\) of \(F\). Rests on Definition 11.21 and Theorem 7.7.
The limit \(F\) need not be a distribution function: mass can escape to infinity, as it does for \(F_{n}=F_{X+n}\), where \(F\equiv0\). Excluding that is precisely what tightness will do.
Derives Theorem A.173. The diagonal argument. Enumerate the rationals, \(\Q=\set{q_{1},q_{2},\dots}\). The numbers \(F_{n}(q_{1})\) lie in the bounded set \([0,1]\), so by Bolzano–Weierstrass (Theorem 7.7) some subsequence, indexed by an infinite set \(N_{1}\subseteq\N\), has \(F_{n}(q_{1})\) convergent along \(N_{1}\); call the limit \(G(q_{1})\). Having chosen \(N_{1}\supseteq N_{2}\supseteq\dots\supseteq N_{m}\) with \(F_{n}(q_{i})\rightarrow G(q_{i})\) along \(N_{m}\) for every \(i\le m\), apply Theorem 7.7 again to the bounded sequence \(\left(F_{n}(q_{m+1})\right)_{n\in N_{m}}\) to get \(N_{m+1}\subseteq N_{m}\) and a limit \(G(q_{m+1})\). Let \(n_{k}\) be the \(k\)th element of \(N_{k}\), in increasing order and chosen larger than \(n_{k-1}\), which is possible because each \(N_{k}\) is infinite. For each fixed \(i\), all but the first \(i-1\) terms of \((n_{k})\) lie in \(N_{i}\), so
Each \(F_{n}\) is non-decreasing with values in \([0,1]\), so \(G\) inherits both properties on \(\Q\).
The limit function. Define \(F(x)=\inf\set{G(q)\mid q\in\Q,\ q>x}\). It is non-decreasing and \([0,1]\)-valued because \(G\) is. It is right-continuous: given \(x\) and \(\varepsilon>0\), pick a rational \(q>x\) with \(G(q)<F(x)+\varepsilon\); every \(y\in(x,q)\) then has \(F(y)\le G(q)<F(x)+\varepsilon\), and monotonicity gives \(F(y)\downarrow F(x)\) as \(y\downarrow x\).
Two inequalities are used repeatedly and follow at once from the definition and the monotonicity of \(G\): for rationals \(r<s\),
The first holds because \(s>r\) makes \(G(s)\) one of the numbers whose infimum is \(F(r)\); the second because every rational \(q>s\) has \(G(q)\ge G(r)\), so the infimum defining \(F(s)\) is at least \(G(r)\).
Convergence at continuity points. Let \(x\) be a continuity point of \(F\) and let \(\varepsilon>0\). By left-continuity of \(F\) at \(x\) there is \(y<x\) with \(F(y)>F(x)-\varepsilon\); choose rationals \(r,s\) with \(y<r<s<x\). Then \(F_{n_{k}}(x)\ge F_{n_{k}}(s)\rightarrow G(s)\), and by Equation (A.358) twice, \(G(s)\ge G(r)\ge F(y)\), whence
For the other side, continuity gives \(z>x\) with \(F(z)<F(x)+\varepsilon\); choose a rational \(q\in(x,z)\). Then \(F_{n_{k}}(x)\le F_{n_{k}}(q)\rightarrow G(q)\), and \(G(q)\le F(z)\) by the second inequality of Equation (A.358), so \(\limsup_{k}F_{n_{k}}(x)\le F(z)<F(x)+\varepsilon\). As \(\varepsilon>0\) was arbitrary, \(F_{n_{k}}(x)\rightarrow F(x)\).
∎A family \(\set{X_{n}}\) of random variables is tight when for every \(\varepsilon>0\) there is \(K<\infty\) with \(\Pr(\abs{X_{n}}>K)\le\varepsilon\) for every \(n\). Rests on Definition 11.21.
If in Theorem A.173 the underlying variables are tight, then the limit \(F\) satisfies \(F(x)\rightarrow0\) as \(x\rightarrow-\infty\) and \(F(x)\rightarrow1\) as \(x\rightarrow+\infty\): it is a distribution function. Rests on Theorem A.173 and Definition A.174.
Derives Lemma A.175. Given \(\varepsilon>0\) take \(K\) as in Definition A.174. A monotone function has at most countably many discontinuities — the open intervals \(\left(F(x^{-}),F(x^{+})\right)\) belonging to distinct jumps are disjoint and each contains a rational — so continuity points are dense and we may pick continuity points \(x>K\) and \(x'<-K\). Then \(F(x)=\lim_{k}F_{n_{k}}(x)\ge\liminf_{k}\Pr(\abs{X_{n_{k}}}\le K) \ge1-\varepsilon\) and \(F(x')=\lim_{k}F_{n_{k}}(x')\le\varepsilon\). Monotonicity extends both to all larger, respectively smaller, arguments.
∎Let \(F_{n},F\) be distribution functions with \(F_{n}(x)\rightarrow F(x)\) at every continuity point of \(F\), and let \(g:\R\rightarrow\R\) be bounded and continuous. Then
where \(X_{n},X\) have distribution functions \(F_{n},F\). Rests on Theorems 7.25 and A.173.
Derives Lemma A.176. Write \(\norm{g}_{\infty}=\sup\abs{g}\) and let \(\varepsilon>0\). Since \(F\) is a distribution function and its continuity points are dense, choose continuity points \(a<b\) with \(F(a)<\varepsilon\) and \(1-F(b)<\varepsilon\); then for \(n\) large \(F_{n}(a)<2\varepsilon\) and \(1-F_{n}(b)<2\varepsilon\). The contributions to \(\avg{g(X_{n})}\) and \(\avg{g(X)}\) from outside \((a,b]\) are therefore at most \(2\varepsilon\norm{g}_{\infty}\) and \(\varepsilon\norm{g}_{\infty}\) in modulus.
On the compact interval \([a,b]\) the function \(g\) is uniformly continuous (Theorem 7.25), so there is \(\delta>0\) with \(\abs{g(x)-g(y)}<\varepsilon\) whenever \(\abs{x-y}<\delta\) inside \([a,b]\). Choose continuity points \(a=x_{0}<x_{1}<\dots<x_{m}=b\) of \(F\) with \(x_{j}-x_{j-1}<\delta\) — possible because continuity points are dense. Then for any distribution function \(H\),
because on each \((x_{j-1},x_{j}]\) the integrand differs from \(g(x_{j})\) by less than \(\varepsilon\) and the increments of \(H\) sum to at most \(1\). Applying this to \(H=F_{n}\) and to \(H=F\), and using \(F_{n}(x_{j})\rightarrow F(x_{j})\) at each of the finitely many partition points, gives
Since \(\varepsilon>0\) was arbitrary, the limit is \(0\).
∎The inversion formula
Let \(X\) have distribution function \(F\) and characteristic function \(\varphi\), and let \(a<b\). Then
the integrand being extended by its limit \(b-a\) at \(t=0\). In particular, if \(a\) and \(b\) are continuity points of \(F\) the right-hand side is \(F(b)-F(a)\). Rests on Definition 11.42 and Lemma A.172.
Derives Theorem A.177. Write \(J_{T}\) for the integral on the left. For real \(u,v\),
so with \(u=X-a\), \(v=X-b\) the integrand of \(J_{T}\), written as \(\left(\ee^{\ii t(X-a)}-\ee^{\ii t(X-b)}\right)/(\ii t)\) under the expectation, is bounded by the constant \(b-a\). Fubini's theorem (Theorem A.170) — the first of the two quoted uses — therefore applies:
Now \(\ee^{\ii tu}/(\ii t)=\left(\sin tu\right)/t -\ii\left(\cos tu\right)/t\), and the cosine terms are odd in \(t\), so they cancel over the symmetric interval. Hence the inner integral is
which by Lemma A.172 is bounded by \(4C_{0}\le16\) uniformly in \(T\) and in \(X\), and converges as \(T\rightarrow\infty\) to \(\pi\left[\sgn(X-a)-\sgn(X-b)\right]\) pointwise. Dominated convergence (Theorem A.169) — the second quoted use — gives
The bracket takes the value \(0\) when \(X<a\) or \(X>b\), the value \(2\) when \(a<X<b\), and the value \(1\) when \(X=a\) or \(X=b\). So
which, using \(F(x^{-})=\Pr(X<x)\) and \(F(x)=\Pr(X\le x)\), is exactly the right-hand side of Equation (A.360).
∎If two random variables have the same characteristic function, they have the same distribution function. Rests on Theorem A.177.
Derives Corollary A.178. Let \(F_{1},F_{2}\) have the common characteristic function \(\varphi\). The left-hand side of Equation (A.360) depends only on \(\varphi\), so \(F_{1}(b)-F_{1}(a)=F_{2}(b)-F_{2}(a)\) whenever \(a,b\) are continuity points of both. Continuity points of either are dense (the argument of Lemma A.175), so the continuity points common to both are dense as well, their complement being a union of two countable sets. Letting \(a\rightarrow-\infty\) through such points gives \(F_{1}(b)=F_{2}(b)\) at every common continuity point \(b\); and every \(x\in\R\) is the limit from the right of such points, so right-continuity of both functions gives \(F_{1}=F_{2}\) everywhere.
∎Tightness from continuity at the origin
For any random variable \(X\) with characteristic function \(\varphi_{X}\) and any \(u>0\),
Rests on Definition 11.42.
Derives Lemma A.179. The integrand \(1-\cos(tX)\) is bounded by \(2\), so Fubini's theorem (Theorem A.170) — the third quoted use — allows the expectation and the \(t\)-integral to be exchanged:
the inner integral being elementary and the quotient being read as \(2\) when \(X=0\). Put \(y=uX\). The function \(1-\left(\sin y\right)/y\) is non-negative for every real \(y\), since \(\abs{\sin y}\le\abs{y}\); and when \(\abs{y}\ge2\) the bound \(\abs{\sin y}\le1\) gives \(1-\left(\sin y\right)/y\ge1-1/\abs{y}\ge\tfrac12\), so that \(2\left(1-\left(\sin y\right)/y\right)\ge1\) there. Hence the integrand of the expectation dominates \(\indic{\abs{uX}\ge2}\) pointwise, and taking expectations gives Equation (A.365).
∎Let \(\varphi_{n}\rightarrow\varphi\) pointwise on \(\R\) with \(\varphi\) continuous at \(0\). Then the corresponding family \(\set{X_{n}}\) is tight. Rests on Lemma A.179 and Definition A.174.
Derives Lemma A.180. First, \(\varphi(0)=\lim_{n}\varphi_{n}(0)=1\) by Proposition 11.43(i). Let \(\varepsilon>0\). Continuity of \(\varphi\) at \(0\) gives \(u>0\) with \(\abs{1-\varphi(t)}<\varepsilon/4\) for \(\abs{t}\le u\), whence
The functions \(1-\Re\varphi_{n}\) are bounded by \(2\) on \([-u,u]\) and converge pointwise to \(1-\Re\varphi\), so dominated convergence (Theorem A.169) — the fourth and last quoted use, here applied to the uniform measure on \([-u,u]\) — gives
so there is \(N\) with the left side below \(\varepsilon\) for \(n\ge N\). By Lemma A.179, \(\Pr(\abs{X_{n}}\ge2/u)\le \varepsilon\) for \(n\ge N\). Each of the finitely many remaining variables \(X_{1},\dots,X_{N-1}\) is individually tight, its distribution function tending to \(0\) and \(1\) at the two ends, so there is \(K_{0}\) with \(\Pr(\abs{X_{n}}>K_{0})\le\varepsilon\) for \(n<N\). Take \(K=\max(2/u,K_{0})\).
∎This lemma is the whole reason the hypothesis of Theorem A.168 mentions the origin, and it shows the hypothesis cannot be dropped: the Cauchy-type variables \(X_{n}=nY\) with \(Y\) standard Cauchy have \(\varphi_{n}(t)=\ee^{-n\abs{t}}\), which converges pointwise to the function equal to \(1\) at \(t=0\) and \(0\) elsewhere — not continuous at the origin — and the mass escapes to infinity, the distributions converging to nothing.
Proof of the theorem
Proof of Theorem A.168. Derives Theorem A.168. By Lemma A.180 the family \(\set{X_{n}}\) is tight.
Every subsequence has a convergent further subsequence with the same limit. Let \((n_{j})\) be any subsequence. By Helly's theorem (Theorem A.173) it has a further subsequence \((n_{j_{k}})\) along which \(F_{n_{j_{k}}}\rightarrow F\) at every continuity point of some non-decreasing right-continuous \(F:\R\rightarrow[0,1]\); by Lemma A.175 and tightness, \(F\) is a genuine distribution function. Let \(X\) be a random variable with distribution function \(F\). Applying the Helly–Bray lemma (Lemma A.176) to the bounded continuous functions \(x\mapsto\cos(tx)\) and \(x\mapsto\sin(tx)\), for each fixed \(t\), gives
But by hypothesis \(\varphi_{n_{j_{k}}}(t)\rightarrow\varphi(t)\). Limits in \(\C\) are unique, so \(\varphi=\varphi_{X}\): the limit function is a characteristic function, which is the first assertion of the theorem. Moreover, by Corollary A.178 the distribution function \(F\) is determined by \(\varphi\) alone, so every subsequential Helly limit obtained this way is the same function \(F\).
Hence the whole sequence converges. Let \(x\) be a continuity point of \(F\) and suppose \(F_{n}(x)\) does not converge to \(F(x)\). Then there are \(\eta>0\) and a subsequence \((n_{j})\) with \(\abs{F_{n_{j}}(x)-F(x)}\ge\eta\) for every \(j\). By the previous paragraph applied to \((n_{j})\) there is a further subsequence along which the distribution functions converge, at every continuity point of \(F\) and in particular at \(x\), to \(F\) — contradicting the standing inequality. Therefore \(F_{n}(x)\rightarrow F(x)\) at every continuity point of \(F\), which is \(X_{n}\rightarrow X\) in distribution.
∎If \(X_{n}\rightarrow X\) in distribution then \(\varphi_{n}(t)\rightarrow\varphi_{X}(t)\) for every \(t\), and the limit is continuous everywhere. Rests on Lemma A.176 and Theorem A.168.
Derives Corollary A.181. The first claim is Equation (A.366) read for the full sequence, the Helly–Bray lemma applying directly since \(F_{n}\rightarrow F_{X}\) at continuity points by hypothesis. For continuity of \(\varphi_{X}\): Equation (A.361) read with \(u=X\), \(v=0\) and \(t=h\) gives \(\abs{\ee^{\ii hX}-1}\le\abs{h}\abs{X}\), while the modulus of a difference of two unit complex numbers is at most \(2\), so \(\abs{\ee^{\ii(t+h)X}-\ee^{\ii tX}}=\abs{\ee^{\ii hX}-1} \le\min(\abs{hX},2)\) and hence \(\abs{\varphi_{X}(t+h)-\varphi_{X}(t)}\le\avg{\min(\abs{hX},2)}\), a bound independent of \(t\) that tends to \(0\) with \(h\) — by the same truncation as in Lemma 11.44: split the expectation at \(\abs{X}=K\), bound the inner part by \(\abs{h}K\) and the outer by \(2\Pr(\abs{X}>K)\), and choose \(K\) then \(h\). So \(\varphi_{X}\) is in fact uniformly continuous.
∎Lévy's Continuity Theorem discharges the proof obligation of Theorem 11.45, the one analytic input that Section 11.3.1 of Probability and Statistics had to quote. With it, the derivation of the central limit theorem Theorem 11.46 in Section 11.3.2 is complete from the axioms of the chapter: Lemma 11.44 supplies the second-order expansion, Equation (11.36) the telescoping estimate, and this appendix the passage back from transforms to distributions. The same passage is the last step of The Lindeberg–Feller Central Limit Theorem, and Theorem A.177 is the starting point of The Berry–Esseen Inequality. Two theorems of Lebesgue theory are assumed and are named in Remark A.171; nothing else is.
The Lindeberg–Feller Central Limit Theorem
This appendix proves Theorem 11.48 of Probability and Statistics: a triangular array of independent, individually negligible summands whose tails satisfy Lindeberg's condition has an asymptotically Gaussian normalised sum. It is the form of the central limit theorem the error theory of Measurement, SI Units, and the Theory of Errors actually needs, because the disturbances that perturb a measurement — a thermal drift, a vibration, a quantisation step, a reading error — are independent but certainly not identically distributed, so the Lindeberg–Lévy theorem Theorem 11.46 does not apply to them.
The argument is the characteristic-function argument of Section 11.3.2 carried out uniformly in the summand index. Three of its four ingredients are already in the chapter: the elementary properties of characteristic functions (Proposition 11.43), the exponential remainder bound Equation (11.31), and the telescoping product estimate Equation (11.36). The fourth is Lévy's continuity theorem, proved in Lévy's Continuity Theorem. Nothing else is assumed, and in particular the Lindeberg condition is shown here to imply the uniform asymptotic negligibility that the expansion needs — that implication is not an extra hypothesis.
Statement and notation
For each \(n\) let \(X_{n,1},\dots,X_{n,k_{n}}\) be independent random variables with \(\avg{X_{n,k}}=0\) and finite variances \(\sigma_{n,k}^{2}=\operatorname{var}X_{n,k}\), and put \(s_{n}^{2}=\sum_{k=1}^{k_{n}}\sigma_{n,k}^{2}\), assumed strictly positive. If for every \(\varepsilon>0\)
then \(s_{n}^{-1}\sum_{k=1}^{k_{n}}X_{n,k}\rightarrow\mathcal{N}(0,1)\) in distribution. Rests on Definition 11.42, Proposition 11.43 and Definition 11.28.
It is convenient to normalise once and for all. Put
and write \(\varphi_{n,k}\) for the characteristic function of \(Y_{n,k}\). The event \(\set{\abs{X_{n,k}}>\varepsilon s_{n}}\) is the event \(\set{\abs{Y_{n,k}}>\varepsilon}\), so Equation (A.367) reads
Every quantity in Equation (A.368) is a pure number: the \(X_{n,k}\) carry whatever SI unit the measurement error carries, and \(s_{n}\) carries the same one, so the ratio is dimensionless and so is the argument \(t\) of the characteristic functions below.
Lindeberg's condition implies uniform asymptotic negligibility
Under Equation (A.369),
Rests on Equation (A.367).
Derives Lemma A.184. Fix \(\varepsilon>0\) and \(k\). Splitting the expectation at \(\abs{Y_{n,k}}=\varepsilon\) and bounding \(Y_{n,k}^{2}\) by \(\varepsilon^{2}\) on the inner piece,
The last term is one summand of the non-negative sum \(L_{n}(\varepsilon)\), hence at most \(L_{n}(\varepsilon)\), and the bound is the same for every \(k\):
Letting \(n\rightarrow\infty\) gives \(\limsup_{n}\max_{k}\tau_{n,k}^{2}\le\varepsilon^{2}\), and \(\varepsilon>0\) was arbitrary.
∎This is Equation (11.39) of Probability and Statistics, proved. It is what licenses treating each factor of the product below as close to \(1\), and it is used twice: to know \(1-\tfrac12 t^{2}\tau_{n,k}^{2}\) lies in \([0,1]\), and to control the sum of squares in Lemma A.187.
The uniform second-order expansion
Lemma 11.44 expands one characteristic function to second order, with a remainder \(o(t^{2})\) that depends on the distribution. Here there is a different distribution for every \((n,k)\) and the remainders must be summed, so what is needed is a bound on the total remainder, uniform in the array. The Lindeberg condition is exactly what supplies it.
For every \(t\in\R\),
Derives Lemma A.185. Apply Equation (11.31) with \(n=2\) and \(x=tY_{n,k}\):
Take expectations and use \(\abs{\avg{Z}}\le\avg{\abs{Z}}\); since \(\avg{Y_{n,k}}=0\) and \(\avg{Y_{n,k}^{2}}=\tau_{n,k}^{2}\) the left side becomes the \(k\)th term of \(\Delta_{n}(t)\), so
Fix \(\varepsilon>0\) and split the expectation at \(\abs{Y_{n,k}}=\varepsilon\), using a different branch of the minimum on each piece. Where \(\abs{Y_{n,k}}\le\varepsilon\) the first branch gives \(\tfrac16\abs{t}^{3}\abs{Y_{n,k}}^{3} \le\tfrac16\abs{t}^{3}\varepsilon\,Y_{n,k}^{2}\); where \(\abs{Y_{n,k}}>\varepsilon\) the second gives \(t^{2}Y_{n,k}^{2}\). Hence
From Equation (A.373) (split the expectation at \(\abs{Y_{n,k}}=\varepsilon\) and use one branch of the minimum on each piece). Summing over \(k\) and using \(\sum_{k}\tau_{n,k}^{2}=1\) from Equation (A.368) and the definition Equation (A.369),
For fixed \(t\) and \(\varepsilon\) the second term tends to \(0\), so \(\limsup_{n}\Delta_{n}(t)\le\varepsilon\abs{t}^{3}/6\); and \(\varepsilon>0\) was arbitrary.
∎The contrast with Lemma 11.44 is worth stating plainly. There the remainder was \(o(t^{2})\) as \(t\rightarrow0\) for a single distribution, and the identically distributed case could rescale one such estimate \(n\) times. Here \(t\) is fixed and \(n\) grows, and the estimate Equation (A.375) is a statement about the array as a whole: the cubic branch of the minimum is used on the bulk of each summand and the quadratic branch on its tail, and Equation (A.367) is precisely the hypothesis that the tails contribute nothing in total.
Two product estimates
For every \(t\in\R\) there is \(N(t)\) such that for \(n\ge N(t)\)
Rests on Equation (11.36), Lemma A.185 and Lemma A.184.
Derives Lemma A.186. Set \(a_{k}=\varphi_{n,k}(t)\) and \(b_{k}=1-\tfrac12 t^{2}\tau_{n,k}^{2}\). By Proposition 11.43(i), \(\abs{a_{k}}\le1\). For the \(b_{k}\): by Lemma A.184 there is \(N(t)\) with \(\tfrac12 t^{2}\max_{k}\tau_{n,k}^{2}\le1\) for \(n\ge N(t)\), and then every \(b_{k}\) lies in \([0,1]\). The telescoping inequality Equation (11.36) applies to such factors and gives \(\abs{\prod a_{k}-\prod b_{k}}\le\sum_{k}\abs{a_{k}-b_{k}}\), which is \(\Delta_{n}(t)\) by Equation (A.372).
∎For every \(t\in\R\),
Rests on Equation (11.36), Lemma A.184 and Theorem 7.38.
Derives Lemma A.187. Write \(z_{k}=\tfrac12 t^{2}\tau_{n,k}^{2}\ge0\), so that \(\sum_{k}z_{k}=\tfrac12t^{2}\) by Equation (A.368) and \(\max_{k}z_{k}\rightarrow0\) by Lemma A.184. Take \(n\) large enough that \(\max_{k}z_{k}\le1\); then both \(1-z_{k}\) and \(\ee^{-z_{k}}\) lie in \([0,1]\) and Equation (11.36) applies:
Taylor's theorem with Lagrange remainder (Theorem 7.38) applied to \(z\mapsto\ee^{-z}\) gives \(\ee^{-z}=1-z+\tfrac12z^{2}\ee^{-\xi}\) with \(\xi\) between \(0\) and \(z\), so for \(z\ge0\) one has \(\abs{\ee^{-z}-1+z}\le\tfrac12z^{2}\). Hence
the middle step bounding each term by \(z_{k}^{2}/2\), factoring out the largest \(z_{k}\) and using \(\sum_{k}z_{k}=t^{2}/2\); while \(\prod_{k}\ee^{-z_{k}}=\ee^{-\sum_{k}z_{k}}=\ee^{-t^{2}/2}\) exactly.
∎Proof of the theorem
Proof of Theorem A.183. Derives Theorem A.183. Write \(T_{n}=s_{n}^{-1}\sum_{k}X_{n,k}=\sum_{k}Y_{n,k}\). The \(Y_{n,k}\), \(k=1,\dots,k_{n}\), are independent, so Proposition 11.43(iii), extended from two factors to \(k_{n}\) by induction, gives
Fix \(t\in\R\). By the triangle inequality,
whose first term tends to \(0\) by Lemma A.186 and whose second tends to \(0\) by Lemma A.187. Hence
From Equations (A.376), (A.377) and (A.379) (replace each factor by its quadratic approximation, then the product of those by the exponential of the sum). The limit is continuous at \(t=0\), and it is the characteristic function of the standard Gaussian: as shown in the derivation of Theorem 11.46, the function \(g(t)=\int\ee^{\ii tu}\ee^{-u^{2}/2}\dd u/\sqrt{2\pi}\) satisfies \(g'(t)=-t\,g(t)\) with \(g(0)=1\), whence \(g(t)=\ee^{-t^{2}/2}\). Theorem A.168 therefore gives \(T_{n}\rightarrow\mathcal{N}(0,1)\) in distribution.
∎Theorem 11.46 follows from Theorem A.183. Rests on Theorem A.183.
Derives Corollary A.188. Let \(X_{1},X_{2},\dots\) be independent and identically distributed with mean \(\mu\) and variance \(\sigma^{2}\in(0,\infty)\), and set \(k_{n}=n\), \(X_{n,k}=X_{k}-\mu\). Then \(\sigma_{n,k}^{2}=\sigma^{2}\) and \(s_{n}^{2}=n\sigma^{2}\), so \(Z_{n}=s_{n}^{-1}\sum_{k}X_{n,k}\) is the statistic of Equation (11.33). The Lindeberg sum is
which tends to \(0\) as \(n\rightarrow\infty\) because \(\avg{(X-\mu)^{2}}=\sigma^{2}\) is finite, a finite integral being one whose tail contribution vanishes — the same elementary fact used at Equation (11.32). The threshold \(\varepsilon\sigma\sqrt{n}\) grows without bound, so the truncation level recedes and the tail expectation goes to zero for each fixed \(\varepsilon\).
∎If there is \(\delta>0\) with
then Equation (A.367) holds, and hence so does the conclusion of Theorem A.183. Rests on Theorem A.183 and Equation (A.367).
Derives Corollary A.189. On the event \(\abs{Y_{n,k}}>\varepsilon\) one has \(1<\abs{Y_{n,k}}^{\delta}/\varepsilon^{\delta}\), so \(Y_{n,k}^{2}\indic{\abs{Y_{n,k}}>\varepsilon} \le\varepsilon^{-\delta}\abs{Y_{n,k}}^{2+\delta}\) pointwise. Taking expectations and summing over \(k\),
which tends to \(0\) for each fixed \(\varepsilon>0\) by Equation (A.381).
∎The condition is sufficient and not necessary, and Probability and Statistics exhibits both failures of the converse: summands equal to \(\pm\sqrt{n}\) with probability \(1/(2n)\) have vanishing variance shares yet Lindeberg sum \(1\) for every \(n\), and adjoining one \(\mathcal{N}(0,n)\) summand to \(n\) standard Gaussians breaks Equation (A.367) while leaving the normalised sum exactly \(\mathcal{N}(0,1)\). What Lemma A.184 establishes is the one implication the error theory uses — no single disturbance may carry a fixed fraction of the total variance — and it is an implication, not a converse: a dominant source violates Equation (A.367), but the absence of a dominant source does not by itself deliver a Gaussian. Feller's companion theorem, which closes the loop inside the negligible class, is not proved here [Feller:1971].
The Lindeberg–Feller Central Limit Theorem discharges the proof obligation of Theorem 11.48, stated in Section 11.3.2 of Probability and Statistics and used there to justify treating the aggregate of many small, non-identically distributed measurement disturbances as Gaussian — the licence on which the standard uncertainty of Measurement, SI Units, and the Theory of Errors rests. Two by-products are worth noting: Equation (11.39) of the chapter is Lemma A.184 here, proved rather than asserted, and Lyapunov's condition, which the chapter quotes, is Corollary A.189. The only input this appendix does not derive is Lévy's continuity theorem, which is Lévy's Continuity Theorem.
The Berry–Esseen Inequality
This appendix proves the uniform rate Equation (11.40) quoted in Remark 11.49 of Probability and Statistics: the Kolmogorov distance between the distribution of a standardised sum of \(n\) independent identically distributed variables and the standard Gaussian is at most a constant times \(\rho/(\sigma^{3}\sqrt{n})\), where \(\rho\) is the third absolute central moment. The central limit theorem Theorem 11.46 is a limit statement and an uncertainty budget needs a rate; this is the rate, and it is what converts “the errors are asymptotically Gaussian” into a number a reader can check.
The proof has three parts, in the order Esseen gave them. First a smoothing inequality, which bounds the largest vertical gap between a distribution function and a smooth comparison function by an integral of the difference of their Fourier–Stieltjes transforms over a finite interval, plus a penalty for truncating the interval (Esseen's smoothing inequality). Then an estimate of that difference for the case at hand, valid on \(\abs{t}\lesssim\sqrt{n}\,\sigma^{3}/\rho\) and carrying a Gaussian decay factor (The characteristic function of the standardised sum). Then the integration, which balances the two terms and fixes an admissible constant (Integration, and the constant).
Notation.
Throughout this appendix \(\varphi(x)=\ee^{-x^{2}/2}/\sqrt{2\pi}\) is the standard Gaussian density and \(\Phi\) its distribution function, as in Definition 11.28; characteristic functions are written \(\psi\), to keep the two apart. For a distribution function \(F\) we write \(\psi_{F}(t)=\int\ee^{\ii tx}\,\dd F(x)\), which is \(\avg{\ee^{\ii tX}}\) for a variable \(X\) with that distribution function (Definition 11.42). The one analytic input is the inversion formula Theorem A.177 of Lévy's Continuity Theorem, and through it the two Lebesgue-theory statements named in Remark A.171; everything else below is derived.
Statement
Let \(X_{1},X_{2},\dots\) be independent and identically distributed with mean \(\mu\), variance \(\sigma^{2}\in(0,\infty)\) and finite third absolute central moment \(\rho=\avg{\abs{X-\mu}^{3}}\). Put \(Z_{n}=\left(\sigma\sqrt{n}\right)^{-1}\sum_{i=1}^{n}(X_{i}-\mu)\) and let \(F_{n}\) be its distribution function. Then for every \(n\ge1\)
Rests on Definition 11.42, Definition 11.28 and Equation (11.31).
The constant produced here is honest but crude; Remark A.201 says what is known about the smallest admissible one and why this proof does not reach it. The quantity \(\rho/\sigma^{3}\) is a pure number — both moments carry the cube of the SI unit of \(X\) — as it must be, since the left-hand side is a difference of probabilities.
The Fejér kernel
For every \(a\in\R\),
Rests on Lemma A.172 and Equation (7.28).
Derives Lemma A.193. For \(a=0\) both sides vanish. For \(a>0\) the substitution \(u=ax\) (Equation (7.27)) turns the left side into \(a\int_{-\infty}^{\infty}\left(1-\cos u\right)u^{-2}\dd u\), and the integrand is even, so it suffices to show \(\int_{0}^{\infty}\left(1-\cos u\right)u^{-2}\dd u=\pi/2\). Integrate by parts (Equation (7.28)) with \(u^{-2}\dd u=-\dd\left(u^{-1}\right)\):
by Lemma A.172. Both boundary terms vanish: at infinity because the numerator is bounded by \(2\), and at zero because \(1-\cos u\le u^{2}/2\). Oddness of \(\cos(ax)\) in \(a\) — it is even — gives the case \(a<0\) from the case \(-a>0\).
∎For \(T>0\) define
extended by \(K_{T}(0)=T/(2\pi)\). Then \(K_{T}\ge0\), \(K_{T}\) is even, \(\int_{\R}K_{T}=1\), and
Moreover \(\int_{\abs{x}>\eta}K_{T}(x)\,\dd x\le4/(\pi T\eta)\) for every \(\eta>0\). Rests on Lemma A.193.
Derives Lemma A.194. Non-negativity and evenness are immediate, and Equation (A.383) with \(a=T\) gives \(\int K_{T}=\pi T/(\pi T)=1\). For the tail bound use \(1-\cos\le2\), so that \(K_{T}(x)\le2/(\pi Tx^{2})\) and
For Equation (A.385): the sine part of \(\ee^{\ii tx}\) pairs with an even function and integrates to zero, so the transform is \(\int\cos(tx)K_{T}(x)\dd x\), which is real. Write the numerator of the integrand, \(\left(1-\cos(Tx)\right)\cos(tx)\), using the product formula \(\cos A\cos B=\tfrac12\left[\cos(A-B)+\cos(A+B)\right]\):
the constants cancelling because \(-1+\tfrac12+\tfrac12=0\). Dividing by \(\pi Tx^{2}\), integrating, and applying Equation (A.383) three times,
From Equations (A.383) and (A.386) (divide by \(\pi Tx^{2}\) and integrate each of the three pieces). For \(\abs{t}\le T\) the two moduli on the right are \(T+t\) and \(T-t\), so the bracket is \(\pi(T-\abs{t})\) and the value is \(1-\abs{t}/T\). For \(t>T\) they are \(T+t\) and \(t-T\), whose half-sum is \(\pi t=\pi\abs{t}\), and the bracket vanishes; the case \(t<-T\) is the mirror image.
∎Esseen's smoothing inequality
The inversion formula of Lévy's Continuity Theorem recovers a distribution function from its transform. What is wanted here is the difference of two of them, at a single point, and in a form that needs no limit.
Let \(h:\R\rightarrow\C\) be continuous and vanish outside a bounded interval. Then \(\int_{\R}h(t)\ee^{-\ii ta}\,\dd t\rightarrow0\) as \(\abs{a}\rightarrow\infty\). Rests on Theorem 7.25.
Derives Lemma A.195. Let \(h\) vanish outside \([-T,T]\) and put \(A(a)=\int_{\R}h(t)\ee^{-\ii ta}\dd t\). Since \(\ee^{-\ii\left(t+\pi/a\right)a}=-\ee^{-\ii ta}\), the substitution \(t\mapsto t+\pi/a\) (Equation (7.27)) gives \(A(a)=-\int_{\R}h(t-\pi/a)\ee^{-\ii ta}\dd t\), so
The function \(h\) is uniformly continuous on \(\R\) (Theorem 7.25 on a compact interval containing the support with room to spare, and \(h\equiv0\) outside), so the integrand tends to \(0\) uniformly and the interval length stays bounded.
∎Let \(F_{1},F_{2}\) be distribution functions with transforms \(\psi_{1},\psi_{2}\), and suppose \(h(t)=\left(\psi_{1}(t)-\psi_{2}(t)\right)/(\ii t)\), extended by its limit at \(t=0\), is continuous on \(\R\) and vanishes outside \([-T,T]\). Then at every point \(x\) at which both \(F_{1}\) and \(F_{2}\) are continuous,
Rests on Theorem A.177 and Lemma A.195.
Derives Lemma A.196. Let \(a<x\) be a common continuity point. Applying Equation (A.360) to \(F_{1}\) and to \(F_{2}\) and subtracting — both limits exist, so their difference is the limit of the difference — gives, for every \(T'\ge T\),
and since \(h\) vanishes outside \([-T,T]\) the integral does not depend on \(T'\) and no limit is needed. Split it into its two terms. By Lemma A.195 the term carrying \(\ee^{-\ii ta}\) tends to \(0\) as \(a\rightarrow-\infty\) through continuity points, while \(F_{1}(a)\rightarrow0\) and \(F_{2}(a)\rightarrow0\). What survives is Equation (A.388).
∎If \(X\) has \(\avg{\abs{X}}<\infty\) and characteristic function \(\psi_{X}\), then \(\psi_{X}(t)=1+\ii t\avg{X}+o(\abs{t})\) as \(t\rightarrow0\). Rests on Equation (11.31) and Definition 11.42.
Derives Lemma A.197. Equation (11.31) at order one gives \(\abs{\ee^{\ii x}-1-\ii x}\le\min\left(x^{2}/2,\,2\abs{x}\right)\). Put \(x=tX\), take expectations, use \(\abs{\avg{Z}}\le\avg{\abs{Z}}\) and divide by \(\abs{t}\):
the last step splitting the expectation at \(\abs{X}=K\) and using one branch of the minimum on each piece, exactly as at Equation (11.32). Given \(\varepsilon>0\) choose \(K\) with the second term below \(\varepsilon/2\) — possible because \(\avg{\abs{X}}\) is finite — and then \(\abs{t}<\varepsilon/K^{2}\).
∎Let \(F\) be a distribution function and \(G\) a non-decreasing differentiable function with \(G(-\infty)=0\), \(G(+\infty)=1\) and \(\sup_{x}G'(x)\le m<\infty\); suppose both have a finite first absolute moment. Write \(\psi_{F},\psi_{G}\) for their transforms and \(\Delta=\sup_{x}\abs{F(x)-G(x)}\). Then for every \(T>0\)
Derives Theorem A.198. Write \(H=F-G\), so \(\abs{H}\le1\) and \(\Delta\le1\) is finite; if \(\Delta=0\) there is nothing to prove, and if the integral on the right is infinite the inequality is trivial, so assume both finite and \(\Delta>0\). Let \(W\) have density \(K_{T}\), independently of everything else, and let \(F_{T},G_{T}\) be the distribution functions of \(U+W\) and \(V+W\), where \(U\) has distribution function \(F\) and \(V\) has \(G\). Both are continuous, being convolutions with a density, and by Proposition 11.43(iii) and Equation (A.385) their transforms are \(\psi_{F}\,\hat{K}_{T}\) and \(\psi_{G}\,\hat{K}_{T}\) with \(\hat{K}_{T}(t)=\left(1-\abs{t}/T\right)\) on \([-T,T]\) and \(0\) outside. Since \(K_{T}\) is even,
The transform side. The function \(\left(\psi_{F}-\psi_{G}\right)\hat{K}_{T}/(\ii t)\) is continuous — at \(t=0\) because Lemma A.197, applicable since both first absolute moments are finite, gives \(\psi_{F}(t)-\psi_{G}(t)=\ii t\left(\avg{U}-\avg{V}\right) +o(\abs{t})\), and at \(t=\pm T\) because \(\hat{K}_{T}\) vanishes there — and is supported in \([-T,T]\). So Lemma A.196 applies, and since \(\abs{\hat{K}_{T}}\le1\),
The kernel side. Put \(\eta=\Delta/(2m)\) and \(\kappa=\int_{\abs{s}>\eta}K_{T}(s)\,\dd s\), so that \(\kappa\le4/(\pi T\eta)=8m/(\pi T\Delta)\) by Lemma A.194. Let \(\varepsilon\in(0,\Delta/2)\). Two cases, according to which sign of \(H\) realises the supremum.
Suppose first \(\sup_{x}H(x)=\Delta\) and pick \(b\) with \(H(b)>\Delta-\varepsilon\). For \(y\ge b\) the mean value theorem (Theorem 7.35) gives \(G(y)\le G(b)+m(y-b)\), while \(F(y)\ge F(b)\); hence \(H(y)\ge H(b)-m(y-b)\). Put \(x_{\ast}=b+\eta\). For \(\abs{s}\le\eta\) the point \(x_{\ast}+s\) lies at distance \(\eta+s\in[0,2\eta]\) to the right of \(b\), so
Insert this in Equation (A.390). The term \(-ms\) integrates to zero against the even kernel over the symmetric set \(\abs{s}\le\eta\); outside it, \(H\ge-\Delta\). Therefore
From Equations (A.390) and (A.392) (split the convolution at \(\abs{s}=\eta\), use the wedge bound inside and \(H\ge-\Delta\) outside, then insert \(\kappa\le8m/(\pi T\Delta)\)).
If instead \(\sup_{x}\left(-H(x)\right)=\Delta\), pick \(b\) with \(H(b)<-\Delta+\varepsilon\) and put \(x_{\ast}=b-\eta\). For \(y\le b\) the same mean value theorem gives \(G(y)\ge G(b)-m(b-y)\) and \(F(y)\le F(b)\), so \(H(y)\le H(b)+m(b-y)\); for \(\abs{s}\le\eta\) the point \(x_{\ast}+s\) lies at distance \(\eta-s\in[0,2\eta]\) to the left of \(b\) and \(H(x_{\ast}+s)\le-\Delta/2+\varepsilon-ms\). The same two steps give \(H_{T}(x_{\ast})\le-\Delta/2+\varepsilon+12m/(\pi T)\).
In both cases \(\abs{H_{T}(x_{\ast})}\ge\Delta/2-\varepsilon-12m/(\pi T)\). Combining with Equation (A.391), letting \(\varepsilon\downarrow0\) and multiplying by \(2\) gives Equation (A.389).
∎The characteristic function of the standardised sum
Let \(\avg{Y}=0\), \(\avg{Y^{2}}=1\) and \(\beta=\avg{\abs{Y}^{3}}<\infty\). Then \(\beta\ge1\). Rests on Lemma 11.63.
Derives Lemma A.199. The discriminant argument of Lemma 11.63, applied to the nowhere-negative quadratic \(s\mapsto\avg{(sV+W)^{2}}\) without centring — which is all that proof uses, as Proposition 11.24 already records — gives \(\avg{VW}^{2}\le\avg{V^{2}}\avg{W^{2}}\) for any \(V,W\) of finite second moment. Take \(V=\abs{Y}^{1/2}\) and \(W=\abs{Y}^{3/2}\): \(1=\avg{Y^{2}}^{2}\le\avg{\abs{Y}}\,\avg{\abs{Y}^{3}} =\avg{\abs{Y}}\,\beta\). Take \(V=\abs{Y}\), \(W=1\): \(\avg{\abs{Y}}^{2}\le\avg{Y^{2}}=1\). Combining, \(1\le\beta\).
∎Let \(Y,Y_{1},Y_{2},\dots\) be independent and identically distributed with \(\avg{Y}=0\), \(\avg{Y^{2}}=1\) and \(\beta=\avg{\abs{Y}^{3}}<\infty\), and let \(Z_{n}=n^{-1/2}\sum_{i=1}^{n}Y_{i}\). Then
Rests on Equation (11.31), Lemma A.199 and Proposition 11.43.
Derives Lemma A.200. Write \(\psi\) for the common characteristic function of the \(Y_{i}\) and \(s=t/\sqrt{n}\), so that \(\psi_{Z_{n}}(t)=\psi(s)^{n}\) by Proposition 11.43(ii) and (iii). Fix \(t\) with \(\abs{t}\le T_{n}\); then \(\beta\abs{s}\le\tfrac12\), and since \(\beta\ge1\) by Lemma A.199 also \(\abs{s}\le\tfrac12\).
Step 1: the quadratic expansion. Equation (11.31) at order two, with \(x=sY\), expectations taken and \(\avg{Y}=0\), \(\avg{Y^{2}}=1\) used, gives
Since \(\beta\abs{s}\le\tfrac12\) one has \(\beta\abs{s}^{3}/6\le s^{2}/12\), so
In particular \(\psi(s)\neq0\), and \(\psi(s)^{n}=\ee^{n\ln\psi(s)}\) with the principal logarithm — for an integer power the branch is immaterial.
Step 2: the logarithm. For \(\abs{w}\le\tfrac12\) the principal logarithm satisfies
Now \(n\omega+\tfrac12t^{2}=nr\), because \(ns^{2}=t^{2}\), and \(\abs{nr}\le\beta\abs{t}^{3}/(6\sqrt{n})\) by Equation (A.395). Also, using Equation (A.396) and then \(\abs{t}/n\le1/(2\beta\sqrt{n})\le\beta/(2\sqrt{n})\) — the last step being \(1/\beta\le\beta\), valid since \(\beta\ge1\) —
Adding, and using \(\tfrac16+\tfrac{49}{288}=\tfrac{97}{288}<\tfrac12\),
From Equations (A.395), (A.396) and (A.397) (add the cubic remainder of the expansion to the quadratic error of the logarithm, using \(\beta\abs{t}\le\sqrt{n}/2\)).
Step 3: exponentiate. On \(\abs{t}\le T_{n}\),
so by Equation (A.398) the real part of \(n\ln\psi(s)\) is at most \(-t^{2}/2+t^{2}/4=-t^{2}/4\), and the same bound holds trivially for \(\Re\left(-t^{2}/2\right)\). For complex \(a,b\),
since \(\Re\left(b+u(a-b)\right)=(1-u)\Re b+u\Re a\) is a convex combination. Taking \(a=n\ln\psi(s)\) and \(b=-t^{2}/2\) and inserting Equation (A.398) gives Equation (A.394).
∎Integration, and the constant
Proof of Theorem A.192. Derives Theorem A.192. Put \(Y_{i}=(X_{i}-\mu)/\sigma\), so the \(Y_{i}\) are independent and identically distributed with mean \(0\), variance \(1\) and \(\beta=\avg{\abs{Y}^{3}}=\rho/\sigma^{3}\), and \(Z_{n}=n^{-1/2}\sum_{i}Y_{i}\) is the statistic of the theorem. Apply Theorem A.198 with \(F=F_{n}\), \(G=\Phi\) and \(T=T_{n}=\sqrt{n}/(2\beta)\). The comparison function qualifies: \(\Phi\) is non-decreasing and differentiable with \(\Phi'=\varphi\le1/\sqrt{2\pi}\), so \(m=1/\sqrt{2\pi}\); and both distributions have finite first moments, namely zero. Then, by Lemma A.200,
From Equations (A.389) and (A.394) (insert the transform estimate into the smoothing inequality at \(T=T_{n}=\sqrt{n}/(2\beta)\)). Extend the first integral to all of \(\R\), which only increases it, and evaluate it by the substitution \(t=2u\) together with the Gaussian moment \(\int u^{2}\ee^{-u^{2}}\dd u=\sqrt{\pi}/2\) (Proposition 11.29, rescaled):
The first term of Equation (A.400) is therefore at most \(\left(4\sqrt{\pi}/\pi\right)\beta/(2\sqrt{n}) =2\beta/\sqrt{\pi n}\), and
since \(2/\sqrt{\pi}=1.1284\) and \(48/\left(\pi\sqrt{2\pi}\right)=6.0954\). That is Equation (A.382). Note that no step required \(n\) to be large: the inequality holds for every \(n\ge1\), being vacuous — because \(\Delta_{n}\le1\) always — until \(n\) exceeds about \(52\beta^{2}\).
∎The value \(C=7.23\) derived here is admissible and is not sharp, and it is worth saying where the loss is. Almost all of it is in the smoothing inequality: of the two terms in Equation (A.401), the truncation penalty \(24m/(\pi T)\) contributes \(6.10\) and the transform integral only \(1.13\). The factor \(24\) in Equation (A.389) comes from the crude estimate \(1-\cos\le2\) used for the tail of the Fejér kernel and from the wedge construction's factor \(\tfrac32\); sharpening either changes the arithmetic but not the structure. The interval cannot simply be lengthened, because Equation (A.394) is proved only for \(\abs{t}\le\sqrt{n}/(2\beta)\) — beyond that the logarithm of \(\psi(t/\sqrt{n})\) need not even be defined — so \(T=T_{n}\) is the best this argument allows.
Determining the smallest admissible \(C\) is a separate literature and is not attempted here. Berry's paper [Berry:1941] establishes finiteness, which is the qualitative content and all that Remark 11.49 needs; Esseen obtained the same result independently in the same period; and the sharpest published value for identically distributed summands is \(C\le0.4690\) [Shevtsova:2014]. Since \(\rho/\sigma^{3}\ge1\) always (Lemma A.199), no value of \(C\) below that can make the bound say anything at small \(n\), which is the practical point Remark 11.49 draws.
The Berry–Esseen Inequality discharges the proof obligation of Equation (11.40), stated in Remark 11.49 of Probability and Statistics as the rate that turns the central limit theorem into a usable statement about a finite sample. The consequences drawn there stand unchanged: the bound is uniform in \(z\), so its relative accuracy is worst in the tails, exactly where a coverage factor \(k=3\) operates; and the \(n^{-1/2}\) decay is slow enough that several hundred contributions are needed before the inequality itself certifies a few per cent. What this appendix adds is that both statements are now consequences of something proved, with an explicit constant, rather than of something quoted.
Wilks' Theorem in the Multi-Parameter Case
This appendix completes the derivation of Theorem 11.88 of Probability and Statistics in \(k\) parameters with \(r\) constraints. The chapter's argument is correct at the level of the quadratic approximation and two of its steps are asserted rather than proved:
-
[(i)] the remainder in the \(k\)-dimensional Taylor expansion of the log-likelihood must be shown uniform on a neighbourhood of the true parameter that shrinks at rate \(n^{-1/2}\), since the expansion is evaluated not at a fixed point but at two estimators that move with the data;
-
[(ii)] both estimators — the unconstrained one and the one constrained to \(\Theta_{0}\) — must be shown to lie in that neighbourhood with probability tending to one, since that is what licenses replacing the constraint surface \(\Theta_{0}\) by its tangent space at \(\theta_{0}\), the step written “\(\Theta_{0}\simeq\theta_{0}+E\)” at Equation (11.85).
Both are proved below, from exactly the regularity conditions (R1)–(R5) of Definition 11.59 and nothing stronger, and the chi-squared limit with \(r\) degrees of freedom is then assembled from them. The one-parameter case with a simple null needs none of this — there \(\Theta_{0}\) is a point, there is no surface to replace — which is why Probability and Statistics carries that derivation in full and defers only this one.
Setting and notation
Let \(X_{1},\dots,X_{n}\) be independent and identically distributed with density \(f(\,\cdot\,;\theta_{0})\), let \(\Theta\subseteq\R^{k}\) and let (R1)–(R5) hold at \(\theta_{0}\), read in \(k\) parameters: (R4) asks that \(\theta\mapsto f(x;\theta)\) be three times continuously differentiable near \(\theta_{0}\) with the first two derivatives available under the integral, and (R5) that the information matrix \(I=I_{1}(\theta_{0})\) of Definition 11.60 be finite and positive definite and that a single integrable \(M(x)\) dominate every third partial derivative \(\pp^{3}\ln f(x;\theta)/\pp\theta^{a}\pp\theta^{b}\pp\theta^{c}\) throughout a ball \(\bar B(\theta_{0},a_{0})\subseteq\Theta\). Write \(\lambda>0\) for the smallest eigenvalue of \(I\) and \(\norm{A}=\sup_{\norm{\vect v}=1}\norm{A\vect v}\) for the operator norm of a matrix.
Put \(\ell_{n}(\theta)=\sum_{i=1}^{n}\ln f(X_{i};\theta)\) and
All three are averages of independent identically distributed terms: \(\vect{Z}_{n}\) of the scores, whose common mean is zero and whose covariance is \(I\) (Lemma 11.61); \(J_{n}\) of the observation Hessians, whose common mean is \(-I\) by Equation (11.48); and \(\bar{M}_{n}\) of \(M\), whose common mean \(\bar{M}=\avg{M}\) is finite by (R5). The statistic is a pure number, the SI dimensions of \(\theta\) cancelling in the likelihood ratio of Equation (11.80), and so is every quantity above: \(\nabla\ell_{n}\) carries the reciprocal of the unit of \(\theta\) and \(I\) its square, so the combinations formed below — \(\vect{u}\transpose I\vect{u}\) with \(\vect{u}=\sqrt{n}(\theta-\theta_{0})\), and \(\vect{Z}_{n}\transpose I^{-1}\vect{Z}_{n}\) — are dimensionless.
The two estimators are the local ones, defined exactly as Probability and Statistics defines its root of the score equation. Fix \(a_{1}\in(0,a_{0}]\), to be chosen in Lemma A.211, and let
Both maxima are attained: \(\ell_{n}\) is continuous and both sets are closed and bounded, hence compact, so the extreme value theorem (Theorem 7.24) applies. Both sets contain \(\theta_{0}\), so
the second because \(\theta_{0}\in\Theta_{0}\) by hypothesis. The statistic under study is \(q=-2\left[\ell_{n}(\hat{\theta}_{0,n}) -\ell_{n}(\hat{\theta}_{n})\right]\ge0\).
The null set is a \(C^{2}\) submanifold near \(\theta_{0}\), which is used in exactly one form and is therefore stated in that form.
\(\Theta_{0}\) is a \(C^{2}\) submanifold of dimension \(k-r\) through \(\theta_{0}\) with tangent space \(E\) when there are a linear subspace \(E\subseteq\R^{k}\) with \(\dim E=k-r\), a neighbourhood \(W\) of \(\theta_{0}\), an open \(V\ni\vect{0}\) in \(E\), and a twice continuously differentiable map \(h:V\rightarrow E^{\perp}\) with \(h(\vect{0})= \vect{0}\) and \(Dh(\vect{0})=0\), such that
Rests on Definition 11.87.
Theorem 11.88 states the hypothesis as “\(H_{0}\) imposes \(r\) smooth, functionally independent constraints”, i.e. \(\Theta_{0}\) is the level set \(\set{g(\theta)=\vect{0}}\) of a \(C^{2}\) map \(g:\Theta\rightarrow\R^{r}\) with \(Dg(\theta_{0})\) of rank \(r\). The equivalence of that description with Equation (A.405) is the implicit function theorem in several variables, which Differentiable Manifolds, Tensors, and Curvature records as owed rather than proved. This appendix therefore takes Definition A.203 as the hypothesis: it is the form the proof uses, it is one of the standard equivalent definitions of a submanifold, and adopting it keeps the argument free of a theorem the treatise has not yet built. A reader who has the implicit function theorem may read the two hypotheses as the same one.
What is quoted
Let \(W_{1},W_{2},\dots\) be independent and identically distributed with \(\avg{\abs{W}}<\infty\). Then \(n^{-1}\sum_{i=1}^{n}W_{i}\rightarrow\avg{W}\) in probability [Billingsley:1995]. Rests on Corollary 11.51.
Let \(\vect{W}_{n},\vect{W}\) be random vectors in \(\R^{k}\). Then \(\vect{W}_{n}\rightarrow\vect{W}\) in distribution if and only if \(\vect{a}\transpose\vect{W}_{n}\rightarrow \vect{a}\transpose\vect{W}\) in distribution for every \(\vect{a}\in\R^{k}\). Moreover, if \(\vect{W}_{n}\rightarrow\vect{W}\) in distribution and \(g:\R^{k}\rightarrow\R\) is continuous, then \(g(\vect{W}_{n})\rightarrow g(\vect{W})\) in distribution [Billingsley:1995]. Rests on Definition 11.21 and Lemma 11.52.
Three imports, all of them already relied on by Probability and Statistics itself, and no more. Theorem A.205 is the weak law in the form that assumes only integrability; Corollary 11.51 is the Chebyshev form and needs a finite variance, which (R5) does not supply for the Hessian summands or for \(M\) — the chapter names this gap explicitly in the derivation of Theorem 11.67 and quotes the same source. Theorem A.206 is what promotes the scalar central limit theorem Theorem 11.46 to the vector statement \(\vect{Z}_{n}\rightarrow\mathcal{N}(0,I)\) and lets a continuous function be applied under a limit in distribution; the chapter uses both silently in the \(k\)-parameter clause of Theorem 11.67 and proves the one-dimensional special case of the second at Equation (11.44). The third import is the implicit function theorem, and it is avoided rather than used, by taking Definition A.203 as the hypothesis (Remark A.204). Everything else below — the uniform remainder, the localisation of both estimators, the tangent-space replacement and the degrees of freedom — is derived here.
\(\norm{J_{n}+I}\rightarrow0\) and \(\bar{M}_{n}\rightarrow\bar{M}\) in probability, and \(\vect{Z}_{n}\rightarrow\mathcal{N}(0,I)\) in distribution. In particular each component of \(\vect{Z}_{n}\), and hence \(\norm{\vect{Z}_{n}}\), is bounded in probability. Rests on Lemma 11.61, Theorem A.205 and Theorem A.206.
Derives Lemma A.208. Each entry of \(J_{n}\) is the average of \(n\) independent copies of \(\pp^{2}\ln f(X;\theta_{0})/\pp\theta^{a}\pp\theta^{b}\), integrable by (R4) with mean \(-I_{ab}\) by Equation (11.48), so Theorem A.205 gives \(J_{n}\rightarrow-I\) entrywise in probability; a matrix norm is a continuous function of finitely many entries, so \(\norm{J_{n}+I}\rightarrow0\) in probability. The same theorem applied to \(M\), integrable by (R5), gives the second claim.
For the third, fix \(\vect{a}\in\R^{k}\), \(\vect{a}\neq\vect{0}\). Then \(\vect{a}\transpose\vect{Z}_{n} =n^{-1/2}\sum_{i}\vect{a}\transpose\vect{U}(\theta_{0};X_{i})\) with \(\vect{U}=\nabla\ln f\) the score, a normalised sum of independent identically distributed variables with mean zero (Lemma 11.61) and variance \(\vect{a}\transpose I\vect{a}\), finite and strictly positive by (R5). Theorem 11.46 gives \(\vect{a}\transpose\vect{Z}_{n}\rightarrow \mathcal{N}(0,\vect{a}\transpose I\vect{a})\), and Theorem A.206 converts this into \(\vect{Z}_{n}\rightarrow\mathcal{N}(0,I)\). Boundedness in probability of each component is Equation (11.43), and \(\norm{\vect{Z}_{n}}\le\sum_{a}\abs{Z_{n}^{a}}\) is a finite sum of variables each bounded in probability.
∎The uniform quadratic expansion
This is item (i). The content is not that a remainder exists — Taylor gives that at every fixed \(\theta\) — but that a single random variable, not depending on \(\theta\), bounds it throughout the ball.
For every \(\theta\in\bar B(\theta_{0},a_{0})\), writing \(\delta=\theta-\theta_{0}\),
The bound is uniform in \(\theta\): apart from the explicit factor \(\norm{\delta}^{3}\) it involves only \(\bar{M}_{n}\), which does not depend on \(\theta\). Rests on Definition 11.59 and Theorem 7.38.
Derives Lemma A.209. Fix \(\theta\) and set \(g(u)=\ell_{n}(\theta_{0}+u\delta)\) for \(u\in[0,1]\); the segment lies in \(\bar B(\theta_{0},a_{0})\) by convexity of the ball, so \(g\) is three times continuously differentiable by (R4). Taylor's theorem with Lagrange remainder (Theorem 7.38) gives \(g(1)=g(0)+g'(0)+\tfrac12g''(0)+\tfrac16g'''(\xi)\) for some \(\xi\in(0,1)\). By the chain rule (Proposition 7.72),
which are the first two displayed terms, and
Each third partial derivative of \(\ell_{n}\) is a sum of \(n\) terms, one per observation, and (R5) bounds each of them by \(M(X_{i})\) throughout \(\bar B(\theta_{0},a_{0})\) — in particular at the unknown intermediate point \(\theta_{0}+\xi\delta\), which is what makes the estimate independent of \(\xi\) and hence of \(\theta\). Since \(\abs{\delta^{a}}\le\norm{\delta}\) for each component and the sum Equation (A.407) has \(k^{3}\) terms,
and \(R_{n}(\theta)=\tfrac16g'''(\xi)\).
∎Define, for \(\vect{u}\in\R^{k}\),
Then for every \(K>0\) and every \(n\) with \(K/\sqrt{n}\le a_{0}\),
where \(\epsilon_{n}(K)=K^{2}\norm{J_{n}+I} +\tfrac13k^{3}K^{3}\bar{M}_{n}/\sqrt{n}\) tends to \(0\) in probability. Rests on Lemmas A.208 and A.209.
Derives Corollary A.210. Put \(\delta=\vect{u}/\sqrt{n}\) in Equation (A.406) and multiply by \(-2\):
because \(n\delta\transpose J_{n}\delta =\vect{u}\transpose J_{n}\vect{u}\) and \(\sqrt{n}\vect{Z}_{n}\transpose\delta =\vect{Z}_{n}\transpose\vect{u}\). Subtracting Equation (A.408) leaves \(-\vect{u}\transpose\left(J_{n}+I\right)\vect{u}-2R_{n}\), whose first part is at most \(\norm{J_{n}+I}\norm{\vect{u}}^{2}\) in modulus and whose second is at most \(\tfrac13k^{3}\bar{M}_{n}\norm{\vect{u}}^{3}/\sqrt{n}\) by Equation (A.406) with \(\norm{\delta}=\norm{\vect{u}}/ \sqrt{n}\). Both are increasing in \(\norm{\vect{u}}\), so the supremum over the ball is attained at the bound. That \(\epsilon_{n}(K)\rightarrow0\) in probability is Lemma A.208: the first term is \(K^{2}\) times something tending to \(0\), and the second is \(\bar{M}_{n}\), bounded in probability, times \(\tfrac13k^{3}K^{3}n^{-1/2}\rightarrow0\).
∎Localisation of both estimators
This is item (ii). Note that it is proved for any closed subset of the ball containing \(\theta_{0}\), so that the unconstrained and the constrained estimators are covered by one argument; nothing about the geometry of \(\Theta_{0}\) is used here.
Choose \(a_{1}\in(0,a_{0}]\) with \(\tfrac16k^{3}\left(\bar{M}+1\right)a_{1}\le\lambda/8\), and let \(A\subseteq\bar B(\theta_{0},a_{1})\) be any closed set with \(\theta_{0}\in A\). Let \(\hat{\theta}\) maximise \(\ell_{n}\) over \(A\). Then
and consequently \(\sqrt{n}\norm{\hat{\theta}-\theta_{0}}\) is bounded in probability and \(\hat{\theta}\rightarrow\theta_{0}\) in probability. In particular both estimators Equation (A.403) satisfy this. Rests on Lemmas A.208 and A.209.
Derives Lemma A.211. Write \(\hat{\delta}=\hat{\theta}-\theta_{0}\) and \(d=\norm{\hat{\delta}}\le a_{1}\). Since \(\theta_{0}\in A\) we have \(\ell_{n}(\hat{\theta})\ge\ell_{n}(\theta_{0})\), so Equation (A.406) gives
Bound the three terms. The first is at most \(\sqrt{n}\norm{\vect{Z}_{n}}d\) by the Cauchy–Schwarz inequality for vectors. For the second, write \(\hat{\delta}\transpose J_{n}\hat{\delta} =-\hat{\delta}\transpose I\hat{\delta} +\hat{\delta}\transpose\left(J_{n}+I\right)\hat{\delta}\); the first piece is at most \(-\lambda d^{2}\), since \(\lambda\) is the smallest eigenvalue of the positive-definite \(I\), and the second is at most \(\norm{J_{n}+I}d^{2}\). The third is at most \(\tfrac16nk^{3}\bar{M}_{n}d^{3}\). Hence
From Equations (A.406) and (A.411) (bound the linear term by Cauchy–Schwarz, split the Hessian term into \(-I\) plus its deviation, and insert the uniform remainder bound).
Let \(E_{n}\) be the event on which both \(\norm{J_{n}+I}\le\lambda/4\) and \(\bar{M}_{n}\le\bar{M}+1\); by Lemma A.208, \(\Pr(E_{n})\rightarrow1\). On \(E_{n}\), and using \(d\le a_{1}\) with the choice of \(a_{1}\),
Substituting in Equation (A.412) and cancelling \(n\lambda d^{2}/4\) from both sides leaves \(\tfrac14n\lambda d^{2}\le\sqrt{n}\norm{\vect{Z}_{n}}d\); dividing by \(\tfrac14n\lambda d\) when \(d>0\) — and noting the conclusion is trivial when \(d=0\) — gives \(\sqrt{n}\,d\le4\norm{\vect{Z}_{n}}/\lambda\), which is Equation (A.410). Boundedness in probability follows because \(\norm{\vect{Z}_{n}}\) is bounded in probability (Lemma A.208), and consistency because \(\sqrt{n}d=O_{P}(1)\) forces \(d\rightarrow0\) in probability.
∎Replacing the surface by its tangent space
Under Definition A.203 there are \(c_{0}<\infty\) and \(a_{2}\in(0,a_{1}]\) such that
-
if \(\theta\in\Theta_{0}\) and \(\norm{\theta-\theta_{0}}\le a_{2}\), then the orthogonal projection \(\vect{w}\) of \(\theta-\theta_{0}\) onto \(E\) satisfies \(\norm{\left(\theta-\theta_{0}\right)-\vect{w}} \le c_{0}\norm{\theta-\theta_{0}}^{2}\);
-
if \(\vect{e}\in E\) and \(\norm{\vect{e}}\le a_{2}\), there is \(\theta\in\Theta_{0}\) with \(\norm{\left(\theta-\theta_{0}\right)-\vect{e}} \le c_{0}\norm{\vect{e}}^{2}\).
Rests on Definition A.203 and Theorem 7.38.
Derives Lemma A.212. Choose \(a_{2}\in(0,a_{1}]\) small enough that the closed ball \(\bar B(\theta_{0},2a_{2})\) lies in \(W\) and that \(\set{\vect{e}\in E\mid\norm{\vect{e}}\le2a_{2}}\subseteq V\); let \(c_{0}=\tfrac12\sup\norm{D^{2}h}\) over that compact set, finite because \(h\) is twice continuously differentiable and the supremum of a continuous function on a compact set is attained (Theorem 7.24). Taylor's theorem applied to \(u\mapsto h(u\vect{e})\) on \([0,1]\), with \(h(\vect{0})=\vect{0}\) and \(Dh(\vect{0})=0\) killing the first two terms, gives
(ii) Take \(\theta=\theta_{0}+\vect{e}+h(\vect{e})\), which lies in \(\Theta_{0}\) by Equation (A.405); then \((\theta-\theta_{0})-\vect{e}=h(\vect{e})\) and Equation (A.413) applies.
(i) Such a \(\theta\) lies in \(W\), so \(\theta-\theta_{0}=\vect{e}+h(\vect{e})\) for some \(\vect{e}\in V\); the two summands are orthogonal, so \(\vect{w}=\vect{e}\) and \(\norm{\vect{e}}\le\norm{\theta-\theta_{0}}\le a_{2}\). Then \((\theta-\theta_{0})-\vect{w}=h(\vect{e})\), and Equation (A.413) together with \(\norm{\vect{e}}\le\norm{\theta-\theta_{0}}\) gives the bound.
∎Write \(\hat{\vect{u}}=\sqrt{n}(\hat{\theta}_{n}-\theta_{0})\) and \(\hat{\vect{u}}_{0}=\sqrt{n}(\hat{\theta}_{0,n}-\theta_{0})\). Then
both in probability. Rests on Corollary A.210, Lemma A.211 and Lemma A.212.
Derives Lemma A.213. Since \(I\) is positive definite, \(\Lambda_{n}\) is a strictly convex quadratic; completing the square with \(\vect{m}=I^{-1}\vect{Z}_{n}\),
so the unconstrained minimum is attained at \(\vect{m}\) and the minimum over the subspace \(E\) at the \(I\)-orthogonal projection \(\vect{p}\) of \(\vect{m}\) onto \(E\). Both \(\norm{\vect{m}}\) and \(\norm{\vect{p}}\) are bounded in probability, being bounded linear images of \(\vect{Z}_{n}\). Also \(\Lambda_{n}\) is Lipschitz on each ball: for \(\norm{\vect{u}},\norm{\vect{u}'}\le K\),
which is bounded in probability for fixed \(K\).
Fix \(\varepsilon>0\). By Lemma A.211 and the boundedness just noted, there is \(K\) such that with probability at least \(1-\varepsilon\) for all large \(n\) the four vectors \(\hat{\vect{u}},\hat{\vect{u}}_{0},\vect{m},\vect{p}\) all lie in \(\bar B(\vect{0},K)\); call that event \(A_{n}\) and intersect it with the event that \(\epsilon_{n}(K+1)\le\varepsilon\), which also has probability approaching \(1\) by Corollary A.210. Work on the intersection, and abbreviate \(D_{n}(\theta)=-2\left[\ell_{n}(\theta)-\ell_{n}(\theta_{0})\right]\), so that Equation (A.409) reads \(\abs{D_{n}(\theta_{0}+\vect{u}/\sqrt{n})-\Lambda_{n}(\vect{u})} \le\varepsilon\) for every \(\norm{\vect{u}}\le K+1\). The extra unit of radius is there because one auxiliary point below, the image in \(\Theta_{0}\) of the tangential minimiser, misses \(\bar B(\vect{0},K)\) by \(O(n^{-1/2})\); all other points used lie in \(\bar B(\vect{0},K)\). Write \(L_{n}=L_{n}(K+1)\) for the Lipschitz constant Equation (A.416) at that radius.
First limit. That \(\Lambda_{n}(\hat{\vect{u}})\ge \min_{\R^{k}}\Lambda_{n}\) is immediate. For the reverse, put \(\theta_{\mathrm{m}}=\theta_{0}+\vect{m}/\sqrt{n}\), which lies in \(\bar B(\theta_{0},a_{1})\) once \(K/\sqrt{n}\le a_{1}\), so \(\ell_{n}(\hat{\theta}_{n})\ge\ell_{n}(\theta_{\mathrm{m}})\) and hence \(D_{n}(\hat{\theta}_{n})\le D_{n}(\theta_{\mathrm{m}})\). Therefore
so the first difference in Equation (A.414) is at most \(2\varepsilon\) on an event of probability approaching \(1-\varepsilon\).
Second limit. For the lower bound let \(\vect{w}\) be the orthogonal projection of \(\hat{\vect{u}}_{0}\) onto \(E\). By Lemma A.212(i) applied to \(\theta=\hat{\theta}_{0,n}\in\Theta_{0}\), whose distance to \(\theta_{0}\) is at most \(K/\sqrt{n}\le a_{2}\) for large \(n\),
and \(\norm{\vect{w}}\le\norm{\hat{\vect{u}}_{0}}\le K\), so Equation (A.416) gives \(\Lambda_{n}(\hat{\vect{u}}_{0})\ge\Lambda_{n}(\vect{w}) -L_{n}c_{0}K^{2}/\sqrt{n}\ge\min_{E}\Lambda_{n} -L_{n}c_{0}K^{2}/\sqrt{n}\).
For the upper bound apply Lemma A.212(ii) to \(\vect{p}/\sqrt{n}\), of norm at most \(K/\sqrt{n}\le a_{2}\): there is \(\theta_{\mathrm{p}}\in\Theta_{0}\) with \(\norm{\left(\theta_{\mathrm{p}}-\theta_{0}\right) -\vect{p}/\sqrt{n}}\le c_{0}K^{2}/n\), hence \(\norm{\sqrt{n}\left(\theta_{\mathrm{p}}-\theta_{0}\right)-\vect{p}} \le c_{0}K^{2}/\sqrt{n}\) and \(\theta_{\mathrm{p}}\in \Theta_{0}\cap\bar B(\theta_{0},a_{1})\) for large \(n\). Since \(\hat{\theta}_{0,n}\) maximises \(\ell_{n}\) over that set, \(D_{n}(\hat{\theta}_{0,n})\le D_{n}(\theta_{\mathrm{p}})\), and running the same three-step chain as before, then moving from \(\sqrt{n}(\theta_{\mathrm{p}}-\theta_{0})\) to \(\vect{p}\) by Equation (A.416),
Both correction terms tend to \(0\) in probability, \(L_{n}\) being bounded in probability, and \(\varepsilon>0\) was arbitrary.
∎The chi-squared limit
Let (R1)–(R5) hold at \(\theta_{0}\) in \(k\) parameters, let the sample be independent and identically distributed of size \(n\), and let \(\Theta_{0}\ni\theta_{0}\) satisfy Definition A.203 with \(\dim E=k-r\). Then the local likelihood-ratio statistic \(q=-2\left[\ell_{n}(\hat{\theta}_{0,n}) -\ell_{n}(\hat{\theta}_{n})\right]\) built from the estimators Equation (A.403) converges in distribution to \(\chi^{2}_{r}\). Rests on Lemma A.213, Lemma 11.83 and Theorem A.206.
Derives Theorem A.214. Write \(D_{n}\) as in Lemma A.213, so that \(q=D_{n}(\hat{\theta}_{0,n})-D_{n}(\hat{\theta}_{n})\) — the common reference value \(\ell_{n}(\theta_{0})\) cancels. By Corollary A.210 each \(D_{n}\) equals the corresponding \(\Lambda_{n}\) up to \(\epsilon_{n}(K)\), and by Lemma A.213 each \(\Lambda_{n}\) equals the corresponding minimum up to a quantity tending to zero in probability. Hence
From Equations (A.414) and (A.415) (replace each statistic by the minimum of the quadratic form and cancel the common constant).
The second equality is Equation (A.415), whose additive constant \(-\vect{Z}_{n}\transpose I^{-1}\vect{Z}_{n}\) is the same in both minima and cancels.
Now substitute \(\vect{V}_{n}=I^{1/2}\vect{m}=I^{-1/2}\vect{Z}_{n}\) and \(S=I^{1/2}E\), where \(I^{1/2}\) is any invertible matrix with \(I^{1/2}\left(I^{1/2}\right)\transpose=I\), symmetric here; the existence of one is the weak form of the spectral theorem recorded in Remark 11.86. Since \(I^{1/2}\) is invertible, \(\dim S=\dim E=k-r\), and
with \(P\) the orthogonal projector onto \(S^{\perp}\), of dimension \(r\): the nearest point of a subspace to a given vector is its orthogonal projection, and what is left over is the projection onto the orthogonal complement.
By Lemma A.208, \(\vect{Z}_{n}\rightarrow \mathcal{N}(0,I)\) in distribution; applying Theorem A.206 to the linear map \(\vect{z}\mapsto I^{-1/2}\vect{z}\) — for each \(\vect{a}\), \(\vect{a}\transpose I^{-1/2}\vect{Z}_{n}\rightarrow \mathcal{N}(0,\vect{a}\transpose I^{-1/2}II^{-1/2}\vect{a}) =\mathcal{N}(0,\norm{\vect{a}}^{2})\) — gives \(\vect{V}_{n}\rightarrow\vect{V}\sim\mathcal{N}(0,\identity)\) in distribution. The map \(\vect{v}\mapsto\norm{P\vect{v}}^{2}\) is continuous, so the continuous-mapping clause of Theorem A.206 gives \(\norm{P\vect{V}_{n}}^{2}\rightarrow\norm{P\vect{V}}^{2}\) in distribution, and \(\norm{P\vect{V}}^{2}\sim\chi^{2}_{r}\) by Lemma 11.83. Finally, adding an \(o_{P}(1)\) term does not change a limit in distribution, by the sum clause of Lemma 11.52. Hence \(q\rightarrow\chi^{2}_{r}\).
∎Equation (11.80) defines \(q\) through the global suprema of the likelihood over \(\Theta\) and over \(\Theta_{0}\), while Theorem A.214 is proved for the local maximisers Equation (A.403). The gap is the same one Remark 11.69 records for Theorem 11.67: identifying a consistent local maximiser with the global one is Wald's theorem [Wald:1949], under conditions this treatise does not assume, and a likelihood with several local maxima can have a consistent root far from its highest peak at any finite \(n\). Stating the theorem for the local statistic is therefore not a weakening introduced here — it is the same statement the chapter's one-parameter derivation proves, made explicit. The number \(r\) is unaffected: it is \(\operatorname{codim}E\), a property of the constraint alone.
It is worth marking the two places, because they are the whole content of this appendix. Uniformity — item (i) — is Equation (A.406): the remainder is bounded by \(\tfrac16nk^{3}\bar{M}_{n}\norm{\delta}^{3}\) with a coefficient that does not depend on \(\theta\), which is what (R5)'s single dominating function \(M\) buys and what a pointwise Taylor expansion would not. It is used at two moving points, \(\hat{\theta}_{n}\) and \(\hat{\theta}_{0,n}\), neither of which is known in advance; a \(\theta\)-dependent remainder would say nothing about either. Localisation — item (ii) — is Equation (A.410): both estimators sit within \(O_{P}(n^{-1/2})\) of \(\theta_{0}\), so the surface is only ever probed at distance \(O(n^{-1/2})\), where Lemma A.212 bounds its departure from the tangent space by \(O(n^{-1})\) — one order smaller, which after the \(\sqrt{n}\) rescaling of Equation (A.417) is the \(O(n^{-1/2})\) that vanishes. That is the honest content of writing \(\Theta_{0}\simeq\theta_{0}+E\), and it is why second-order smoothness of the constraint, and not merely differentiability, is assumed.
Wilks' Theorem in the Multi-Parameter Case discharges the proof obligation left at Section 11.5.4 of Probability and Statistics. With it, Corollary 11.89 — one parameter of interest, any number of nuisance parameters, one degree of freedom — is a corollary of a proved theorem rather than of an argued one, and so is the practice it licenses: a collider search reads its significance off a one-degree-of-freedom distribution however many nuisance parameters describe the background and the detector [Cowan:2011]. The two failure modes remain exactly as Remark 11.90 and Remark 11.91 describe them, and both are visible in the proof above: a parameter on the boundary violates (R3), so Equation (A.403) is no longer an interior maximisation and Lemma A.211 loses the two-sided estimate; a parameter unidentified under the null makes \(I\) singular, so \(\lambda=0\) and every step from Equation (A.412) onward fails at once.
Rice's Formula for the Expected Number of Upcrossings
This appendix proves Theorem 11.102 of Probability and Statistics: for a smooth Gaussian field of unit variance the expected number of upcrossings of a level \(z\) factorises into \(\ee^{-z^{2}/2}/(2\pi)\) times an integral of the standard deviation of the derivative, so that the level enters only through a Gaussian factor. That factorisation is what the whole look-elsewhere machinery of Section 11.6 rests on: it is what makes Theorem 11.103 true, and hence what allows a search to count upcrossings at a low threshold in a few hundred simulated experiments and extrapolate to a threshold at which nothing could ever be simulated. The formula is section 3.3 of the second instalment of Rice's memoir on random noise [Rice:1945]; the first instalment [Rice:1944] carries the shot effect and the power spectra and not this.
The proof is in three moves. A purely deterministic counting identity expresses the number of upcrossings of a \(C^{1}\) function as the limit of a normalised integral over the set where the function is near the level — Kac's device (Kac's counting integral). The expectation of that integral factorises, because a field of constant variance is uncorrelated with its own derivative at each point, and for jointly Gaussian variables uncorrelated means independent (The joint law of the field and its derivative). And the two elementary Gaussian integrals that remain are evaluated (Evaluation). The passage from the integral identity to its expectation is where Lebesgue theory enters, and it is the only thing quoted; it is stated precisely in Remark A.224.
Hypotheses and statement
Let \(M=[m_{1},m_{2}]\) be a compact interval and \(X\) a real-valued process on \(M\) satisfying:
-
[(H1)] Gaussian, standardised: every finite linear combination \(\sum_{j}a_{j}X(n_{j})\), \(n_{j}\in M\), is Gaussian with mean zero; and the covariance function \(r(n,n')=\avg{X(n)X(n')}\) satisfies \(r(m,m)=1\) for every \(m\in M\).
-
[(H2)] Smooth covariance: \(r\) is twice continuously differentiable on \(M\times M\).
-
[(H3)] Smooth paths: almost every sample path of \(X\) is continuously differentiable on \(M\), and for each \(m\) the difference quotients \(\left(X(m+h)-X(m)\right)/h\) converge to \(X'(m)\) in mean square as \(h\rightarrow0\).
-
[(H4)] Non-degeneracy and non-tangency: writing \(\sigma_{1}(m)^{2}=\operatorname{var}X'(m)\), the function \(\sigma_{1}\) is continuous and strictly positive on \(M\); and for the level \(z\) under consideration, almost surely there is no \(m\in M\) with \(X(m)=z\) and \(X'(m)=0\).
Under Definition A.218, the expected number of upcrossings of the level \(z\) by \(X\) on \(M\), in the sense of Definition 11.100, is
The two sides carry the same SI dimension, namely none. The field is standardised, so \(X\) and the level \(z\) are pure numbers; the parameter \(m\) carries whatever unit the scanned quantity carries — in the application of Section 11.6 it is a mass — and \(X'\), being a derivative with respect to it, carries the reciprocal of that unit. So \(\sigma_{1}(m)\,\dd m\) is dimensionless, as a count must be. This is also the quickest check that no factor has been lost: the level enters only through \(\ee^{-z^{2}/2}\), which could carry no unit, and the geometry only through \(\int\sigma_{1}\,\dd m\), which is the total number of “independent looks” the scan contains.
The joint law of the field and its derivative
Under (H1)–(H4), for each fixed \(m\in M\) the pair \(\left(X(m),X'(m)\right)\) is jointly Gaussian with mean zero and covariance matrix \(\operatorname{diag}\left(1,\sigma_{1}(m)^{2}\right)\). The two components are therefore independent, with joint density
Rests on Definition A.218 and Corollary A.178.
Derives Lemma A.220. Joint Gaussianity. Fix \(a,b\in\R\) and write \(D_{h}=\left(X(m+h)-X(m)\right)/h\), a linear combination of two values of \(X\); by (H1) the variable \(S_{h}=aX(m)+bD_{h}\) is Gaussian with mean zero, say \(S_{h}\sim\mathcal{N}(0,v_{h})\). Put \(S=aX(m)+bX'(m)\). By (H3), \(D_{h}\rightarrow X'(m)\) in mean square, so \(S_{h}\rightarrow S\) in mean square and hence in probability, and therefore in distribution. Mean-square convergence also forces \(v_{h}=\avg{S_{h}^{2}}\rightarrow\avg{S^{2}}=v\), since \(\abs{\avg{S_{h}^{2}}^{1/2}-\avg{S^{2}}^{1/2}} \le\avg{(S_{h}-S)^{2}}^{1/2}\) by the triangle inequality for the second-moment norm. The characteristic function of \(S_{h}\) is \(\ee^{-v_{h}t^{2}/2}\) (Proposition 11.29 and the computation of the Gaussian transform in the derivation of Theorem 11.46), which converges to \(\ee^{-vt^{2}/2}\) for every \(t\). By Corollary A.181 the characteristic function of \(S\) is that limit, and by Corollary A.178 \(S\) is therefore \(\mathcal{N}(0,v)\). As \(a\) and \(b\) were arbitrary, the pair is jointly Gaussian.
The covariance vanishes. By the Cauchy–Schwarz inequality (Lemma 11.63, applied without centring, both variables having mean zero), \(\abs{\avg{X(m)\left(D_{h}-X'(m)\right)}} \le\avg{X(m)^{2}}^{1/2}\avg{\left(D_{h}-X'(m)\right)^{2}}^{1/2} \rightarrow0\), so
where \(\pp_{2}\) is the derivative in the second slot, which exists and is continuous by (H2). Now \(r(m,m)=1\) for every \(m\) by (H1); differentiating this identity in \(m\) by the chain rule (Proposition 7.72) gives \(\pp_{1}r(m,m)+\pp_{2}r(m,m)=0\), while \(r\) is symmetric, \(r(n,n')=r(n',n)\), so \(\pp_{1}r(m,m)=\pp_{2}r(m,m)\). Hence both vanish, and \(\avg{X(m)X'(m)}=0\) by Equation (A.422). Together with \(\operatorname{var}X(m)=r(m,m)=1\) and \(\operatorname{var}X'(m)=\sigma_{1}(m)^{2}>0\) from (H4), the covariance matrix is \(\operatorname{diag}(1,\sigma_{1}^{2})\), which is invertible.
Independence. A jointly Gaussian pair with mean zero and invertible covariance \(\Sigma\) has density \(\left(2\pi\right)^{-1}\left(\det\Sigma\right)^{-1/2} \exp\left(-\tfrac12\vect{y}\transpose\Sigma^{-1}\vect{y}\right)\); with \(\Sigma\) diagonal the quadratic form splits and the density factorises as Equation (A.421), which by Definition 11.25 is independence.
∎The lemma is the structural heart of Equation (A.420) and is worth stating in words, because it is what Theorem 11.103 exploits: constancy of the variance is what makes field and derivative independent, and once they are, the level \(z\) can influence only the first factor of Equation (A.421) while the derivative supplies the same average slope at every level. Had the field not been standardised, the two would be correlated wherever \(\dd\operatorname{var}X/\dd m\neq0\) and no such factorisation would exist.
Kac's counting integral
The next lemma contains no probability at all. It is the observation that a transversal upcrossing contributes exactly \(2\delta\) to \(\int\indic{\abs{g-z}<\delta}(g')^{+}\), whatever \(\delta\) is, while a downcrossing contributes nothing.
Let \(g:M\rightarrow\R\) be continuously differentiable and let \(z\in\R\) satisfy \(g(m_{1})\neq z\), \(g(m_{2})\neq z\), and \(g'(m)\neq0\) at every \(m\) with \(g(m)=z\). Then \(g\) has finitely many upcrossings of \(z\), their number \(N_{z}(g)\) in the sense of Definition 11.100 is finite, and there is \(\delta_{0}(g)>0\) with
where \(y^{+}=\max(y,0)\). In particular the left side converges to \(N_{z}(g)\) as \(\delta\downarrow0\). Rests on Definition 11.100, Theorem 7.43 and Theorem 7.24.
Derives Lemma A.221. Finiteness. Let \(Z=\set{m\in M\mid g(m)=z}\), a closed subset of the compact \(M\) and hence compact; by hypothesis it misses both endpoints. If \(p\in Z\) then \(g'(p)\neq0\), and \(g'\) is continuous, so \(g\) is strictly monotone on a neighbourhood of \(p\) and takes the value \(z\) there only at \(p\): every point of \(Z\) is isolated. A compact set all of whose points are isolated is finite — the singletons form an open cover with no proper subcover. Write \(Z=\set{p_{1}<\dots<p_{J}}\), all interior to \(M\).
A separating radius. For each \(j\) choose \(\eta>0\), the same for all \(j\), small enough that the closed intervals \(U_{j}=[p_{j}-\eta,p_{j}+\eta]\) are pairwise disjoint, contained in \((m_{1},m_{2})\), and such that \(g'\) keeps the sign of \(g'(p_{j})\) throughout \(U_{j}\); continuity of \(g'\) and finiteness of \(Z\) make this possible. On \(M\setminus\bigcup_{j}\left(p_{j}-\eta,p_{j}+\eta\right)\) — a compact set on which \(g-z\) does not vanish — the continuous function \(\abs{g-z}\) attains a strictly positive minimum \(\mu\) (Theorem 7.24). Let \(\delta_{0}=\min\left(\mu,\;\min_{j}\abs{g(p_{j}-\eta)-z},\; \min_{j}\abs{g(p_{j}+\eta)-z}\right)\), which is positive because \(g\) is strictly monotone on each \(U_{j}\) with \(g(p_{j})=z\).
The integral, for \(\delta<\delta_{0}\). The set \(A_{\delta}=\set{m\in M\mid\abs{g(m)-z}<\delta}\) is contained in \(\bigcup_{j}U_{j}\), since \(\delta<\mu\). Fix \(j\). On \(U_{j}\) the function \(g\) is strictly monotone and continuous, so \(A_{\delta}\cap U_{j}\) is the open interval \(\left(\alpha_{j},\beta_{j}\right)\) whose endpoints are the two solutions of \(g=z\mp\delta\) inside \(U_{j}\); these exist and lie in the interior of \(U_{j}\) by the intermediate value theorem (Theorem 7.23) and the choice \(\delta<\delta_{0}\).
If \(g'>0\) on \(U_{j}\) — an upcrossing at \(p_{j}\), by Definition 11.100, since \(g\) passes from below \(z\) to above it there — then \(\left(g'\right)^{+}=g'\) on \(U_{j}\) and, by the second fundamental theorem of calculus (Theorem 7.43),
If instead \(g'<0\) on \(U_{j}\), then \(\left(g'\right)^{+}=0\) there and the contribution is \(0\); that is a downcrossing and is correctly not counted. Summing Equation (A.424) over the finitely many cells, the integral in Equation (A.423) equals \(2\delta\) times the number of \(j\) with \(g'(p_{j})>0\), which is \(N_{z}(g)\); dividing by \(2\delta\) gives the identity, with a value independent of \(\delta\).
∎The three conditions imposed on \(g\) are exactly what (H4) and the continuity of the one-dimensional Gaussian law provide for almost every path. That \(X(m_{i})\neq z\) almost surely is immediate: \(X(m_{i})\) is standard Gaussian and \(\Pr(X(m_{i})=z)=0\) by Proposition 11.22. That no crossing is tangential is the second half of (H4), and it is assumed rather than derived — see Remark A.224.
Taking the expectation
Let \(h(m,\omega)\ge0\) be jointly measurable on \(M\times\Omega\). Then \(\omega\mapsto\int_{M}h(m,\omega)\dd m\) and \(m\mapsto\avg{h(m,\,\cdot\,)}\) are measurable and
both sides possibly infinite. And if \(Y_{\delta}\rightarrow Y\) almost surely as \(\delta\downarrow0\) with \(\abs{Y_{\delta}}\le D\) for a single \(D\) with \(\avg{D}<\infty\), then \(\avg{Y_{\delta}}\rightarrow\avg{Y}\) [Billingsley:1995]. Rests on Definition 11.4.
Two things, and they are the only ones.
First, Theorem A.223. Real Analysis builds the Riemann–Darboux integral and no more, so neither Tonelli's theorem nor dominated convergence is available from anything earlier in this treatise; both belong to Lebesgue theory, and the same two statements are quoted for the same reason in Remark A.171. The interchange Equation (A.425) is unavoidable here: the object being averaged is an integral over \(M\) of a quantity that depends on the whole path, and there is no elementary truncation — of the kind used at Equation (11.32) to avoid dominated convergence in Lemma 11.44 — that reaches it. The dominated-convergence step is the passage from Equation (A.423), which holds path by path only for \(\delta<\delta_{0}(g)\) with a threshold that depends on the path and has no positive lower bound over all paths, to a single limit under the expectation. The dominating variable is \(D=\sup_{0<\delta<1}\left(2\delta\right)^{-1}\int_{M} \indic{\abs{X-z}<\delta}\left(X'\right)^{+}\dd m\), whose integrability is part of what is being assumed.
Second, the non-tangency half of (H4). It is stated as a hypothesis rather than derived. It is known to follow from \(\sigma_{1}^{2}>0\) together with the smoothness already assumed, by an argument (Bulinskaya's lemma) that bounds the probability of the field being near \(z\) while its derivative is near \(0\) and requires the same Lebesgue theory; this treatise does not build it, and no elementary substitute is offered. It is not a vacuous assumption: a field with a genuine tangency at level \(z\) would touch the level without crossing it, and Lemma A.221 would fail on that path — the set \(\set{\abs{g-z}<\delta}\) would not be a finite union of monotone cells.
Everything else — the joint law Lemma A.220, the counting identity Lemma A.221, the two Gaussian integrals of Evaluation, and the assembly — is derived here.
Evaluation
If \(W\sim\mathcal{N}(0,\sigma^{2})\) with \(\sigma>0\), then \(\avg{W^{+}}=\sigma/\sqrt{2\pi}\). Rests on Definition 11.28 and Equation (7.27).
Derives Lemma A.225. By Equation (11.20) and the substitution \(w=\sigma y\) (Equation (7.27)),
Write \(\varphi(v)=\ee^{-v^{2}/2}/\sqrt{2\pi}\) for the standard Gaussian density. Then for every \(z\in\R\) and \(\delta>0\) there is \(\zeta_{\delta}\in[z-\delta,z+\delta]\) with
and the bound does not depend on \(m\). Rests on Definition 11.28 and Theorem 7.35.
Derives Lemma A.226. By (H1) the variable \(X(m)\) is standard Gaussian for every \(m\), so \(\Pr(\abs{X(m)-z}<\delta)=\int_{z-\delta}^{z+\delta}\varphi(v)\dd v\), independently of \(m\). Since \(\varphi\) is continuous, the mean value theorem for integrals — apply Theorem 7.35 to the antiderivative \(\Phi\), whose derivative is \(\varphi\) by Proposition 11.22 — gives \(\int_{z-\delta}^{z+\delta}\varphi=2\delta\,\varphi(\zeta_{\delta})\) for some \(\zeta_{\delta}\) in the interval. Applying Theorem 7.35 once more, this time to \(\varphi\) itself, \(\abs{\varphi(\zeta_{\delta})-\varphi(z)}\le\delta\sup_{v} \abs{\varphi'(v)}\), and \(\varphi'(v)=-v\varphi(v)\) has modulus maximised at \(v=\pm1\), where it equals \(\ee^{-1/2}/\sqrt{2\pi}\).
∎Proof of Theorem A.219. Derives Theorem A.219. By (H3) and Remark A.222, almost every sample path satisfies the hypotheses of Lemma A.221, so almost surely
Take expectations. The dominated-convergence clause of Theorem A.223 gives \(\avg{N_{z}(X)}=\lim_{\delta\downarrow0}\avg{Y_{\delta}}\), and the Tonelli clause — the integrand is non-negative — gives, for each fixed \(\delta\),
At each fixed \(m\) the indicator is a function of \(X(m)\) alone and the positive part a function of \(X'(m)\) alone, and those two are independent by Lemma A.220; so the expectation factorises, and by Lemmas A.225 and A.226,
From Equations (A.421) and (A.428) (factorise the expectation at each \(m\) by independence, then evaluate the level factor and the mean positive part). The point \(\zeta_{\delta}\) does not depend on \(m\) (Lemma A.226), so it comes outside the integral:
the integral being finite because \(\sigma_{1}\) is continuous on the compact \(M\) (H4) and therefore bounded and Riemann integrable (Theorem 7.40). Letting \(\delta\downarrow0\) and using the uniform estimate of Equation (A.426),
since \(\varphi(z)/\sqrt{2\pi}=\ee^{-z^{2}/2}/(2\pi)\). That is Equation (A.420).
∎If in addition \(X\) is stationary, so that \(r(n,n')\) depends only on \(n-n'\), then \(\sigma_{1}\) is the constant \(\sqrt{\lambda_{2}}\) with \(\lambda_{2}=-r''(0)\), and
Rests on Theorem A.219.
Derives Corollary A.227. With \(D_{h}\) as in Lemma A.220, the second-moment norm is continuous under mean-square convergence, so \(\avg{D_{h}D_{h'}}\rightarrow\avg{X'(m)^{2}}\) as \(h,h'\rightarrow0\). On the other hand
a second-order difference quotient of the twice continuously differentiable \(r\), which converges to \(\pp_{1}\pp_{2}r(m,m)\). Hence \(\sigma_{1}(m)^{2}=\pp_{1}\pp_{2}r(m,m)\) in general. Now write \(r(n,n')=\varrho(n-n')\) with \(\varrho\) twice continuously differentiable by (H2) and \(\varrho(0)=1\); then \(\pp_{1}r=\varrho'(n-n')\) and \(\pp_{2}\pp_{1}r=-\varrho''(n-n')\), so \(\sigma_{1}(m)^{2}=-\varrho''(0)=\lambda_{2}\), independent of \(m\), and the integral in Equation (A.420) is \(L\sqrt{\lambda_{2}}\).
∎Rice's Formula for the Expected Number of Upcrossings discharges the proof obligation of Theorem 11.102 in Section 11.6.2 of Probability and Statistics, the last of the results that section had to import. With it, the chain the look-elsewhere effect rests on is complete: Davies' bound Theorem 11.101 was already derived there from the intermediate value theorem and Markov's inequality; Theorem 11.103 — the exponential level dependence \(\avg{N_{u}}=\avg{N_{u_{0}}}\ee^{-(u-u_{0})/2}\), which is what lets a search estimate a trials factor from upcrossings counted at a one-sigma threshold — follows from Equation (A.420) at \(z=\sqrt{u}\); and Corollary 11.105 follows from those two. What is assumed and not proved is named in Remark A.224: two theorems of Lebesgue theory, and the absence of tangential crossings. The extension of Theorem 11.103 to a scan over several parameters remains, as Remark 11.104 says, reported and not derived — and nothing in this treatise uses it.
The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators
This appendix proves Theorem 12.44 of Hilbert Spaces: a compact self-adjoint operator on a Hilbert space has an orthonormal system of eigenvectors, with real eigenvalues accumulating only at \(0\), in terms of which the operator is a norm- convergent sum of rank-one pieces. It is the one infinite-dimensional theorem that reproduces the finite-dimensional spectral theorem Theorem 5.75 without weakening any of its conclusions, and it is the theorem Hilbert proved for symmetric integral kernels [Hilbert:1912] and Schmidt recast in the language used here [Schmidt:1907].
Everything below rests on material already available: the numerical-radius formula \(\norm{A}=\sup_{\norm{x}=1}\abs{\braket{x}{Ax}}\) (Proposition 12.43), Bessel's inequality (Proposition 12.27), the projection theorem (Theorem 12.18), and the equivalence of compactness and sequential compactness in a metric space (Theorem 6.31). No weak topology is used; the closing Remark A.236 says precisely what the standard weak-compactness argument would have supplied and why the elementary route replaces it.
Statement
Let \(\mathcal{H}\neq\set{0}\) be a Hilbert space and let \(A\in\mathcal{B}(\mathcal{H})\) be compact and self-adjoint (Definition 12.41). Then there is a finite or countable orthonormal system \(\set{e_{n}}\) in \(\mathcal{H}\) and real numbers \(\lambda_{n}\neq0\) with \(\abs{\lambda_{1}}\geq\abs{\lambda_{2}}\geq\cdots\) such that
the series converging in the norm of \(\mathcal{H}\). If the system is infinite then \(\lambda_{n}\longrightarrow0\). Every eigenspace belonging to a non-zero eigenvalue is finite dimensional, \(Ae_{n}=\lambda_{n}e_{n}\), and \(\norm{A}=\max_{n}\abs{\lambda_{n}}\) when \(A\neq0\). Finally
so that adjoining any orthonormal basis of \(\ker A\) to \(\set{e_{n}}\) produces an orthonormal basis of \(\mathcal{H}\) consisting of eigenvectors of \(A\). Rests on Definition 12.41, Proposition 12.43 and Theorem 12.30.
The proof occupies the rest of the section: the sequential form of compactness (Compactness in sequential form), the attainment of the supremum and the first eigenvector (The supremum is attained: the first eigenvector), the recursion (The recursion), and the passage to the limit (Passage to the limit).
Compactness in sequential form
\(A\in\mathcal{B}(\mathcal{H})\) is compact if and only if every bounded sequence \((x_{n})\) in \(\mathcal{H}\) admits a subsequence \((x_{n_{k}})\) for which \((Ax_{n_{k}})\) converges in \(\mathcal{H}\). Rests on Definition 12.41 and Theorem 6.31.
Derives Lemma A.230. Suppose \(A\) is compact and let \(\norm{x_{n}}\leq c\) for all \(n\); we may assume \(c>0\). The vectors \(Ax_{n}/c\) lie in the image of the closed unit ball, hence in the set \(K=\overline{A\set{x\mid\norm{x}\leq1}}\), which is compact by hypothesis. A compact subset of the metric space \(\mathcal{H}\) is sequentially compact (Theorem 6.31), so some subsequence \(Ax_{n_{k}}/c\) converges in \(K\), and therefore \(Ax_{n_{k}}\) converges.
Conversely, assume the sequential property and let \((y_{m})\) be a sequence in \(K\). Each \(y_{m}\) is a limit of points \(Ax\) with \(\norm{x}\leq1\), so we may choose \(x_{m}\) with \(\norm{x_{m}}\leq1\) and \(\norm{y_{m}-Ax_{m}}<1/m\). By hypothesis some \(Ax_{m_{k}}\) converges, and then \(y_{m_{k}}\) converges to the same limit, which lies in the closed set \(K\). Hence \(K\) is sequentially compact, hence compact by Theorem 6.31.
∎Let \(A\) be compact and self-adjoint and let \(M\subseteq\mathcal{H}\) be a closed subspace with \(AM\subseteq M\). Then \(M\) is a Hilbert space, and the restriction \(A|_{M}\) is a compact self-adjoint operator on \(M\) with \(\norm{A|_{M}}\leq\norm{A}\). Moreover \(AM^{\perp}\subseteq M^{\perp}\). Rests on Lemma A.230, Theorem 12.18 and Definition 12.41.
Derives Lemma A.231. A closed subspace of a complete metric space is complete, so \(M\) is a Hilbert space with the inherited inner product. Symmetry of \(A|_{M}\) is inherited, and by Proposition 12.42 a symmetric operator defined on all of the Hilbert space \(M\) is self-adjoint there. Compactness is Lemma A.230: a bounded sequence in \(M\) is a bounded sequence in \(\mathcal{H}\), so \((Ax_{n})\) has a convergent subsequence, whose limit lies in the closed set \(M\). The norm bound is immediate, the supremum being taken over a smaller set. For the last claim, let \(y\in M^{\perp}\) and \(z\in M\); then \(Az\in M\), so \(\braket{z}{Ay}=\braket{Az}{y}=0\), and \(Ay\in M^{\perp}\).
∎The supremum is attained: the first eigenvector
This is the step Hilbert Spaces left open. The numerical-radius formula Equation (12.26) exhibits \(\norm{A}\) as a supremum; for a compact operator the supremum is a maximum, and the maximising vector is an eigenvector.
Let \(A\in\mathcal{B}(\mathcal{H})\) be compact and self-adjoint with \(A\neq0\). Then there are \(e\in\mathcal{H}\) with \(\norm{e}=1\) and \(\lambda\in\R\) with \(\abs{\lambda}=\norm{A}\) such that \(Ae=\lambda e\). Rests on Proposition 12.43, Lemma A.230 and Proposition 12.42.
Derives Lemma A.232. By Proposition 12.43 there are unit vectors \(x_{n}\) with
Each number \(\braket{x_{n}}{Ax_{n}}\) is real, because \(A\) is self-adjoint (Proposition 12.42). A real sequence whose absolute values converge to \(\norm{A}\) has a subsequence converging to \(+\norm{A}\) or to \(-\norm{A}\): infinitely many of its terms are non-negative or infinitely many are negative, and along the corresponding indices the absolute value carries the sign. Relabel, and let
Expand, using \(\lambda\) real, \(\norm{x_{n}}=1\), and \(\braket{Ax_{n}}{x_{n}}=\braket{x_{n}}{Ax_{n}}\):
Now \(\norm{Ax_{n}}\leq\norm{A}=\abs{\lambda}\), so the right-hand side of Equation (A.437) is at most \(2\lambda^{2}-2\lambda\braket{x_{n}}{Ax_{n}}\), which tends to \(0\) by Equation (A.436). Hence
This much uses only self-adjointness; it says that \(\lambda\) is an approximate eigenvalue. Compactness converts the approximation into an eigenvector.
The sequence \((x_{n})\) is bounded, so by Lemma A.230 there is a subsequence with \(Ax_{n_{k}}\longrightarrow y\) for some \(y\in\mathcal{H}\). Since \(\lambda\neq0\) we may write
by Equation (A.438). The norm is continuous (Corollary 12.5), so \(\norm{e}=\lim\norm{x_{n_{k}}} =1\); and \(A\) is continuous (Proposition 12.36), so \(Ax_{n_{k}}\longrightarrow Ae\). But \(Ax_{n_{k}}\longrightarrow y=\lambda e\). Limits in a metric space are unique, so \(Ae=\lambda e\).
∎The recursion
Let \(A\) be compact and self-adjoint. There are a finite or countable orthonormal system \(\set{e_{n}}\) and real numbers \(\lambda_{n}\neq0\) with \(Ae_{n}=\lambda_{n}e_{n}\) and \(\abs{\lambda_{1}}\geq\abs{\lambda_{2}}\geq\cdots\), such that, writing \(\mathcal{H}_{1}=\mathcal{H}\) and
one has \(\norm{A_{N+1}}=\abs{\lambda_{N+1}}\) at every stage at which the construction continues, and the construction stops at stage \(N\) exactly when \(A_{N}=0\), i.e. when \(\mathcal{H}_{N}\subseteq\ker A\). Rests on Lemma A.232, Lemma A.231 and Theorem 12.18.
Derives Lemma A.233. Set \(\mathcal{H}_{1}=\mathcal{H}\) and \(A_{1}=A\). Suppose \(\mathcal{H}_{N}\) and \(A_{N}\) have been defined, with \(\mathcal{H}_{N}\) a closed \(A\)-invariant subspace and \(A_{N}=A|_{\mathcal{H}_{N}}\) compact and self-adjoint on it (Lemma A.231). If \(A_{N}=0\) the construction stops, and then every \(x\in\mathcal{H}_{N}\) has \(Ax=0\). If \(A_{N}\neq0\), apply Lemma A.232 inside the Hilbert space \(\mathcal{H}_{N}\): it produces a unit vector \(e_{N}\in\mathcal{H}_{N}\) and a real \(\lambda_{N}\) with \(\abs{\lambda_{N}}=\norm{A_{N}}\neq0\) and \(Ae_{N}=A_{N}e_{N}=\lambda_{N}e_{N}\).
The vector \(e_{N}\) lies in \(\mathcal{H}_{N}\), hence is orthogonal to \(e_{1},\dots,e_{N-1}\), so the system stays orthonormal. Put \(\mathcal{H}_{N+1}=\mathcal{H}_{N}\cap\set{e_{N}}^{\perp} =\set{e_{1},\dots,e_{N}}^{\perp}\), which is closed (Proposition 12.17). It is \(A\)-invariant: for \(x\in\mathcal{H}_{N+1}\) and \(j\leq N\),
using self-adjointness and \(\lambda_{j}\in\R\); and \(Ax\in\mathcal{H}_{N}\) because \(\mathcal{H}_{N}\) is invariant. Finally \(\norm{A_{N+1}}\leq\norm{A_{N}}\) because \(A_{N+1}\) is a restriction of \(A_{N}\), whence \(\abs{\lambda_{N+1}}\leq\abs{\lambda_{N}}\); the construction therefore delivers a non-increasing sequence of moduli, and it is exactly the statement \(\norm{A_{N}}=\abs{\lambda_{N}}\) that Lemma A.232 supplies at each stage.
∎If the construction of Lemma A.233 does not stop, then \(\lambda_{n}\longrightarrow0\). Moreover, for every \(\lambda\neq0\) the eigenspace \(\ker(A-\lambda\identity)\) is finite dimensional. Rests on Lemmas A.230 and A.233.
Derives Lemma A.234. Both statements follow from one computation. Let \((f_{n})\) be an orthonormal sequence with \(Af_{n}=\nu_{n}f_{n}\). For \(n\neq m\), orthogonality gives
Eigenvalues. Apply this to \(f_{n}=e_{n}\), \(\nu_{n}=\lambda_{n}\). If \((\lambda_{n})\) does not tend to \(0\) there are \(\varepsilon>0\) and infinitely many indices with \(\abs{\lambda_{n}}\geq\varepsilon\); along them Equation (A.441) gives \(\norm{Ae_{n}-Ae_{m}}^{2}\geq2\varepsilon^{2}\), so \((Ae_{n})\) has no Cauchy subsequence, hence no convergent subsequence. Since \((e_{n})\) is bounded, this contradicts Lemma A.230. As the moduli are non-increasing, \(\lambda_{n}\longrightarrow0\).
Multiplicity. If \(\ker(A-\lambda\identity)\) were infinite dimensional for some \(\lambda\neq0\), Gram–Schmidt (Proposition 12.23) would produce an infinite orthonormal sequence \((f_{n})\) inside it, all with \(\nu_{n}=\lambda\), and Equation (A.441) would give \(\norm{Af_{n}-Af_{m}}^{2}=2\abs{\lambda}^{2}>0\): the same contradiction.
∎Passage to the limit
Proof of Theorem A.229. Derives Theorem A.229. Let \(\set{e_{n}}\), \(\set{\lambda_{n}}\) be the system of Lemma A.233. Fix \(x\in\mathcal{H}\) and put
Then \(\braket{e_{j}}{x_{N}}=\braket{e_{j}}{x}-\braket{e_{j}}{x}=0\) for \(j\leq N\), so \(x_{N}\in\mathcal{H}_{N+1}\), and Bessel's inequality (Proposition 12.27) gives
Applying \(A\) to Equation (A.442) and using \(Ae_{n}=\lambda_{n}e_{n}\),
the last equality because \(x_{N}\in\mathcal{H}_{N+1}\). Therefore, by Lemma A.233 and Equation (A.443),
If the construction stops at stage \(N\) then \(A_{N}=0\) and Equation (A.444) already gives Equation (A.433) as a finite sum. If it does not stop, \(\abs{\lambda_{N+1}}\longrightarrow0\) by Lemma A.234, so the right-hand side of Equation (A.445) tends to \(0\) and Equation (A.433) holds with convergence in norm. The eigenvalue relation \(Ae_{n}=\lambda_{n}e_{n}\), the ordering of the moduli and the finiteness of the non-zero eigenspaces are Lemmas A.233 and A.234. For \(A\neq0\), \(\norm{A}=\norm{A_{1}}=\abs{\lambda_{1}} =\max_{n}\abs{\lambda_{n}}\), the maximum being attained because the moduli decrease.
It remains to prove Equation (A.434). Write \(M=\overline{\text{span}\set{e_{n}}}\), a closed subspace, so that \(\mathcal{H}=M\oplus M^{\perp}\) by Theorem 12.18. If \(x\in M^{\perp}\) then \(\braket{e_{n}}{x}=0\) for every \(n\), and Equation (A.433) gives \(Ax=0\): thus \(M^{\perp}\subseteq\ker A\). Conversely if \(Ax=0\) then for every \(n\),
and \(\lambda_{n}\neq0\) forces \(\braket{e_{n}}{x}=0\), so \(x\in M^{\perp}\); hence \(\ker A=M^{\perp}\), which is Equation (A.434). Every vector of \(\ker A\) is an eigenvector with eigenvalue \(0\), so an orthonormal basis of \(\ker A\) adjoined to \(\set{e_{n}}\) is an orthonormal system of eigenvectors whose closed span is \(M\oplus\ker A=\mathcal{H}\): an orthonormal basis of \(\mathcal{H}\) in the sense of Definition 12.29.
∎The last sentence of Theorem A.229 needs an orthonormal basis of the closed subspace \(\ker A\), and that is the only point at which the theorem is not constructive. If \(\mathcal{H}\) is separable — the case of every physical application, and the standing assumption of Theorem 12.33 — so is \(\ker A\), and Gram–Schmidt applied to a countable dense set produces the basis (Corollary 12.24). For a non-separable \(\mathcal{H}\) the existence of an orthonormal basis of \(\ker A\) is a maximality statement proved with Zorn's lemma (Logic, Sets, and Maps); the system \(\set{e_{n}}\) itself, on which the whole content of the theorem rests, is always countable and always constructed, never chosen.
The route usually taken to Lemma A.232 [Reed:1972] passes through two statements about the weak topology, neither of which this treatise develops: that the closed unit ball of a Hilbert space is weakly sequentially compact, so that a maximising sequence has a weakly convergent subsequence \(x_{n}\rightharpoonup e\); and that a compact operator carries weakly convergent sequences to norm convergent ones, so that \(Ax_{n}\longrightarrow Ae\) and the supremum is attained at \(e\). Both are true and both are proved in [Reed:1972], chapter VI.
The proof given above needs neither. The reason is Equation (A.438): self-adjointness alone forces the maximising sequence to satisfy \(Ax_{n}-\lambda x_{n}\longrightarrow0\) in norm, and once that is known the sequential form of compactness, Lemma A.230, is enough to extract a genuine eigenvector — the vector \(e\) is recovered from \(Ax_{n_{k}}\) by dividing by \(\lambda\neq0\), not by any weak limit. The argument therefore quotes nothing beyond Proposition 12.43 and Theorem 6.31, and Hilbert Spaces may drop the reservation with which Theorem 12.44 was first stated.
The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators discharges the proof obligation of Theorem 12.44 (Section 12.3.1). The theorem is what makes the integral operators Equation (12.28) of Partial Differential Equations — the Green functions of a boundary-value problem — carry a discrete spectrum of modes with a complete set of eigenfunctions, and it is the exact infinite-dimensional counterpart of Equation (5.105). It is also the last point at which the counterpart is exact: Example 12.56 exhibits a bounded self-adjoint operator, not compact, with no eigenvector at all, and the repair for that case is the projection-valued measure of The Spectral Theorem for a Bounded Self-Adjoint Operator.
The Spectral Theorem for a Bounded Self-Adjoint Operator
This appendix proves Theorem 12.59 of Hilbert Spaces, in both of the forms stated there: the existence and uniqueness of a projection-valued measure \(E\) on \(\R\), supported on \(\sigma(A)\), with \(A=\int_{\sigma(A)}\lambda\,\dd E(\lambda)\); and the unitary equivalence of \(A\), on a separable Hilbert space, with multiplication by a bounded real function on a finite measure space. It is the theorem three of the quantum postulates of The Postulates of Quantum Mechanics are statements about (Remark 12.63), and it is the replacement, forced by Example 12.56, for the eigenbasis that Theorem 5.75 supplies in finite dimension.
The route is the classical one of von Neumann [vonNeumann:1930] [vonNeumann:1932] in the arrangement of Reed and Simon [Reed:1972]: first the polynomial calculus and the identity \(\norm{p(A)}=\sup_{\sigma(A)}\abs{p}\), which is what makes the construction possible at all; then the continuous calculus, obtained by completing the polynomials in the supremum norm; then, for each pair of vectors, a complex measure on \(\sigma(A)\); then the extension of the calculus to bounded Borel functions, whose values on indicator functions are the projections \(E(\Omega)\); and finally the multiplication form. Two classical theorems are used and are not proved in this treatise — Stone–Weierstrass and Riesz–Markov. Each is stated precisely where it is used, and Remark A.251 collects them.
Throughout, \(A\in\mathcal{B}(\mathcal{H})\) is self-adjoint, \(\mathcal{H}\neq\set{0}\), and \(\sigma(A)\subseteq\R\) is compact and non-empty (Theorems 12.54 and 12.55). We write \(C(\sigma(A))\) for the continuous complex functions on \(\sigma(A)\) with the supremum norm \(\norm{f}_{\infty}=\sup_{\lambda\in\sigma(A)}\abs{f(\lambda)}\), and \(B(\sigma(A))\) for the bounded Borel functions with the same norm.
Statement
Let \(A\in\mathcal{B}(\mathcal{H})\) be self-adjoint. Then:
-
there is exactly one projection-valued measure \(E\) on \(\R\) (Definition 12.58) with \(E(\R\setminus\sigma(A))=0\) and
\begin{equation}\tag{A.446} \braket{x}{Ay}=\int_{\sigma(A)}\lambda\, \braket{x}{\dd E(\lambda)y}\ec\qquad \forall\,x,y\in\mathcal{H}\ec \end{equation}and every \(B\in\mathcal{B}(\mathcal{H})\) commuting with \(A\) commutes with every \(E(\Omega)\);
-
if \(\mathcal{H}\) is separable there are a measure space \((M,\mu)\) with \(\mu(M)<\infty\), a bounded real measurable function \(a\) on \(M\), and a unitary \(U\!:\mathcal{H}\longrightarrow L^{2}(M,\mu)\) with
\begin{equation}\tag{A.447} \bigl(UAU^{-1}f\bigr)(m)=a(m)\,f(m)\ec\qquad \forall\,f\in L^{2}(M,\mu)\ep \end{equation}
Rests on Definition 12.58, Theorem 12.55 and Proposition 12.43.
The polynomial calculus
For a polynomial \(p(z)=\sum_{k=0}^{d}c_{k}z^{k}\) with complex coefficients write \(p(A)=\sum_{k}c_{k}A^{k}\in\mathcal{B}(\mathcal{H})\), with \(A^{0}=\identity\), and let \(\bar p(z)=\sum_{k}c_{k}^{\ast}z^{k}\). Since \(A^{\dagger}=A\), Equation (12.24) gives \(p(A)^{\dagger}=\bar p(A)\).
Let \(B\in\mathcal{B}(\mathcal{H})\) be self-adjoint. Then \(\norm{B}\in\sigma(B)\) or \(-\norm{B}\in\sigma(B)\); consequently
Rests on Proposition 12.43, Proposition 12.52 and Definition 12.49.
Derives Lemma A.239. If \(B=0\) then \(\sigma(B)=\set{0}\) by Theorem 12.54 and Proposition 12.52 and the claim is trivial, so let \(B\neq0\). Exactly as in Equation (A.436), choose unit vectors \(x_{n}\) with \(\braket{x_{n}}{Bx_{n}}\longrightarrow\lambda\) where \(\lambda\in\R\) and \(\abs{\lambda}=\norm{B}\); this is possible by Proposition 12.43 and because \(\braket{x_{n}}{Bx_{n}}\) is real (Proposition 12.42). Then, as in Equation (A.437),
If \(B-\lambda\identity\) had a bounded inverse \(R\) we would get \(1=\norm{x_{n}}\leq\norm{R}\norm{(B-\lambda\identity)x_{n}} \longrightarrow0\), which is false; so \(\lambda\in\sigma(B)\) with \(\abs{\lambda}=\norm{B}\). Since \(\sigma(B)\) is contained in the disc of radius \(\norm{B}\) (Proposition 12.52), Equation (A.448) follows.
∎For every polynomial \(p\) and every \(A\in\mathcal{B}(\mathcal{H})\),
Rests on Definition 12.49 and Proposition 12.37.
Derives Lemma A.240. If \(p\) is constant, \(p\equiv c\), then \(p(A)=c\identity\) and both sides are \(\set{c}\) (using \(\sigma(A)\neq\varnothing\), Theorem 12.54). Otherwise let \(\mu\in\C\) and factor the non-constant polynomial \(p(z)-\mu\) over \(\C\) (Theorem 8.19):
so that, substituting \(A\) and using that polynomials in \(A\) commute,
If every factor is invertible, so is the product, and \(\mu\in\rho(p(A))\). Conversely suppose the product \(T=\prod_{j}(A-\alpha_{j}\identity)\) is invertible and fix \(j\). Its inverse \(T^{-1}\) commutes with \(A\) — from \(TA=AT\) one gets \(AT^{-1}=T^{-1}A\) on multiplying by \(T^{-1}\) on both sides — hence with every factor. Writing \(S_{j}=\prod_{k\neq j}(A-\alpha_{k} \identity)\) we have \(T=(A-\alpha_{j}\identity)S_{j}=S_{j} (A-\alpha_{j}\identity)\), so
and \(A-\alpha_{j}\identity\) is invertible with inverse \(S_{j}T^{-1}=T^{-1}S_{j}\). Therefore \(\mu\in\sigma(p(A))\) if and only if some \(\alpha_{j}\in\sigma(A)\); and by Equation (A.451) the \(\alpha_{j}\) are exactly the roots of \(p-\mu\), so this happens exactly when \(\mu=p(\lambda)\) for some \(\lambda\in\sigma(A)\).
∎For every polynomial \(p\) and every self-adjoint \(A\in\mathcal{B}(\mathcal{H})\),
Rests on Lemma A.239, Lemma A.240 and Proposition 12.39.
Derives Proposition A.241. The operator \(p(A)^{\dagger}p(A)=(\bar p\,p)(A)\) is self-adjoint, and \(\bar p\,p\) is a polynomial. By the \(C^{\ast}\) identity Equation (12.25), then Lemma A.239 applied to it, then Lemma A.240,
the last step because \(\sigma(A)\subseteq\R\) (Theorem 12.55), so that \(\bar p(\lambda)=p(\lambda)^{\ast}\) there.
∎The continuous calculus
Let \(K\) be a compact Hausdorff space and let \(\mathcal{A}\subseteq C(K)\) be a subalgebra that contains the constants, separates the points of \(K\), and is closed under complex conjugation. Then \(\mathcal{A}\) is dense in \(C(K)\) for the supremum norm. For \(K\subseteq\R\) compact the polynomials form such a subalgebra, and the statement reduces to the classical Weierstrass approximation theorem quoted in Remark 12.1. Rests on Definition 6.9.
There is exactly one map \(\Phi\!:C(\sigma(A))\longrightarrow\mathcal{B}(\mathcal{H})\) that is linear, multiplicative, sends \(1\longmapsto\identity\) and \(\lambda\longmapsto A\), and is continuous for \(\norm{\cdot}_{\infty}\) and the operator norm. It satisfies, writing \(f(A)=\Phi(f)\),
and \(f(A)\) commutes with every bounded operator commuting with \(A\). Rests on Proposition A.241, Theorem A.242 and Proposition 12.61.
Derives Proposition A.243. Well defined on polynomials. If two polynomials agree on \(\sigma(A)\) their difference \(q\) has \(\norm{q}_{\infty}=0\), so \(q(A)=0\) by Equation (A.453): the map \(p|_{\sigma(A)}\longmapsto p(A)\) depends only on the restriction, and it is isometric.
Extension. The polynomials restricted to the compact set \(\sigma(A)\subseteq\R\) are dense in \(C(\sigma(A))\) by Theorem A.242. An isometric linear map defined on a dense subspace of a normed space, with values in the complete space \(\mathcal{B}(\mathcal{H})\) (Proposition 12.37), extends uniquely to an isometry of the closure: given \(f\), choose polynomials \(p_{n}\) with \(\norm{p_{n}-f}_{\infty}\longrightarrow0\); then \(\norm{p_{n}(A)-p_{m}(A)}=\norm{p_{n}-p_{m}}_{\infty}\) makes \((p_{n}(A))\) Cauchy, and its limit \(\Phi(f)\) is independent of the approximating sequence by the same identity. Linearity and multiplicativity pass to the limit because addition and multiplication are continuous in a Banach algebra (Proposition 12.37), and the first identity of Equation (A.455) is the limit of Equation (A.453). Taking adjoints term by term — legitimate because the adjoint is isometric (Proposition 12.39) — gives \(\Phi(f)^{\dagger}=\Phi(\bar f)\).
Positivity. If \(f\geq0\) on \(\sigma(A)\) then \(g=\sqrt{f}\) is continuous and real, so \(\Phi(f)=\Phi(g)^{2}=\Phi(g)^{\dagger}\Phi(g)\) is positive by Proposition 12.42(iv).
Commutation. If \(BA=AB\) then \(B\) commutes with every polynomial in \(A\), hence with the norm limits \(\Phi(f)\).
Uniqueness is Proposition 12.61, already proved in the chapter.
∎The spectral measures
Let \(K\) be a compact metric space. For every bounded linear functional \(\ell\) on \(C(K)\) there is exactly one regular complex Borel measure \(\nu\) on \(K\) with
and \(\norm{\ell}\) equals the total variation \(\abs{\nu}(K)\). If \(\ell(f)\geq0\) whenever \(f\geq0\), then \(\nu\) is a positive measure and \(\norm{\ell}=\nu(K)=\ell(1)\). Regularity carries with it that \(C(K)\) is dense in \(L^{2}(K,\nu)\) for every positive finite Borel measure \(\nu\). Rests on Definition 6.9.
For \(x,y\in\mathcal{H}\) there is a unique regular complex Borel measure \(\mu_{x,y}\) on \(\sigma(A)\) with
The assignment is linear in \(y\), antilinear in \(x\), obeys \(\mu_{y,x}=\mu_{x,y}^{\ast}\) and \(\abs{\mu_{x,y}}(\sigma(A))\leq\norm{x}\norm{y}\); and \(\mu_{x} :=\mu_{x,x}\) is a positive measure with \(\mu_{x}(\sigma(A))=\norm{x}^{2}\). Rests on Proposition A.243, Theorem A.244 and Theorem 12.46.
Derives Proposition A.245. The map \(f\longmapsto\braket{x}{f(A)y}\) is linear on \(C(\sigma(A))\) and bounded, since by Equation (12.1) and Equation (A.455)
Theorem A.244 supplies a unique regular complex Borel measure \(\mu_{x,y}\) satisfying Equation (A.457), with \(\abs{\mu_{x,y}}(\sigma(A))\leq\norm{x}\norm{y}\) by the norm statement there. Linearity in \(y\) and antilinearity in \(x\) hold for the functionals and therefore, by the uniqueness in Theorem A.244, for the measures. For the conjugation rule,
and since \(f\longmapsto\bar f\) exhausts \(C(\sigma(A))\), uniqueness gives \(\mu_{y,x}=\mu_{x,y}^{\ast}\). Finally, if \(f\geq0\) then \(\braket{x}{f(A)x}\geq0\) by the positivity in Equation (A.455), so \(\mu_{x}\) is a positive measure, and taking \(f=1\) in Equation (A.457) gives \(\mu_{x}(\sigma(A))=\braket{x}{x}=\norm{x}^{2}\).
∎The Borel calculus
There is a unique map \(g\longmapsto g(A)\) from \(B(\sigma(A))\) to \(\mathcal{B}(\mathcal{H})\) extending \(\Phi\) and satisfying
It is linear, multiplicative, satisfies \(g(A)^{\dagger}=\bar g(A)\) and \(\norm{g(A)}\leq\norm{g}_{\infty}\), and \(g(A)\) commutes with every bounded operator commuting with \(A\). Rests on Proposition A.245, Theorem 12.46 and Proposition A.243.
Derives Proposition A.246. Existence of the operator. Fix \(g\in B(\sigma(A))\) and consider
which by Proposition A.245 is antilinear in \(x\) and linear in \(y\). For fixed \(x\) the map \(y\longmapsto b_{g}(x,y)\) is a bounded linear functional, so by Theorem 12.46 there is a unique vector \(w(x)\) with \(b_{g}(x,y)=\braket{w(x)}{y}\); the map \(x\longmapsto w(x)\) is linear — antilinearity in the first slot on both sides — and bounded by \(\norm{g}_{\infty}\). Its adjoint in the sense of Theorem 12.38 is the operator we call \(g(A)\), so that Equation (A.459) holds with \(\norm{g(A)}\leq\norm{g}_{\infty}\). For continuous \(g\) the formula reproduces Equation (A.457), so the map extends \(\Phi\), and it is unique because Equation (A.459) determines all matrix elements. Linearity in \(g\) and \(g(A)^{\dagger}=\bar g(A)\) follow from Equation (A.459) and \(\mu_{y,x}=\mu_{x,y}^{\ast}\).
Multiplicativity, first step. Let \(f\in C(\sigma(A))\) and \(x,y\in\mathcal{H}\). For every \(h\in C(\sigma(A))\),
using multiplicativity of the continuous calculus. The two regular measures \(\mu_{x,f(A)y}\) and \(f\,\mu_{x,y}\) therefore integrate every continuous function alike, so they are equal by the uniqueness in Theorem A.244. Consequently, for \(g\in B(\sigma(A))\),
i.e. \(g(A)f(A)=(gf)(A)\) for \(g\) Borel and \(f\) continuous. Taking adjoints in Equation (A.460) and replacing \((f,g)\) by \((\bar f,\bar g)\) gives \(f(A)g(A)=(fg)(A)\) as well.
Multiplicativity, second step. Now let \(g\in B(\sigma(A))\) be fixed. For every \(h\in C(\sigma(A))\), by the step just proved,
so \(\mu_{x,g(A)y}=g\,\mu_{x,y}\) by uniqueness again. Hence for any two bounded Borel \(g,g'\),
Commutation. Let \(BA=AB\). By Proposition A.243, \(B\) commutes with \(f(A)\) for every continuous \(f\), so for such \(f\)
whence \(\mu_{B^{\dagger}x,y}=\mu_{x,By}\), and evaluating Equation (A.459) on both sides gives \(\braket{x}{Bg(A)y}=\braket{x}{g(A)By}\) for every Borel \(g\).
∎The projection-valued measure
Proof of Theorem A.238, part (1): existence. Derives Theorem A.238. For a Borel set \(\Omega\subseteq\R\) let \(\chi_{\Omega}\) be its indicator function, restricted to \(\sigma(A)\), and put
This is a bounded operator by Proposition A.246.
Projections. \(\chi_{\Omega}^{2}=\chi_{\Omega}\) and \(\bar\chi_{\Omega}=\chi_{\Omega}\), so \(E(\Omega)^{2}=E(\Omega)=E(\Omega)^{\dagger}\): each \(E(\Omega)\) is a projection in the sense of Proposition 12.21.
Normalisation and multiplicativity. \(\chi_{\varnothing}=0\) and \(\chi_{\R}=1\) on \(\sigma(A)\) give \(E(\varnothing)=0\) and \(E(\R)=\identity\); also \(E(\R\setminus\sigma(A))=0\), since that indicator vanishes identically on \(\sigma(A)\), which is the support statement. And \(\chi_{\Omega_{1}\cap\Omega_{2}}=\chi_{\Omega_{1}}\chi_{\Omega_{2}}\) gives \(E(\Omega_{1}\cap\Omega_{2})=E(\Omega_{1})E(\Omega_{2})\) by Equation (A.461).
Countable additivity. Finite additivity is linearity of the calculus: for disjoint \(\Omega_{1},\Omega_{2}\), \(\chi_{\Omega_{1}\cup\Omega_{2}}=\chi_{\Omega_{1}}+\chi_{\Omega_{2}}\). For a disjoint sequence \(\set{\Omega_{n}}\) with union \(\Omega\), write \(S_{N}=\sum_{n\leq N}E(\Omega_{n})=E(\bigcup_{n\leq N}\Omega_{n})\), so that \(E(\Omega)-S_{N}=E(\Omega_{>N})\) with \(\Omega_{>N}=\bigcup_{n>N}\Omega_{n}\). Since \(E(\Omega_{>N})\) is a projection,
using Equation (A.459) with \(g=\chi_{\Omega_{>N}}\) and the countable additivity of the ordinary positive measure \(\mu_{x}\). The series \(\sum_{n}\mu_{x}(\Omega_{n})=\mu_{x}(\Omega)\leq \norm{x}^{2}\) converges, so its tail tends to \(0\), and Equation (12.36) holds. Thus \(E\) is a projection-valued measure in the sense of Definition 12.58.
It represents \(A\). Taking \(g=\chi_{\Omega}\) in Equation (A.459) identifies the complex measure of Definition 12.58 with the measure of Proposition A.245: \(\braket{x}{E(\Omega)y}=\mu_{x,y}(\Omega)\). Hence, taking \(f(\lambda)=\lambda\) in Equation (A.457) — the continuous calculus assigns \(A\) to that function — we obtain Equation (A.446). The commutation statement is the last part of Proposition A.246.
∎Let \(F\) be any projection-valued measure on \(\R\) supported on a compact set \(K\), and for \(g\in B(K)\) define \(\int g\,\dd F\) by \(\braket{x}{(\int g\,\dd F)y}=\int g\,\dd\braket{x}{F y}\). Then \(g\longmapsto\int g\,\dd F\) is linear and multiplicative. Rests on Definition 12.58 and Theorem 12.46.
Derives Lemma A.247. Exactly as in Proposition A.246, the sesquilinear form is bounded by \(\norm{g}_{\infty}\norm{x}\norm{y}\) — the total variation of \(\braket{x}{F(\cdot)y}\) is at most \(\norm{x}\norm{y}\), by the Cauchy–Schwarz inequality applied to \(\braket{x}{F(\Omega)y}=\braket{F(\Omega)x}{F(\Omega)y}\) over a partition — so the operator exists, and linearity is clear. For multiplicativity, note first that for simple functions \(g=\sum_{i}c_{i}\chi_{\Omega_{i}}\) and \(g'=\sum_{j}c'_{j}\chi_{\Omega'_{j}}\) the property \(F(\Omega)F(\Omega')=F(\Omega\cap\Omega')\) gives
A bounded Borel function is a uniform limit of simple Borel functions — partition the disc of radius \(\norm{g}_{\infty}\) into finitely many Borel pieces of diameter \(<\varepsilon\) and take preimages — and \(\norm{\int g\,\dd F}\leq\norm{g}_{\infty}\), so both sides pass to the uniform limit.
∎Proof of Theorem A.238, part (1): uniqueness. Derives Theorem A.238. Let \(E\) and \(E'\) both satisfy Equation (A.446) and be supported on \(\sigma(A)\), and write \(\nu_{x,y},\nu'_{x,y}\) for the associated complex measures. By Lemma A.247, \(\int\lambda^{n}\,\dd E=(\int\lambda\,\dd E)^{n}=A^{n}\) for every \(n\geq0\), and likewise for \(E'\); hence
for every polynomial \(p\). The polynomials are dense in \(C(\sigma(A))\) (Theorem A.242) and both measures are finite, so the two integrate every continuous function alike; by the uniqueness in Theorem A.244, \(\nu_{x,y}=\nu'_{x,y}\). Taking \(\Omega\) Borel and \(x,y\) arbitrary gives \(\braket{x}{E(\Omega)y}=\braket{x}{E'(\Omega)y}\), i.e. \(E=E'\).
∎The multiplication form
For \(x\in\mathcal{H}\) the cyclic subspace generated by \(x\) is
a closed \(A\)-invariant subspace containing \(x=1(A)x\). The vector \(x\) is cyclic for \(A\) if \(\mathcal{H}_{x}=\mathcal{H}\). Rests on Proposition A.243 and Theorem 12.18.
Let \(x\neq0\) be cyclic for \(A\). Then there is a unitary \(U\!:\mathcal{H}\longrightarrow L^{2}(\sigma(A),\mu_{x})\) with \(Uf(A)x=f\) for every \(f\in C(\sigma(A))\), and
Rests on Definition A.248, Proposition A.245 and Theorem A.244.
Derives Lemma A.249. For \(f\in C(\sigma(A))\), using Equation (A.455) and Equation (A.457),
so \(f(A)x\longmapsto f\) is a well-defined linear isometry from the dense subspace \(\set{f(A)x}\) of \(\mathcal{H}\) onto the subspace \(C(\sigma(A))\) of \(L^{2}(\sigma(A),\mu_{x})\) — well defined because \(f(A)x=g(A)x\) forces \(\norm{f-g}_{L^{2}(\mu_{x})}=0\) by Equation (A.466). Its target is dense by the regularity statement of Theorem A.244, and both spaces are complete, so the isometry extends to a unitary \(U\) (a densely defined isometry with dense range extends to a surjective isometry by taking limits, and Equation (12.3) then gives preservation of the inner product). Finally, for \(f\) continuous, \(UA f(A)x=U(\lambda f)(A)x=\lambda f=\lambda\cdot(Uf(A)x)\), which is Equation (A.465) on a dense subspace; both sides are bounded, so it holds everywhere.
∎Let \(\mathcal{H}\) be separable and \(A\) self-adjoint. Then there are finitely or countably many non-zero vectors \(x_{n}\) such that the cyclic subspaces \(\mathcal{H}_{x_{n}}\) are mutually orthogonal and
Rests on Definition A.248, Proposition 12.87 and Definition 12.32.
Derives Lemma A.250. Let \(\set{y_{j}}_{j\in\N}\) be dense in \(\mathcal{H}\) (Definition 12.32). Construct the \(x_{n}\) recursively. Put \(K_{0}=\mathcal{H}\). Given mutually orthogonal cyclic subspaces \(\mathcal{H}_{x_{1}},\dots,\mathcal{H}_{x_{n-1}}\), let \(K_{n-1}\) be the orthogonal complement of their sum. Each \(\mathcal{H}_{x_{i}}\) is \(A\)-invariant, and \(A\) is self-adjoint, so \(K_{n-1}\) is \(A\)-invariant too (Lemma A.231), and it is closed. If \(K_{n-1}=\set{0}\), stop. Otherwise let \(j(n)\) be the least index with \(P_{K_{n-1}}y_{j}\neq0\) — such an index exists, since otherwise every \(y_{j}\) would lie in \(K_{n-1}^{\perp}\) and, by density and closedness, so would every vector, forcing \(K_{n-1}=\set{0}\) — and set \(x_{n}=P_{K_{n-1}}y_{j(n)}\), \(\mathcal{H}_{x_{n}}\subseteq K_{n-1}\).
By construction \(j(1)<j(2)<\cdots\), because \(P_{K_{n-1}}y_{j}=0\) for every \(j<j(n)\) not already used means \(y_{j}\in \mathcal{H}_{x_{1}}\oplus\cdots\oplus\mathcal{H}_{x_{n-1}}\), while \(y_{j(n)}=(\text{its part in that sum})+x_{n}\) also lies in \(\mathcal{H}_{x_{1}}\oplus\cdots\oplus\mathcal{H}_{x_{n}}\). Hence every \(y_{j}\) lies in \(\bigoplus_{n}\mathcal{H}_{x_{n}}\), whose closure is therefore all of \(\mathcal{H}\); that closure is the internal direct sum of Proposition 12.87, which is Equation (A.467). If instead the recursion stops because \(K_{n-1}=\set{0}\), the sum is already \(\mathcal{H}\).
∎Proof of Theorem A.238, part (2). Derives Theorem A.238. Take the decomposition Equation (A.467) and write \(A_{n}\) for the restriction of \(A\) to \(\mathcal{H}_{x_{n}}\), a bounded self-adjoint operator on that Hilbert space for which \(x_{n}\) is cyclic, with \(\sigma(A_{n})\subseteq\sigma(A)\) and with the same functional calculus (the calculus of \(A\) leaves \(\mathcal{H}_{x_{n}}\) invariant). Lemma A.249 gives unitaries \(U_{n}\!:\mathcal{H}_{x_{n}}\longrightarrow L^{2}(\sigma(A),\mu_{n})\), \(\mu_{n}=\mu_{x_{n}}\), carrying \(A\) into multiplication by \(\lambda\).
Let \(M\) be the disjoint union of countably many copies of \(\sigma(A)\), the \(n\)th copy written \(\sigma(A)\times\set{n}\), and define on it the measure
so that, by Proposition A.245, \(\mu(M)=\sum_{n}c_{n}\norm{x_{n}}^{2}=\sum_{n}2^{-n}\leq1<\infty\). Define \(a(\lambda,n)=\lambda\), a bounded real measurable function on \(M\) because \(\sigma(A)\) is compact and real. The map
is unitary from \(\mathcal{H}\) onto \(L^{2}(M,\mu)\): on the \(n\)th block it is an isometry from \(\mathcal{H}_{x_{n}}\) onto \(L^{2}(\sigma(A),c_{n}\mu_{n})\), because \(\int\abs{c_{n}^{-1/2}h}^{2}\,c_{n}\dd\mu_{n} =\int\abs{h}^{2}\dd\mu_{n}\), and \(L^{2}(M,\mu)\) is by construction the orthogonal direct sum of those blocks (Proposition 12.85). Since multiplication by \(\lambda\) is unaffected by the constant \(c_{n}^{-1/2}\), Equation (A.465) gives Equation (A.447), which completes the proof of Theorem A.238 and, with it, of Theorem 12.59.
∎Two classical theorems enter the proof and are not derived in this treatise; every other step above is carried out from Hilbert Spaces and its predecessors.
-
Theorem A.242 (Stone–Weierstrass) is used twice: to extend the polynomial calculus to \(C(\sigma(A))\), and in the uniqueness argument. Only the special case of a compact subset of \(\R\), where it is the Weierstrass approximation theorem, is ever needed; that theorem is already declared as quoted in Remark 12.1 and is used there for Proposition 12.61.
-
Theorem A.244 (Riesz–Markov) is the representation of a bounded functional on \(C(K)\) by a regular complex Borel measure, together with the uniqueness of that measure and the density of \(C(K)\) in \(L^{2}(K,\nu)\). It is a theorem of measure theory, and this treatise develops no measure theory (Remark 12.1); it is what converts the operator inequality Equation (A.458) into a measure, and there is no route to a projection-valued measure that avoids it. Statement and proof: [Reed:1972], chapter IV.
Nothing else is assumed. In particular the polynomial identity Equation (A.453) — the step that makes the whole construction possible, and the one a reader should check first — rests only on the \(C^{\ast}\) identity Equation (12.25) and the numerical-radius formula Equation (12.26), both proved in the chapter.
The Spectral Theorem for a Bounded Self-Adjoint Operator discharges the proof obligation of Theorem 12.59 (Section 12.4.2). The projection-valued measure it constructs is what Definition 12.60 integrates against, so the functional calculus \(f(A)\) of the chapter — unique by Proposition 12.61 — now exists as well as being unique; Proposition A.246 is that existence statement for bounded Borel \(f\). The theorem is used again in Stone's Theorem on One-Parameter Unitary Groups, where the unbounded case is derived from it and the exponential \(\exp(-\ii Ht/\hbar)\) is defined by Equation (A.459), and in The Nuclear Spectral Theorem of Gelfand and Maurin, whose direct-integral decomposition is Lemma A.250 read fibrewise. Physically it is Remark 12.63: the support of \(E\) is the set of possible measured values, and \(\norm{E(\Omega)\psi}^{2}\) is the Born probability.
Stone's Theorem on One-Parameter Unitary Groups
This appendix proves the two halves of Theorem 12.66 (Hilbert Spaces) that Proposition 12.67 left open: that the domain
is dense and that \(H\) is self-adjoint on it, not merely symmetric; and the converse, that every self-adjoint operator — bounded or not — exponentiates to a strongly continuous one-parameter unitary group. Together with Proposition 12.67 this establishes the bijection asserted by Theorem 12.66, which is what makes the two halves of the abstract Schrödinger equation Equation (12.48) equivalent statements [Stone:1932].
Units are carried explicitly, as everywhere in this treatise. The group parameter \(t\) has the unit \(\mathrm{s}\), the generator \(H\) the unit \(\mathrm{J}\), and \(\hbar\) the unit \(\mathrm{J}\,\mathrm{s}\), so that \(Ht/\hbar\) is a pure number; the energy scale \(\varepsilon>0\) introduced in The converse: exponentiating a self-adjoint operator likewise has the unit \(\mathrm{J}\), and the spectral variable \(\lambda\) is an energy. The smoothing function \(\varphi\) of The domain is dense has the unit \(/\mathrm{s}\), so that \(\int\varphi(t)\,\dd t\) is dimensionless.
Statement
Let \(\set{U(t)}_{t\in\R}\) be a strongly continuous one-parameter unitary group on a Hilbert space \(\mathcal{H}\) (Definition 12.64). Then the domain Equation (A.470) is dense and \(H\) is self-adjoint on it. Conversely, if \(H\) is any self-adjoint operator on \(\mathcal{H}\) (Definition 12.72), then \(U(t)=\exp(-\ii Ht/\hbar)\), defined by the functional calculus of Proposition A.246 applied to the projection-valued measure of Proposition A.262, is a strongly continuous one-parameter unitary group whose generator in the sense of Equation (A.470) is exactly \(H\). The two constructions are mutually inverse. Rests on Definition 12.64, Proposition 12.67 and Theorem A.238.
Vector-valued integrals
The smoothing argument needs to integrate a continuous \(\mathcal{H}\)-valued function. Nothing more than the Riemann integral is required, and it is built exactly as on the line.
Let \(G\!:[a,b]\longrightarrow\mathcal{H}\) be continuous. Then the Riemann sums of \(G\) over tagged partitions converge, as the mesh tends to \(0\), to a vector \(\int_{a}^{b}G(t)\,\dd t\in\mathcal{H}\), and
for every \(w\in\mathcal{H}\). If \(B\in\mathcal{B}(\mathcal{H})\) then \(B\int_{a}^{b}G=\int_{a}^{b}BG\). Rests on Definition 12.2 and Proposition 12.4.
Derives Lemma A.254. \(G\) is uniformly continuous on the compact interval \([a,b]\): the Heine–Cantor argument of Theorem 7.25 uses only that the target is a metric space, and \(\mathcal{H}\) is one. So given \(\eta>0\) there is \(\delta>0\) with \(\norm{G(t)-G(s)}<\eta\) whenever \(\abs{t-s}<\delta\). Two tagged partitions of mesh \(<\delta\) have a common refinement, and comparing each with the refinement bounds the difference of their Riemann sums by \(2\eta(b-a)\); the sums therefore form a Cauchy net, which converges because \(\mathcal{H}\) is complete (Definition 12.2). Both identities in Equation (A.471) hold for Riemann sums — the first by linearity of the inner product in its second slot, the second by the triangle inequality — and pass to the limit, the first by continuity of the inner product (Proposition 12.4). The same argument gives \(B\int G=\int BG\) for a bounded \(B\).
∎The domain is dense
Let \(U\) be a strongly continuous one-parameter unitary group, let \(\varphi\!:\R\longrightarrow\R\) be continuously differentiable with support in a compact interval, and for \(x\in\mathcal{H}\) put
Then \(x_{\varphi}\in D(H)\) and
Rests on Lemma A.254, Definition 12.64 and Equation (A.470).
Derives Lemma A.255. The integrand is continuous, by strong continuity of \(U\) and continuity of \(\varphi\), and vanishes outside a compact interval, so Lemma A.254 defines \(x_{\varphi}\). Using \(U(s)U(t)=U(t+s)\) and the substitution \(t\longmapsto t-s\) — a translation, which the Riemann integral of Lemma A.254 respects because it acts on tagged partitions by translation —
Hence, for \(s\neq0\),
By the mean value theorem, \([\varphi(t-s)-\varphi(t)]/s =-\varphi'(t-\theta s)\) for some \(\theta=\theta(t,s)\in(0,1)\), and \(\varphi'\) is continuous with compact support, hence uniformly continuous; so \([\varphi(\cdot-s)-\varphi]/s\longrightarrow-\varphi'\) uniformly as \(s\to0\), with all supports inside one fixed compact interval \([a,b]\) of length \(\ell\) once \(\abs{s}\leq1\). The second estimate of Equation (A.471) and \(\norm{U(t)x}=\norm{x}\) then give
So the limit in Equation (A.470) exists and equals \(-x_{\varphi'}\); multiplying by \(\ii\hbar\) gives Equation (A.473).
∎\(D(H)\) is dense in \(\mathcal{H}\). Rests on Lemma A.255 and Definition 12.64.
Derives Proposition A.256. For \(n\in\N\) define
Each \(\varphi_{n}\) is continuously differentiable — at \(t=\pm1/n\) both \(\varphi_{n}\) and \(\varphi_{n}'\) vanish, so the pieces match — non-negative, supported in \([-1/n,1/n]\), and normalised: substituting \(u=nt\),
a pure number, as it must be: \(\varphi_{n}\) carries the unit \(/\mathrm{s}\) and \(\dd t\) the unit \(\mathrm{s}\).
Fix \(x\in\mathcal{H}\). Since \(\int\varphi_{n}=1\),
by Equation (A.471) and non-negativity of \(\varphi_{n}\). The right-hand side tends to \(0\) as \(n\to\infty\), by strong continuity of \(U\) at \(t=0\) (Definition 12.64). Each \(x_{\varphi_{n}}\) lies in \(D(H)\) by Lemma A.255, so \(x\) is a limit of vectors of \(D(H)\).
∎The generator is closed and self-adjoint
For \(x\in D(H)\) and every \(t\),
Rests on Proposition 12.67 and Lemma A.254.
Derives Lemma A.257. Fix \(w\in\mathcal{H}\) and let \(g(r)=\braket{w}{U(r)x}\). By Equation (12.46) the curve \(r\longmapsto U(r)x\) is differentiable in norm with derivative \(-\ii\hbar^{-1}U(r)Hx\), so \(g\) is differentiable with \(g'(r)=-\ii\hbar^{-1}\braket{w}{U(r)Hx}\), which is continuous in \(r\). The fundamental theorem of calculus (Theorem 7.43) and the first identity of Equation (A.471) give
As \(w\) is arbitrary, Equation (5.31) gives Equation (A.477).
∎\(H\) is a closed operator (Definition 12.70). Rests on Lemma A.257 and Equation (A.470).
Derives Proposition A.258. Let \(x_{n}\in D(H)\) with \(x_{n}\longrightarrow x\) and \(Hx_{n}\longrightarrow z\). Each term of Equation (A.477) converges: the left side to \(U(t)x-x\), since \(U(t)\) is bounded; and the right side to \(-\ii\hbar^{-1}\int_{0}^{t}U(r)z\,\dd r\), because
Hence \(U(t)x-x=-\ii\hbar^{-1}\int_{0}^{t}U(r)z\,\dd r\) for every \(t\). Dividing by \(t\neq0\) and using continuity of \(r\longmapsto U(r)z\) at \(r=0\),
so the limit Equation (A.470) exists for \(x\) and equals \(-\ii\hbar^{-1}z\) before multiplication by \(\ii\hbar\). That is, \(x\in D(H)\) and \(Hx=z\): the graph is closed.
∎\(H\) is self-adjoint on \(D(H)\). Rests on Propositions 12.67, A.256 and A.258.
Derives Proposition A.259. \(H\) is densely defined (Proposition A.256) and symmetric (Equation (12.45)), so \(H\subseteq H^{\dagger}\) and the adjoint exists. Let \(\varepsilon>0\) be an energy.
Step 1: \(\ker(H^{\dagger}\mp\ii\varepsilon)=\set{0}\). Let \(y\in D(H^{\dagger})\) with \(H^{\dagger}y=\ii\varepsilon y\), fix \(x\in D(H)\), and put \(f(t)=\braket{y}{U(t)x}\). By Proposition 12.67, \(U(t)x\in D(H)\) and the curve is differentiable, so
where the second equality is Equation (12.50) and the third uses the antilinearity Equation (5.36), which turns \(\ii\varepsilon\) in the first slot into \(-\ii\varepsilon\) and cancels one factor of \(\ii\) against the prefactor. Hence \(f(t)=f(0)\,\ee^{-\varepsilon t/\hbar}\). But \(\abs{f(t)}\leq\norm{y}\norm{U(t)x}=\norm{y}\norm{x}\) for all \(t\in\R\), while \(\ee^{-\varepsilon t/\hbar}\longrightarrow\infty\) as \(t\longrightarrow-\infty\); so \(f(0)=\braket{y}{x}=0\). As \(x\) ranges over the dense set \(D(H)\), \(y=0\). The case \(H^{\dagger}y=-\ii\varepsilon y\) is the same with \(t\to+\infty\).
Step 2: \(\im(H\pm\ii\varepsilon)=\mathcal{H}\). For \(x\in D(H)\), symmetry makes the cross terms cancel:
If \((H+\ii\varepsilon)x_{n}\longrightarrow u\) then, by Equation (A.479) applied to differences, both \((x_{n})\) and \((Hx_{n})\) are Cauchy; \(H\) is closed (Proposition A.258), so \(x_{n}\longrightarrow x\in D(H)\) with \((H+\ii\varepsilon)x=u\). The image is therefore closed. It is also dense: exactly as in Equation (12.53), \(y\perp\im(H+\ii\varepsilon)\) means \(\braket{y}{Hx}=\braket{\ii\varepsilon y}{x}\) for all \(x\in D(H)\), i.e. \(y\in\ker(H^{\dagger}-\ii\varepsilon)=\set{0}\) by Step 1. A dense closed subspace is everything, and likewise for \(H-\ii\varepsilon\).
Step 3. Let \(y\in D(H^{\dagger})\). By Step 2 there is \(x\in D(H)\) with \((H-\ii\varepsilon)x=(H^{\dagger}-\ii\varepsilon)y\). Since \(H\subseteq H^{\dagger}\), this reads \((H^{\dagger}-\ii\varepsilon)(y-x)=0\), so \(y-x=0\) by Step 1 and \(y=x\in D(H)\). Hence \(D(H^{\dagger})=D(H)\) and \(H=H^{\dagger}\).
∎The converse: exponentiating a self-adjoint operator
The converse needs a projection-valued measure for an operator that is in general unbounded, and Theorem A.238 supplies one only for bounded operators. The gap is closed by the Cayley transform: the unbounded \(H\) is traded for a unitary operator \(V\), for which the construction of The Spectral Theorem for a Bounded Self-Adjoint Operator goes through with three changes displayed below, and the measure is then transported back along an explicit Möbius map.
Let \(H\) be self-adjoint and let \(\varepsilon>0\) be an energy. Then \(H\pm\ii\varepsilon\) maps \(D(H)\) bijectively onto \(\mathcal{H}\), the operator
is unitary, \(\ker(\identity-V)=\set{0}\), and
Rests on Proposition A.259, Definition 12.72 and Proposition 12.73.
Derives Lemma A.260. The identity Equation (A.479) holds for any symmetric \(H\), so \(H\pm\ii\varepsilon\) is injective with \(\norm{(H\pm\ii\varepsilon)x}\geq\varepsilon\norm{x}\); and \(H\) is closed (Proposition 12.73), so, exactly as in Step 2 of Proposition A.259, the image is closed. It is dense because \(\bigl(\im(H\pm\ii\varepsilon)\bigr)^{\perp} =\ker(H^{\dagger}\mp\ii\varepsilon)=\ker(H\mp\ii\varepsilon) =\set{0}\), the middle equality being \(H=H^{\dagger}\) and the last one Equation (A.479) again. Hence \(\im(H\pm\ii\varepsilon)=\mathcal{H}\) and \(H\pm\ii\varepsilon\) is a bijection of \(D(H)\) onto \(\mathcal{H}\). Thus \(V\) is defined on all of \(\mathcal{H}\), and by Equation (A.479)
so \(V\) is isometric; its image is \(\im(H-\ii\varepsilon)=\mathcal{H}\), so it is unitary. Writing \(u=(H+\ii\varepsilon)x\),
which is Equation (A.481) — both the formula for the inverse and, since \(x\) runs over \(D(H)\) as \(u\) runs over \(\mathcal{H}\), the identification of \(\im(\identity-V)\) with \(D(H)\). If \(Vu=u\) then Equation (A.482) gives \(x=0\), hence \(u=(H+\ii\varepsilon)x=0\).
∎Let \(V\in\mathcal{B}(\mathcal{H})\) be unitary and let \(\mathbb{T}=\set{z\in\C\mid\abs{z}=1}\). Then \(\sigma(V)\subseteq\mathbb{T}\), and there is a unique projection-valued measure \(F\) on \(\mathbb{T}\), supported on \(\sigma(V)\), with \(V=\int z\,\dd F(z)\); the associated bounded Borel calculus \(g\longmapsto g(V)\) has all the properties of Proposition A.246. Rests on Theorem A.238, Proposition A.246 and Proposition 12.52.
Derives Proposition A.261. The construction of The Spectral Theorem for a Bounded Self-Adjoint Operator is repeated with Laurent polynomials \(q(z)=\sum_{k=-m}^{m}c_{k}z^{k}\) in place of polynomials, \(q(V)=\sum_{k}c_{k}V^{k}\) with \(V^{-1}=V^{\dagger}\). Three steps need a new argument; everything after them — Propositions A.245 and A.246 and the construction and uniqueness of the projection-valued measure in The projection-valued measure — uses only that \(\Phi\) is an isometric, multiplicative, conjugation-preserving map on \(C(K)\) for a compact \(K\subseteq\C\), and applies word for word.
(a) The spectrum lies on the circle. \(\norm{V}=1\), so \(\sigma(V)\) lies in the closed unit disc (Proposition 12.52). For \(\abs{z}<1\) write \(V-z\identity=V(\identity-zV^{-1})\); since \(\norm{zV^{-1}}=\abs{z}<1\) the bracket is invertible by the geometric series of Proposition 12.52 and \(V\) is invertible, so \(z\in\rho(V)\).
(b) Spectral mapping. Fix \(\nu\in\C\) and put \(r(z)=z^{m}\bigl(q(z)-\nu\bigr)\), an ordinary polynomial. Then \(r(V)=V^{m}\bigl(q(V)-\nu\identity\bigr)\) with \(V^{m}\) invertible, so \(q(V)-\nu\identity\) is invertible exactly when \(r(V)\) is, i.e. exactly when \(0\notin\sigma(r(V))=r(\sigma(V))\) (Lemma A.240). Since \(z\neq0\) on \(\sigma(V)\subseteq\mathbb{T}\), \(r(z)=0\) there means \(q(z)=\nu\); hence \(\sigma(q(V))=q(\sigma(V))\).
(c) The norm identity. With \(\bar q(z)=\sum_{k}c_{k}^{\ast}z^{-k}\) one has \(q(V)^{\dagger}=\bar q(V)\), and on \(\mathbb{T}\), where \(z^{-1}=z^{\ast}\), \(\bar q(z)=q(z)^{\ast}\). So \((\bar q\,q)(V)\) is self-adjoint and, by Equation (12.25), Lemma A.239 and (b),
Laurent polynomials restricted to the compact set \(\sigma(V)\) form a subalgebra of \(C(\sigma(V))\) that contains the constants, is closed under conjugation (by (c)) and separates points (the function \(z\) does), so Theorem A.242 makes it dense; with Equation (A.483) this yields the isometric continuous calculus, and the rest is The Spectral Theorem for a Bounded Self-Adjoint Operator unchanged.
∎Let \(H\) be self-adjoint. There is a projection-valued measure \(E\) on \(\R\) such that, writing \(\mu_{x}(\Omega)=\braket{x}{E(\Omega)x}\),
for \(x\in D(H)\), \(y\in\mathcal{H}\). For bounded Borel \(g\) the operator \(g(H)=\int g\,\dd E\) is bounded with \(\norm{g(H)}\leq\norm{g}_{\infty}\), and \(g\longmapsto g(H)\) is linear, multiplicative and \(\ast\)-preserving. Rests on Lemma A.260, Proposition A.261 and Proposition A.246.
Derives Proposition A.262. Let \(V\) be the Cayley transform Equation (A.480) and \(F\) its projection-valued measure (Proposition A.261). Write
so that \((H+\ii\varepsilon)^{-1}=w(V)\) by Equation (A.481).
\(h\) is real, and \(F(\set{1})=0\). For \(\abs{z}=1\), \(z\neq1\), multiply numerator and denominator by \(1-z^{\ast}\):
so \(h(z)=-2\varepsilon\operatorname{Im}z/\abs{1-z}^{2}\in\R\), with the unit \(\mathrm{J}\). Since \(w\chi_{\set{1}}=0\) identically, multiplicativity gives \(w(V)F(\set{1})=0\); but \(w(V)\) is injective, so \(F(\set{1})=0\).
The measure. Define \(E(\Omega)=F(h^{-1}(\Omega))\) for Borel \(\Omega\subseteq\R\). Preimages respect complements, intersections and countable unions, so \(E\) inherits every clause of Definition 12.58 from \(F\), with \(E(\R)=F(\mathbb{T}\setminus\set{1})=\identity\). By construction \(\mu_{x}\) is the image of \(\mu^{F}_{x}\) under \(h\), so \(\int(g\circ h)\,\dd\mu^{F}_{x}=\int g\,\dd\mu_{x}\) for every Borel \(g\geq0\) or bounded, and \(g(H):=(g\circ h)(V)\) is the calculus claimed.
The domain. A direct computation on \(\mathbb{T}\), using \(\abs{1\pm z}^{2}=2\pm2\operatorname{Re}z\), gives the identity that makes the transport work:
Now \(D(H)=\im w(V)\) by Equation (A.481), and
Indeed, if \(x=w(V)u\) then \(\mu^{F}_{x}=\abs{w}^{2}\mu^{F}_{u}\) — because \(\braket{x}{F(\Omega)x} =\braket{u}{(\bar w\chi_{\Omega}w)(V)u}\) — so \(\int\abs{w}^{-2}\dd\mu^{F}_{x}=\norm{u}^{2}<\infty\). Conversely, if the integral is finite put \(g_{n}=w^{-1}\chi_{\set{\abs{w}\geq1/n}}\), a bounded Borel function, and \(u_{n}=g_{n}(V)x\); then \(\norm{u_{n}-u_{m}}^{2}=\int\abs{g_{n}-g_{m}}^{2}\dd\mu^{F}_{x} \longrightarrow0\) by dominated convergence (Remark 12.1), the dominating function being \(\abs{w}^{-2}\), and \(w(V)u_{n}=\chi_{\set{\abs{w}\geq1/n}}(V)x \longrightarrow x\) because \(F(\set{w=0})=F(\set{1})=0\). Hence \(x\in\im w(V)\). Combining Equations (A.487) and (A.488) with \(\mu^{F}_{x}(\mathbb{T})=\norm{x}^{2}\),
so the two finiteness conditions agree and the first half of Equation (A.484) holds.
The action. Let \(x=w(V)u\in D(H)\). Then, by Equation (A.481),
On the other side, put \(h_{n}=h\,\chi_{\set{\abs{h}\leq n}}\), a bounded Borel function, so that \(h_{n}(H)=\int_{\abs{\lambda}\leq n}\lambda \,\dd E\). Since \(h_{n}w\longrightarrow hw\) pointwise on \(\mathbb{T}\setminus\set{1}\) with \(\abs{h_{n}w}\leq\abs{hw}=\abs{1+z}/2\leq1\), dominated convergence gives
because \(h(z)w(z)=(1+z)/2\) by Equation (A.485). Pairing with \(y\) and passing to the limit gives the second half of Equation (A.484); the integral converges absolutely because \(\int\abs{\lambda}\,\dd\abs{\mu_{y,x}}\leq\norm{y} \bigl(\int\lambda^{2}\dd\mu_{x}\bigr)^{1/2}\), by Cauchy–Schwarz on the partition sums.
∎Proof of Theorem A.253, converse half. Derives Theorem A.253. Let \(H\) be self-adjoint, \(E\) its projection-valued measure (Proposition A.262), and for \(t\in\R\) put
which is legitimate because \(\abs{f_{t}}=1\): the exponent \(\lambda t/\hbar\) is dimensionless, \(\lambda\) being an energy.
Unitary, and a group. \(f_{t}\bar f_{t}=1\) and \(f_{t}f_{s}=f_{t+s}\), \(f_{0}=1\), so multiplicativity and the \(\ast\)-property of the calculus give \(U(t)^{\dagger}U(t)=U(t)U(t)^{\dagger}=\identity\) and Equation (12.41).
Strong continuity. For \(x\in\mathcal{H}\),
by dominated convergence (Remark 12.1): the integrand tends to \(0\) pointwise and is bounded by the constant \(4\), which is \(\mu_{x}\)-integrable because \(\mu_{x}(\R)=\norm{x}^{2}\).
The generator is \(H\). Let \(x\in D(H)\). Then
whose integrand tends to \(0\) pointwise and is bounded by \(4\lambda^{2}/\hbar^{2}\), since \(\abs{\ee^{\ii\theta}-1}\leq\abs{\theta}\) for real \(\theta\); that bound is \(\mu_{x}\)-integrable exactly because \(x\in D(H)\) (Equation (A.484)). Dominated convergence makes Equation (A.492) tend to \(0\), so \(x\) lies in the domain Equation (A.470) of the generator \(G\) of \(U\), with \(Gx=Hx\). Thus \(H\subseteq G\). But \(G\) is self-adjoint by Proposition A.259 and \(H\) is self-adjoint by hypothesis, so
the middle inclusion because taking adjoints reverses inclusions. Hence \(G=H\): a self-adjoint operator has no proper symmetric extension.
∎If \(U\) and \(W\) are strongly continuous one-parameter unitary groups with the same generator \(H\), then \(U=W\). Consequently \(H\longmapsto\exp(-\ii Ht/\hbar)\) and \(U\longmapsto H\) are mutually inverse bijections between the self-adjoint operators on \(\mathcal{H}\) and the strongly continuous one-parameter unitary groups on \(\mathcal{H}\). Rests on Proposition A.259, Proposition 12.67 and Theorem A.253.
Derives Proposition A.263. Fix \(x\in D(H)\) and \(t\), and set \(g(s)=W(t-s)U(s)x\) for \(s\in[0,t]\). By Proposition 12.67, \(U(s)x\in D(H)\) and both curves are differentiable, so the product rule — legitimate because each factor is differentiable in norm and \(W\) is isometric — gives
since \(H\) commutes with \(W(t-s)\) on \(D(H)\) (again Proposition 12.67). Hence \(g\) is constant: \(W(t)x=g(0)=g(t)=U(t)x\). As \(D(H)\) is dense (Proposition A.256) and both operators are bounded, \(U(t)=W(t)\).
For the bijection: Propositions A.256 and A.259 show that every group has a self-adjoint generator, and Theorem A.253 that every self-adjoint operator is the generator of the group Equation (A.490); the paragraph just proved shows that no group is the exponential of two different self-adjoint operators, since the generator is recovered from the group by Equation (A.470).
∎The direct half — Propositions A.256, A.258 and A.259 — quotes nothing at all: it uses the Riemann integral of Lemma A.254, built here, and the scalar fundamental theorem of calculus (Theorem 7.43).
The converse inherits the two quoted inputs of The Spectral Theorem for a Bounded Self-Adjoint Operator — Stone–Weierstrass (Theorem A.242) and Riesz–Markov (Theorem A.244) — through Proposition A.261, and in addition uses the dominated convergence theorem of Lebesgue integration three times: in Equation (A.488), in Equation (A.491) and in Equation (A.492). That theorem is already declared as quoted in Remark 12.1, where the same declaration covers the completeness of \(L^{2}\). The passage from the bounded spectral theorem to the unbounded one is not quoted: it is carried out in Lemma A.260, Proposition A.261 and Proposition A.262, and the algebra that makes it work — \((\identity-V)/(2\ii\varepsilon) =(H+\ii\varepsilon)^{-1}\) and \(\abs{w}^{-2}=\abs{h}^{2}+\varepsilon^{2}\) — is displayed in full.
Stone's Theorem on One-Parameter Unitary Groups discharges the proof obligation of Theorem 12.66 (Section 12.4.3), of which Proposition 12.67 had proved the elementary part. What the theorem buys the physics is stated in Remark 12.68: the equivalence of the differential and the integrated Schrödinger equation Equation (12.48), and the reason a Hamiltonian must be self-adjoint rather than merely symmetric — symmetry gives no unitary group, hence no conserved probability. That is not a technicality, and Example 12.82 is the proof: there the generator has no self-adjoint extension at all, so by Proposition A.263 there is no unitary dynamics to be had, a point settled in general by Theorem 12.80 and Von Neumann's Criterion for Self-Adjoint Extensions, where the Cayley transform used above is developed for an arbitrary closed symmetric operator.
Von Neumann's Criterion for Self-Adjoint Extensions
This appendix proves Theorem 12.80 of Hilbert Spaces: a closed symmetric operator is self-adjoint exactly when both deficiency indices vanish; it admits self-adjoint extensions exactly when the two indices are equal; and the extensions are then in bijective correspondence with the unitary maps from one deficiency subspace to the other. The proof is von Neumann's [vonNeumann:1930], through the Cayley transform, which converts a closed symmetric operator into an isometry between closed subspaces and a self-adjoint operator into a unitary — so that the extension problem for operators becomes the extension problem for isometries, where it is a matter of counting dimensions. The two examples the criterion was made for, momentum on an interval and momentum on a half-line, are recovered in The two worked cases as corollaries, with the one-parameter family of boundary conditions Equation (12.55) produced explicitly from the unitary parameter.
Throughout, \(A\) is a densely defined closed symmetric operator (Definitions 12.70 and 12.72) and \(\mu>0\) is a fixed number carrying the same SI dimension as \(A\), so that \(A\pm\ii\mu\) is dimensionally consistent, as in Definition 12.79; for the momentum operator of The two worked cases, \(\mu\) is a momentum in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\). The deficiency subspaces and indices are those of Equation (12.53),
where \(\dim\) means the cardinality of an orthonormal basis.
Statement
Let \(A\) be a closed symmetric operator with deficiency indices \((n_{+},n_{-})\). Then
-
\(A\) is self-adjoint if and only if \(n_{+}=n_{-}=0\);
-
the self-adjoint extensions of \(A\) are in bijective correspondence with the unitary maps \(W\!:K_{+}\longrightarrow K_{-}\), the extension \(A_{W}\) attached to \(W\) being
\begin{equation}\tag{A.494} D(A_{W})=D(A)\oplus\set{k-Wk\mid k\in K_{+}}\ec\qquad A_{W}\bigl(x+k-Wk\bigr)=Ax+\ii\mu\left(k+Wk\right)\ec \end{equation}the sum of subspaces being direct; such maps exist if and only if \(n_{+}=n_{-}\), and when \(n_{+}=n_{-}=n<\infty\) they form a family parametrised by \(\U(n)\);
-
if \(n_{+}\neq n_{-}\) there is no self-adjoint extension.
The basic identity
For every \(x\in D(A)\),
so \(A\pm\ii\mu\) is injective with \(\norm{(A\pm\ii\mu)x}\geq\mu\norm{x}\), and \(\norm{(A+\ii\mu)x}=\norm{(A-\ii\mu)x}\). If \(A\) is closed, the subspaces \(\im(A\pm\ii\mu)\) are closed, so that
Rests on Definition 12.72, Definition 12.79 and Theorem 12.18.
Derives Lemma A.267. Expanding and using \(\braket{Ax}{x}=\braket{x}{Ax}\), which is symmetry, the two cross terms cancel:
the sign of the third term coming from the antilinearity Equation (5.36) of the first slot. Both consequences are immediate. For closedness of the image, let \((A+\ii\mu)x_{n}\longrightarrow u\). Applying Equation (A.495) to differences shows that both \((x_{n})\) and \((Ax_{n})\) are Cauchy, hence convergent, say \(x_{n}\longrightarrow x\) and \(Ax_{n}\longrightarrow y\); since \(A\) is closed, \(x\in D(A)\) and \(Ax=y\), so \(u=(A+\ii\mu)x\) lies in the image. Equation (A.496) is then the projection theorem Theorem 12.18 together with Equation (A.493).
∎For a closed symmetric \(A\) the numbers \(n_{\pm}\) are the same for every \(\mu>0\). Rests on Lemma A.267 and Definition 12.79.
Derives Lemma A.268. Write \(K_{+}(\mu)=\ker(A^{\dagger}-\ii\mu)\) and suppose \(\dim K_{+}(\nu)<\dim K_{+}(\mu)\) for two positive \(\mu,\nu\) of the same dimension. Then there is a non-zero \(u\in K_{+}(\mu)\cap K_{+}(\nu)^{\perp}\): otherwise the orthogonal projection onto \(K_{+}(\nu)\) would be injective on \(K_{+}(\mu)\), and its adjoint — the projection onto \(K_{+}(\mu)\) restricted to \(K_{+}(\nu)\) — would have dense range in \(K_{+}(\mu)\), forcing \(\dim K_{+}(\mu)\leq\dim K_{+}(\nu)\).
By Equations (A.493) and (A.496), \(K_{+}(\nu)^{\perp}=\im(A+\ii\nu)\), so \(u=(A+\ii\nu)x\) for some \(x\in D(A)\); and \(A^{\dagger}u=\ii\mu u\). Hence, using Equation (12.50) and Equation (5.36),
But Equation (A.495) gives \(\norm{u}\geq\nu\norm{x}\), so \(\abs{\braket{u}{x}}\leq\norm{u}\norm{x}\leq\norm{u}^{2}/\nu\) by Equation (12.1), and Equation (A.497) forces \(\norm{u}^{2}\leq\abs{\nu-\mu}\,\norm{u}^{2}/\nu\), i.e.\ \(\abs{\nu-\mu}\geq\nu\). So whenever \(\abs{\nu-\mu}<\nu\) no such \(u\) exists and \(\dim K_{+}(\mu)\leq\dim K_{+}(\nu)\); exchanging the roles of \(\mu\) and \(\nu\), the two dimensions are equal whenever \(\abs{\nu-\mu}<\min(\mu,\nu)\). The function \(\mu\longmapsto n_{+}(\mu)\) is therefore locally constant on the connected set \((0,\infty)\), hence constant (Theorem 6.14). The argument for \(K_{-}\) is the same with \(\ii\) replaced by \(-\ii\).
∎The Cayley transform
The Cayley transform of a closed symmetric operator \(A\) is the map
Rests on Lemma A.267 and Definition 12.79.
\(V_{A}\) is a well-defined linear isometry of the closed subspace \(K_{+}^{\perp}=\im(A+\ii\mu)\) onto the closed subspace \(K_{-}^{\perp}=\im(A-\ii\mu)\). Moreover \(\identity-V_{A}\) is injective on \(K_{+}^{\perp}\), its image is \(D(A)\) up to the factor \(2\ii\mu\),
and consequently \(A\) is recovered from \(V_{A}\) by
Rests on Definition A.269 and Lemma A.267.
Derives Proposition A.270. \(A+\ii\mu\) is injective on \(D(A)\) (Lemma A.267), so every \(u\in\im(A+\ii\mu)\) is \((A+\ii\mu)x\) for exactly one \(x\) and Equation (A.498) defines \(V_{A}u\) unambiguously; linearity is inherited from that of \(A\). Isometry is the second consequence in Lemma A.267, \(\norm{V_{A}u}=\norm{(A-\ii\mu)x}=\norm{(A+\ii\mu)x}=\norm{u}\), and the image is \(\im(A-\ii\mu)\) by construction. Both subspaces are closed and are the orthogonal complements of \(K_{\pm}\), by Equation (A.496).
The two identities Equation (A.499) are immediate: \((A+\ii\mu)x-(A-\ii\mu)x=2\ii\mu x\) and \((A+\ii\mu)x+(A-\ii\mu)x=2Ax\). The first shows that \((\identity-V_{A})u=0\) forces \(x=0\), hence \(u=0\): injectivity; that \(\im(\identity-V_{A})=D(A)\), since \(x\) runs over \(D(A)\) as \(u\) runs over \(K_{+}^{\perp}\); and, dividing the second identity by the first, Equation (A.500).
∎From an isometry back to an operator
Let \(V\) be an isometry of a closed subspace of \(\mathcal{H}\) into \(\mathcal{H}\) extending \(V_{A}\). Then \(\ker(\identity-V)=\set{0}\). Rests on Proposition A.270 and Definition 12.72.
Derives Lemma A.271. An isometry preserves inner products, by the polarization identity Equation (12.3) applied on its domain. Let \(Vz=z\) and let \(u\in K_{+}^{\perp}\) be arbitrary. Then
using \(Vz=z\) in the third step. By Equation (A.499), \((\identity-V)u=(\identity-V_{A})u\) ranges over \(2\ii\mu\,D(A)\), which is dense because \(A\) is densely defined. Hence \(z\perp\mathcal{H}\) and \(z=0\).
∎Let \(V\) be an isometry of a closed subspace \(D(V)\) onto a closed subspace, with \(\ker(\identity-V)=\set{0}\) and \(\im(\identity-V)\) dense. Then
defines a densely defined symmetric operator \(B\) whose Cayley transform is \(V\). If \(V\) extends \(V_{A}\) then \(B\) extends \(A\). Rests on Proposition A.270, Lemma A.271 and Definition 12.72.
Derives Lemma A.272. \(B\) is well defined and single valued because \(\identity-V\) is injective, and linear because \(V\) is; its domain is dense by hypothesis.
Symmetry. Let \(z,z'\in D(V)\) and put \(x=(\identity-V)z\), \(y=(\identity-V)z'\), so \(Bx=\ii\mu(\identity+V)z\) and \(By=\ii\mu(\identity+V)z'\). Using \(\braket{Vz}{Vz'}=\braket{z}{z'}\) and the antilinearity of the first slot, which turns the prefactor \(\ii\mu\) into \(-\ii\mu\),
while, the prefactor now sitting in the second slot,
So \(\braket{Bx}{y}=\braket{x}{By}\) for all \(x,y\in D(B)\).
Its Cayley transform is \(V\). From Equation (A.501),
so \(\im(B+\ii\mu)=D(V)\), \(\im(B-\ii\mu)=\im V\), and \(V_{B}\) sends \(2\ii\mu z\) to \(2\ii\mu Vz\): that is, \(V_{B}=V\).
Extension. If \(V\supseteq V_{A}\) then, by Equation (A.499), \(D(B)=\im(\identity-V)\supseteq \im(\identity-V_{A})=D(A)\), and on \(D(A)\) the prescription Equation (A.501) is Equation (A.499) itself, so \(B\) agrees with \(A\) there.
∎A closed symmetric operator \(A\) is self-adjoint if and only if its Cayley transform is unitary on all of \(\mathcal{H}\), that is, if and only if \(n_{+}=n_{-}=0\). Rests on Proposition A.270, Lemma A.267 and Definition 12.72.
Derives Proposition A.273. Suppose \(A=A^{\dagger}\). Then \(K_{\pm}=\ker(A^{\dagger}\mp\ii\mu)=\ker(A\mp\ii\mu)=\set{0}\) by Equation (A.495), so \(n_{\pm}=0\) and, by Equation (A.496), \(V_{A}\) is an isometry of \(\mathcal{H}\) onto \(\mathcal{H}\): unitary.
Conversely suppose \(n_{+}=n_{-}=0\), i.e.\ \(\im(A\pm\ii\mu)=\mathcal{H}\). Let \(y\in D(A^{\dagger})\). There is \(x\in D(A)\) with \((A-\ii\mu)x=(A^{\dagger}-\ii\mu)y\), and since \(A\subseteq A^{\dagger}\) this reads \((A^{\dagger}-\ii\mu)(y-x)=0\), i.e. \(y-x\in K_{+}=\set{0}\). Hence \(y=x\in D(A)\), so \(D(A^{\dagger})=D(A)\) and \(A\) is self-adjoint. This is part (1) of Theorem A.266.
∎Proof of the criterion
Proof of Theorem A.266, parts (2) and (3). Derives Theorem A.266. From an extension to a unitary. Let \(B\supseteq A\) be self-adjoint. By Proposition A.273 its Cayley transform \(V_{B}\) is unitary on \(\mathcal{H}\), and it extends \(V_{A}\): indeed \(\im(A+\ii\mu)\subseteq\im(B+\ii\mu)\) and, on that subspace, \(V_{B}(A+\ii\mu)x=(B-\ii\mu)x=(A-\ii\mu)x\). A unitary carrying the closed subspace \(K_{+}^{\perp}\) onto \(K_{-}^{\perp}\) carries the orthogonal complement onto the orthogonal complement, so \(W=V_{B}|_{K_{+}}\) is a unitary map \(K_{+}\longrightarrow K_{-}\).
From a unitary to an extension. Conversely let \(W\!:K_{+}\longrightarrow K_{-}\) be unitary and define, using the decompositions Equation (A.496),
Both summands on the right are orthogonal — \(V_{A}u\in K_{-}^{\perp}\) and \(Wk\in K_{-}\) — so \(\norm{V_{W}(u+k)}^{2} =\norm{u}^{2}+\norm{k}^{2}\) and \(V_{W}\) is an isometry of \(\mathcal{H}\); its image is \(K_{-}^{\perp}\oplus K_{-}=\mathcal{H}\), so \(V_{W}\) is unitary. By Lemma A.271, \(\identity-V_{W}\) is injective, and its image is dense because
where the middle equality holds because for a unitary \(V_{W}\) the relations \(V_{W}z=z\) and \(V_{W}^{\dagger}z=z\) are equivalent. So Lemma A.272 applies: it produces a densely defined symmetric \(A_{W}\supseteq A\) whose Cayley transform is the unitary \(V_{W}\), and \(A_{W}\) is therefore self-adjoint by Proposition A.273.
The formula Equation (A.494). Write \(z=u+k\) with \(u=(A+\ii\mu)x\), \(x\in D(A)\). Then, by Equations (A.499) and (A.504),
so, dividing by \(2\ii\mu\) and renaming \(k/(2\ii\mu)\) as \(k\) — which is a bijection of the subspace \(K_{+}\), and \(W\) is linear — Equation (A.501) becomes exactly Equation (A.494). The sum there is direct: if \(x+k-Wk=0\) then applying \(\identity-V_{W}\) backwards, i.e. using injectivity on \(2\ii\mu x=(\identity-V_{W})u\) and \(k-Wk=(\identity-V_{W})k\), gives \(u+k=0\) with \(u\perp k\), hence \(u=k=0\).
The correspondence is bijective. The two constructions are inverse to each other: from \(A_{W}\) one recovers \(V_{A_{W}}=V_{W}\) (Lemma A.272) and hence \(W=V_{W}|_{K_{+}}\); and from a self-adjoint \(B\supseteq A\) one gets \(W=V_{B}|_{K_{+}}\), whose associated operator has Cayley transform \(V_{B}\) and therefore equals \(B\), since \(B\) is determined by \(V_{B}\) through Equation (A.500).
Counting. A unitary map \(K_{+}\longrightarrow K_{-}\) exists if and only if the two spaces have orthonormal bases of the same cardinality — given such bases, the induced map is unitary; conversely a unitary carries a basis to a basis — that is, if and only if \(n_{+}=n_{-}\). If \(n_{+}\neq n_{-}\) there is none, and by the first paragraph a self-adjoint extension would produce one: this is part (3). When \(n_{+}=n_{-}=n<\infty\), fixing orthonormal bases of \(K_{\pm}\) identifies the unitary maps with the matrix group \(\U(n)\), a transitive and free action of \(\U(n)\) on the set of extensions.
∎The two worked cases
Both examples of Hilbert Spaces use the same differential expression \(P=-\ii\hbar\,\dd/\dd x\) of Equation (12.54), and both use the boundary form Equation (12.57), which for absolutely continuous \(\varphi,\psi\) with \(\varphi',\psi'\in L^{2}\) reads
an integration by parts legitimate for absolutely continuous functions (Theorem 7.43); on \([0,\infty)\) the term at \(L\) is replaced by a limit at infinity, which vanishes for \(L^{2}\) functions with \(L^{2}\) derivative. As Hilbert Spaces records, \(P^{\dagger}\) is the same differential expression on the maximal domain — absolutely continuous \(\psi\) with \(\psi'\in\mathcal{H}\) and no boundary condition [Reed:1972].
Let \(\mathcal{H}=L^{2}([0,L])\) and let \(P\) be Equation (12.54) on the domain \(D_{0}\) of continuously differentiable functions vanishing at both endpoints. Then \(n_{+}=n_{-}=1\); every \(P_{\theta}\) of Equation (12.55) is self-adjoint; these are all the self-adjoint extensions; and the unitary parameter \(W k_{+}=\ee^{\ii\alpha}k_{-}\) of Theorem A.266 corresponds to the boundary phase
a bijection of the circle onto the circle. Rests on Theorem A.266, Example 12.81 and Equation (12.55).
Derives Corollary A.274. Indices. This is the computation of Example 12.81: the equations \(P^{\dagger}\psi=\pm\ii\mu\psi\) have the one-dimensional solution spaces spanned by
both square integrable on the compact interval, so \(n_{+}=n_{-}=1\) and Theorem A.266 gives a family of extensions parametrised by \(\U(1)\). Normalising, \(N_{+}^{-2}=\int_{0}^{L}\ee^{-2\mu x/\hbar}\dd x =\bigl(\hbar/2\mu\bigr)\left(1-q^{2}\right)\) and \(N_{-}^{-2}=\bigl(\hbar/2\mu\bigr)\left(q^{-2}-1\right)\), whence
Each \(P_{\theta}\) is self-adjoint. Symmetry on \(D_{\theta}\) is Equation (A.505): for \(\varphi,\psi\in D_{\theta}\), \(\varphi(L)^{\ast}\psi(L) =\ee^{-\ii\theta}\ee^{\ii\theta}\varphi(0)^{\ast}\psi(0) =\varphi(0)^{\ast}\psi(0)\) and the right-hand side vanishes. For the converse inclusion let \(y\in D(P_{\theta}^{\dagger})\). Since \(D_{0}\subseteq D_{\theta}\), \(y\) belongs to the maximal domain \(D(P^{\dagger})\) and \(P_{\theta}^{\dagger}y=P^{\dagger}y\); so Equation (A.505) with \(\varphi=y\) must vanish for every \(\psi\in D_{\theta}\):
Choosing \(\psi(x)=1+(\ee^{\ii\theta}-1)x/L\), which lies in \(D_{\theta}\) and has \(\psi(0)=1\), gives \(y(L)^{\ast}\ee^{\ii\theta}=y(0)^{\ast}\), i.e.\ \(y(L)=\ee^{\ii\theta}y(0)\) and \(y\in D_{\theta}\). Hence \(D(P_{\theta}^{\dagger})=D_{\theta}\) and \(P_{\theta}\) is self-adjoint.
There are no others. Let \(B\) be a self-adjoint extension of \(P\). Then \(P\subseteq B=B^{\dagger}\subseteq P^{\dagger}\), so \(D(B)\) lies in the maximal domain and Equation (A.505) vanishes on \(D(B)\times D(B)\). Taking \(\varphi=\psi\) gives \(\abs{\psi(L)}^{2}=\abs{\psi(0)}^{2}\) for every \(\psi\in D(B)\): the image \(S\subseteq\C^{2}\) of the boundary map \(\psi\longmapsto(\psi(0),\psi(L))\) is a subspace isotropic for the hermitian form \(\abs{b}^{2}-\abs{a}^{2}\), which has signature \((1,1)\), so \(\dim S\leq1\). If \(\dim S=1\), say \(S=\C\,(1,\ee^{\ii\theta})\), then \(D(B)\subseteq D_{\theta}\), i.e. \(B\subseteq P_{\theta}\); taking adjoints reverses the inclusion, so \(P_{\theta}=P_{\theta}^{\dagger}\subseteq B^{\dagger}=B\) and \(B=P_{\theta}\). If \(S=\set{0}\) then \(D(B)\) is contained in \(D_{\theta}\) for every \(\theta\), so \(B\subseteq P_{\theta}\) and the same argument gives \(B=P_{\theta}\) — contradicting \(S=\set{0}\), since the domain of \(P_{\theta}\) contains \(\psi(x) =1+(\ee^{\ii\theta}-1)x/L\), whose boundary values are not both zero. So every self-adjoint extension is one of the \(P_{\theta}\).
Matching the parameters. Let \(W k_{+}=\ee^{\ii\alpha}k_{-}\) and let \(A\) be the closure of \(P\), a closed symmetric operator with the same adjoint and the same deficiency subspaces — the defining condition Equation (12.50) for \(A^{\dagger}\) extends from \(D_{0}\) to its graph closure by continuity of the inner product. By Equation (A.494), every element of \(D(A_{W})\) is \(\psi=\psi_{0}+c\left(k_{+}-\ee^{\ii\alpha}k_{-}\right)\) with \(\psi_{0}\in D(A)\) and \(c\in\C\). Every \(\psi_{0}\in D(A)\) vanishes at both endpoints: if \(\psi_{n}\in D_{0}\) converges to \(\psi_{0}\) in the graph norm then, for absolutely continuous \(\phi\) on \([0,L]\),
— average \(\phi(x)=\phi(y)+\int_{y}^{x}\phi'\) over \(y\) and use Equation (12.1) — so graph convergence implies uniform convergence and \(\psi_{0}(0)=\lim\psi_{n}(0)=0\), likewise at \(L\). Therefore the boundary values of \(\psi\) come from the second term alone: by Equations (A.507) and (A.508),
using \(k_{+}(L)=N_{+}q\) and \(k_{-}(L)=N_{-}q^{-1}=N_{+}\). The denominator \(1-\ee^{\ii\alpha}q\) never vanishes because \(q<1\), and
so the ratio \(\psi(L)/\psi(0)\) is a phase, which is Equation (A.506). Hence \(D(A_{W})\subseteq D_{\theta}\) with that \(\theta\), and since both operators are self-adjoint, \(A_{W}=P_{\theta}\) by the adjoint argument used above. Finally \(z\longmapsto(q-z)/(1-qz)\) is a Möbius transformation with \(\abs{q}<1\); Equation (A.510) shows it maps the unit circle into itself, and it is injective, so it is a bijection of the circle: as \(\alpha\) runs once around, so does \(\theta\), and the \(\U(1)\) of Theorem A.266 is exactly the circle of boundary conditions Equation (12.55).
∎Let \(\mathcal{H}=L^{2}([0,\infty))\) and let \(P\) be Equation (12.54) on the continuously differentiable functions of compact support in \((0,\infty)\). Then \((n_{+},n_{-})=(1,0)\) and \(P\) has no self-adjoint extension. Rests on Theorem A.266 and Example 12.82.
Derives Corollary A.275. The solutions of \(P^{\dagger}\psi=\pm\ii\mu\psi\) are again \(\psi(x)=C\ee^{\mp\mu x/\hbar}\); on the half-line the decaying one is square integrable, with \(\norm{\psi}^{2} =\abs{C}^{2}\hbar/(2\mu)\), and the growing one is not, so \(n_{+}=1\) and \(n_{-}=0\) (Example 12.82). By part (3) of Theorem A.266 — there is no unitary map from a one-dimensional space onto \(\set{0}\) — there is no self-adjoint extension. By Proposition A.263 there is therefore no strongly continuous unitary group whose generator is a momentum on the half-line, which is the statement that probability cannot be conserved under translations that push it off the end of the half-line (Remark 12.83).
∎The criterion itself — Lemma A.267 through Theorem A.266 — quotes nothing: every step is carried out from the definitions of Section 12.5 and the projection theorem. Two facts are imported in the worked cases of The two worked cases, both of them classical real analysis rather than operator theory. First, that the adjoint of \(P\) on either interval is the same differential expression on the maximal domain of absolutely continuous functions with square-integrable derivative; this is stated without proof in Example 12.81 as well, with the reference [Reed:1972], and identifying it requires the du Bois-Reymond lemma on weak derivatives. Second, that integration by parts Equation (A.505) is valid for absolutely continuous functions, which is Theorem 7.43 applied to the product. Neither is needed for Theorem A.266.
Von Neumann's Criterion for Self-Adjoint Extensions discharges the proof obligation of Theorem 12.80 (Section 12.5.2), and with Corollaries A.274 and A.275 it completes the two examples that section is built around — in particular the assertion of Example 12.81 that \(D_{\theta}\) exhausts the possibilities, which is proved above by the isotropic-subspace count and not merely asserted. Read with Stone's Theorem on One-Parameter Unitary Groups, the criterion is the precise statement of what Remark 12.77 calls the physics of a domain: the deficiency indices decide whether a symmetric differential expression admits a unitary dynamics at all, and when they are equal and non-zero it is the boundary condition — the ring, or the ring threaded by a flux — that selects one, the mathematics offering the whole \(\U(n)\) and choosing nothing.
The Nuclear Spectral Theorem of Gelfand and Maurin
This appendix proves Theorem 12.107 of Hilbert Spaces — that a self-adjoint operator acting continuously on the test space of a Gelfand triple possesses a complete set of generalized eigenvectors — and says exactly what the proof assumes. It is the theorem that licenses the completeness relation \(\int\ketbra{\lambda}{\lambda}\,\dd\mu(\lambda)=\identity\) of Remark 12.108, and with it Dirac's formalism [Dirac:1930b] in the form physics actually uses it. The result is due to Gelfand and his collaborators [Gelfand:1964], where it appears in volume 4 of Generalized Functions; Maurin obtained it independently at about the same time, and the theorem carries both names.
The architecture has three stages, and only the second imports anything. First, the spectral theorem of The Spectral Theorem for a Bounded Self-Adjoint Operator and Stone's Theorem on One-Parameter Unitary Groups is read as a direct-integral decomposition: \(\mathcal{H}\) is realised as a space of square-integrable fields \(\lambda\longmapsto\hat x(\lambda)\) over the spectrum, in which \(A\) acts by multiplication by \(\lambda\). At that stage the assignment \(x\longmapsto\hat x(\lambda)\) is defined only up to a \(\mu\)-null set, and only as a map of equivalence classes: for a fixed \(\lambda\) it is not a map on \(\mathcal{H}\) at all, so there are no eigenvectors yet, generalized or otherwise. Second, nuclearity of the test space is used — and this is the only import — to produce, for \(\mu\)-almost every \(\lambda\), a genuine continuous map \(\Phi\longrightarrow\mathcal{H}_{\lambda}\) that is a version of the fibre map. Third, each vector in the fibre is paired with that map, giving a continuous linear functional on \(\Phi\), and the multiplication form of \(A\) makes it a generalized eigenvector of eigenvalue \(\lambda\); the isometry of the first stage, read fibrewise, is the completeness relation.
Hypotheses and statement
The standing hypotheses of Definition 12.103 are made explicit here, because the proof uses each of them.
A countably Hilbert space is a vector space \(\Phi\) whose topology is given by an increasing sequence of norms \(\norm{\cdot}_{1}\leq\norm{\cdot}_{2}\leq\cdots\), each coming from an inner product, and which is complete for the resulting metric; \(\Phi_{p}\) denotes the Hilbert space completion of \(\Phi\) in \(\norm{\cdot}_{p}\). It is nuclear if for every \(p\) there is \(q\geq p\) such that the canonical map \(\Phi_{q}\longrightarrow\Phi_{p}\) is nuclear, that is, of the form \(u\longmapsto\sum_{k}s_{k}\braket{a_{k}}{u}\,b_{k}\) with \(\sum_{k}\abs{s_{k}}<\infty\) and \(\set{a_{k}}\), \(\set{b_{k}}\) orthonormal. Rests on Definitions 12.2 and 12.103.
Let \(\Phi\subseteq\mathcal{H}\subseteq\Phi'\) be a Gelfand triple (Definition 12.103) in which \(\mathcal{H}\) is separable and \(\Phi\) is a separable countably Hilbert nuclear space, densely and continuously embedded in \(\mathcal{H}\). Let \(A\) be self-adjoint on \(\mathcal{H}\) and map \(\Phi\) continuously into itself. Then there exist a finite Borel measure \(\mu\) on \(\sigma(A)\), separable Hilbert spaces \(\mathcal{H}_{\lambda}\) and, for \(\mu\)-almost every \(\lambda\), a continuous linear map \(\Lambda_{\lambda}\!:\Phi\longrightarrow\mathcal{H}_{\lambda}\) such that
-
\(\Lambda_{\lambda}(A\varphi)=\lambda\,\Lambda_{\lambda}(\varphi)\) for every \(\varphi\in\Phi\) and \(\mu\)-almost every \(\lambda\);
-
for every \(\xi\in\mathcal{H}_{\lambda}\) the functional \(F_{\lambda,\xi}(\varphi) =\braket{\xi}{\Lambda_{\lambda}\varphi}_{\mathcal{H}_{\lambda}}\) belongs to \(\Phi'\) and is a generalized eigenvector of \(A\) with eigenvalue \(\lambda\) (Definition 12.105);
-
writing \(\set{\xi_{n}}\) for an orthonormal basis of \(\mathcal{H}_{\lambda}\) and \(F_{\lambda,n}=F_{\lambda,\xi_{n}}\),
\begin{equation}\tag{A.511} \braket{\varphi}{\psi}=\int_{\sigma(A)}\sum_{n} F_{\lambda,n}(\varphi)^{\ast}\,F_{\lambda,n}(\psi)\, \dd\mu(\lambda)\ec\qquad \forall\,\varphi,\psi\in\Phi\ep \end{equation}
When the spectrum of \(A\) is simple — when \(\mathcal{H}\) has a single cyclic vector — every \(\mathcal{H}_{\lambda}\) is one dimensional, the sum in Equation (A.511) has one term, and the identity is Equation (12.69) exactly as stated in Hilbert Spaces. Rests on Definition 12.105, Definition 12.103 and Theorem A.238.
The direct integral
Let \(A\) be self-adjoint on a separable \(\mathcal{H}\), with projection-valued measure \(E\) (Theorem A.238 and Proposition A.262). There are a finite Borel measure \(\mu\) on \(\R\), supported on \(\sigma(A)\), and an isometry
such that for \(x\in D(A)\) one has \(\widehat{Ax}(\lambda)=\lambda\,\hat x(\lambda)\) for \(\mu\)-almost every \(\lambda\). Rests on Theorem A.238, Lemma A.250 and Proposition A.262.
Derives Proposition A.280. Cyclic decomposition. The construction of Lemma A.250 applies verbatim with the bounded Borel calculus of Proposition A.246 — or, for unbounded \(A\), of Proposition A.262 — in place of the continuous one: it used only that the cyclic subspaces \(\mathcal{H}_{n}=\overline{\set{g(A)x_{n}\mid g\ \text{bounded Borel}}}\) are \(A\)-invariant with \(A\)-invariant orthogonal complements, and that \(\mathcal{H}\) is separable. So \(\mathcal{H}=\bigoplus_{n}\mathcal{H}_{n}\) with each \(x_{n}\) cyclic in \(\mathcal{H}_{n}\).
Each block. As in Lemma A.249, \(\norm{g(A)x_{n}}^{2}=\int\abs{g}^{2}\dd\mu_{n}\) with \(\mu_{n}(\Omega)=\braket{x_{n}}{E(\Omega)x_{n}}\), so \(g(A)x_{n}\longmapsto g\) extends to a unitary \(U_{n}\!:\mathcal{H}_{n}\longrightarrow L^{2}(\R,\mu_{n})\) — the bounded Borel functions are dense in \(L^{2}(\mu_{n})\), since the simple functions already are — and \(U_{n}AU_{n}^{-1}\) is multiplication by \(\lambda\).
One measure. Put \(\mu=\sum_{n}2^{-n}\norm{x_{n}}^{-2}\mu_{n}\), a Borel measure with \(\mu(\R)\leq1\), supported on \(\sigma(A)\) because every \(\mu_{n}\) is. Each \(\mu_{n}\) is absolutely continuous with respect to \(\mu\), so the Radon–Nikodym theorem (Remark A.284) provides densities \(\rho_{n}=\dd\mu_{n}/\dd\mu\geq0\).
The isometry. For \(x=\sum_{n}v_{n}\) with \(v_{n}\in\mathcal{H}_{n}\) write \(h_{n}=U_{n}v_{n}\in L^{2}(\mu_{n})\) and define
Then, exchanging sum and integral by monotone convergence (Remark 12.1),
which is Equation (A.512); in particular \(\hat x(\lambda)\) lies in \(\ell^{2}(\N)\) for \(\mu\)-almost every \(\lambda\). Since \(A\) acts on each block as multiplication by \(\lambda\), and multiplication commutes with the fibrewise rescaling by \(\sqrt{\rho_{n}}\), \(\widehat{Ax}(\lambda)=\lambda\hat x(\lambda)\) almost everywhere for \(x\in D(A)\).
∎The map \(J\) is the whole content of the spectral theorem written fibrewise, and it is also the whole of the difficulty: \(\hat x\) is an equivalence class of fields modulo \(\mu\)-null sets, so the expression \(\hat x(\lambda)\) for one fixed \(\lambda\) has no meaning, and there is nothing yet that could be evaluated on a test function.
What nuclearity buys
Let \(\Phi\) be a countably Hilbert nuclear space (Definition A.278) continuously embedded in a Hilbert space \(\mathcal{H}\). Then there is an index \(q\) such that the embedding extends to a Hilbert–Schmidt map \(J_{q}\!:\Phi_{q}\longrightarrow\mathcal{H}\), that is, one for which
for one, hence for every, orthonormal basis \(\set{e_{k}}\) of \(\Phi_{q}\). Rests on Definition A.278.
The Schwartz triple Equation (12.66) is the case to keep in mind: \(\mathcal{S}(\R)\) is countably Hilbert for the norms built from the harmonic-oscillator Hermite expansion, and Equation (A.514) holds there because the Hermite coefficients of a Schwartz function decay faster than any power.
Under the hypotheses of Theorem A.279 there are, for \(\mu\)-almost every \(\lambda\), a constant \(C(\lambda)<\infty\) and a linear map \(\Lambda_{\lambda}\!:\Phi_{q}\longrightarrow\ell^{2}(\N)\) with
such that for each fixed \(\varphi\in\Phi\) one has \(\Lambda_{\lambda}\varphi=\hat\varphi(\lambda)\) for \(\mu\)-almost every \(\lambda\). In particular \(\Lambda_{\lambda}\) is continuous for the topology of \(\Phi\). Rests on Theorem A.281 and Proposition A.280.
Derives Proposition A.282. Let \(q\) and \(\set{e_{k}}\) be as in Theorem A.281. Applying Equation (A.512) to each \(e_{k}\) and summing,
the exchange of sum and integral being monotone convergence for a series of non-negative terms (Remark 12.1). A non-negative function with finite integral is finite almost everywhere, so
This is the entire role of nuclearity: without Equation (A.514) the left-hand side of Equation (A.516) is a divergent series and Equation (A.517) says nothing.
Fix such a \(\lambda\). For \(\varphi\in\Phi_{q}\) expand \(\varphi=\sum_{k}c_{k}e_{k}\) with \(c_{k}=\braket{e_{k}}{\varphi}_{q}\) and \(\sum_{k}\abs{c_{k}}^{2}=\norm{\varphi}_{q}^{2}\) (Theorem 12.30), and put
The series converges absolutely in \(\ell^{2}(\N)\), since by Cauchy–Schwarz
which is also Equation (A.515). Linearity is clear.
It remains to identify \(\Lambda_{\lambda}\varphi\) with the fibre of \(\varphi\). Fix \(\varphi\in\Phi\) and let \(S_{N}=\sum_{k\leq N}c_{k}e_{k}\). Then \(S_{N}\longrightarrow\varphi\) in \(\Phi_{q}\), hence in \(\mathcal{H}\) because the embedding is bounded, and \(J\) is isometric, so
A sequence converging in \(L^{2}(\mu)\) has a subsequence converging \(\mu\)-almost everywhere (Remark A.284), so along it \(\hat S_{N}(\lambda)\longrightarrow\hat\varphi(\lambda)\) for almost every \(\lambda\). But \(\hat S_{N}(\lambda)\) is the \(N\)th partial sum of Equation (A.518), which converges to \(\Lambda_{\lambda}\varphi\) for every \(\lambda\) satisfying Equation (A.517). Limits are unique, so \(\Lambda_{\lambda}\varphi=\hat\varphi(\lambda)\) almost everywhere. Finally \(\norm{\cdot}_{q}\) is one of the norms defining the topology of \(\Phi\), so Equation (A.515) is continuity on \(\Phi\).
∎Generalized eigenvectors, and completeness
Proof of Theorem A.279. Derives Theorem A.279. Take \(\mu\), \(\mathcal{H}_{\lambda}=\ell^{2}(\N)\) and \(\Lambda_{\lambda}\) from Propositions A.280 and A.282.
(1) The eigenvalue relation. Since \(A\) maps \(\Phi\) into \(\Phi\), in particular \(\Phi\subseteq D(A)\), and Proposition A.280 gives, for each fixed \(\varphi\in\Phi\),
the exceptional null set depending on \(\varphi\). Let \(\set{\varphi_{j}}\) be dense in \(\Phi\), which exists because \(\Phi\) is separable, and let \(Z\) be the union of the countably many exceptional sets attached to the \(\varphi_{j}\) together with the null set outside which Equation (A.517) holds; \(Z\) is \(\mu\)-null. For \(\lambda\notin Z\) the two maps \(\varphi\longmapsto\Lambda_{\lambda}(A\varphi)\) and \(\varphi\longmapsto\lambda\Lambda_{\lambda}\varphi\) are continuous on \(\Phi\) — the first because \(A\) is continuous from \(\Phi\) to \(\Phi\) and \(\Lambda_{\lambda}\) is continuous on \(\Phi\), the second by Equation (A.515) — and they agree on the dense set \(\set{\varphi_{j}}\), hence everywhere. This is (1), now with one null set for all \(\varphi\) at once.
(2) The functionals. For \(\lambda\notin Z\) and \(\xi\in\mathcal{H}_{\lambda}\) put \(F_{\lambda,\xi}(\varphi) =\braket{\xi}{\Lambda_{\lambda}\varphi}\). It is linear in \(\varphi\), because \(\Lambda_{\lambda}\) is linear and the inner product is linear in its second slot (Equation (5.34)), and continuous, since \(\abs{F_{\lambda,\xi}(\varphi)}\leq\norm{\xi}C(\lambda) \norm{\varphi}_{q}\) by Equation (12.1) and Equation (A.515); so \(F_{\lambda,\xi}\in\Phi'\) (Definition 12.103). By (1) and the reality of \(\lambda\),
which is Equation (12.68): \(F_{\lambda,\xi}\) is a generalized eigenvector of \(A\) with eigenvalue \(\lambda\).
(3) Completeness. Polarising the isometry Equation (A.512) — by Equation (12.3), an isometry of Hilbert spaces preserves inner products — gives, for \(\varphi,\psi\in\Phi\),
Let \(\set{\xi_{n}}\) be the standard orthonormal basis of \(\ell^{2}(\N)\), the same for every \(\lambda\), and \(F_{\lambda,n}=F_{\lambda,\xi_{n}}\). Parseval's identity Equation (12.20) in the fibre reads
using \(\Lambda_{\lambda}\varphi=\hat\varphi(\lambda)\) almost everywhere and \(\braket{\hat\varphi}{\xi_{n}} =\braket{\xi_{n}}{\hat\varphi}^{\ast}=F_{\lambda,n}(\varphi)^{\ast}\). Substituting into Equation (A.520) gives Equation (A.511). Since \(\mu\) is supported on \(\sigma(A)\), the integral may be written over \(\sigma(A)\).
Simple spectrum. If a single cyclic vector generates \(\mathcal{H}\) the decomposition of Proposition A.280 has one block, \(\ell^{2}(\N)\) is replaced by \(\C\), and Equation (A.511) collapses to \(\braket{\varphi}{\psi}=\int F_{\lambda}(\varphi)^{\ast} F_{\lambda}(\psi)\,\dd\mu\), which is Equation (12.69).
∎For the Schwartz triple Equation (12.66) and \(A=P=-\ii\hbar\,\dd/\dd x\), the spectrum is simple, \(\mu\) may be taken to be Lebesgue measure on \(\R\) in the variable \(p\), and \(\Lambda_{p}\varphi\) is the Fourier transform Equation (12.67), so that \(F_{p}=F_{p}\) of Proposition 12.106. Statement (2) of Theorem A.279 is then the computation already carried out there, and Equation (A.511) is the Plancherel identity of Fourier Analysis and Integral Transforms. The plane waves are a complete set of generalized eigenvectors of the momentum, and no one of them is a state. Rests on Theorem A.279, Proposition 12.106 and Example 12.104.
This is the one long proof of Hilbert Spaces that rests on a theory the treatise does not build, and the reader is owed a precise account of where.
-
Theorem A.281 — that a countably Hilbert nuclear space embeds in \(\mathcal{H}\) through some Hilbert–Schmidt map — is assumed. It is the definition of nuclearity (Definition A.278) plus the elementary fact that the composition of a nuclear map with a bounded one is Hilbert–Schmidt, but the theory of nuclear spaces itself, including the proof that \(\mathcal{S}(\R)\) is nuclear, is not developed anywhere in this treatise. Reference of record: [Gelfand:1964], volume 4, chapter I. Everything that nuclearity is used for is Equations (A.516) and (A.517) and nothing else — which is why the architecture above is worth having even though the input is quoted: the reader can see exactly how much is being bought, and how little.
-
Three standard theorems of measure theory are used: the Radon–Nikodym theorem, in Proposition A.280; the monotone convergence theorem, twice; and the fact that a sequence convergent in \(L^{2}\) has an almost-everywhere convergent subsequence. The treatise develops no measure theory and says so in Remark 12.1, where the convergence theorems are already declared.
-
The spectral theorem is not quoted: the direct integral of Proposition A.280 is built from The Spectral Theorem for a Bounded Self-Adjoint Operator for bounded \(A\) and from Proposition A.262 for unbounded \(A\) — the momentum operator of Example A.283 is the unbounded case.
-
Two hypotheses of Theorem A.279 are suppressed in the chapter's statement Theorem 12.107 and are used here: separability of \(\mathcal{H}\), without which the cyclic decomposition need not be countable, and separability of \(\Phi\), without which the single null set of Equation (A.519) cannot be assembled. Both hold for the Schwartz triple. A third point of honesty: for a spectrum of multiplicity greater than one the completeness relation carries the sum over \(n\) displayed in Equation (A.511), which Equation (12.69) suppresses; the two agree exactly when the spectrum is simple.
Nothing above claims more. In particular the theorem does not say that the \(F_{\lambda,n}\) are unique, nor that every generalized eigenvector arises this way, nor that \(\mu\) is canonical — only its measure class is — and no use made of the theorem in this treatise requires any of those.
The Nuclear Spectral Theorem of Gelfand and Maurin discharges the proof obligation of Theorem 12.107 (Section 12.6.3). What it licenses is the working rule of Remark 12.108: a ket belonging to a point of continuous spectrum is an element of \(\Phi'\), the pairing \(\braket{\lambda}{\varphi}\) is the number \(F_{\lambda}(\varphi)\) for a test function \(\varphi\), and the “normalisation” \(\braket{\lambda}{\lambda'} =\delta(\lambda-\lambda')\) is shorthand for Equation (A.511) — an identity between functionals, never between numbers. Proposition 12.102 is the proof that no other reading is available, and Example A.283 is the case every scattering calculation of Scattering Theory silently uses.
The Implicit Function Theorem
This appendix proves the one input that Theorem 13.59 of Differentiable Manifolds, Tensors, and Curvature takes from multivariable calculus and that Real Analysis does not state: where the derivative of a system of equations with respect to one block of variables is invertible, the system can be solved locally for that block as a differentiable function of the remaining variables, and the solution is as smooth as the system. It is the theorem that turns a level set into a graph, and through the regular value theorem it is what makes a sphere, a hyperboloid, a constraint surface or a mass shell a manifold at all.
What is assumed as already available is exactly the several-variable differential calculus of Section 7.10.2: partial derivatives and the gradient (Definition 7.65), the class \(C^{1}\) (Definition 7.66), differentiability at a point with the uniqueness of the differential (Definitions 7.67 and 7.70), the theorem that a \(C^{1}\) map is differentiable (Theorem 7.68), the chain rule (Proposition 7.72), the one-variable mean value theorem (Theorem 7.35), the Cauchy criterion (Theorem 7.8) and the geometric series (Proposition 7.46). No fixed-point theorem is imported: the contraction is carried out explicitly, in the manner of the Picard iteration of Theorem 9.8, with the geometric estimate written out. Nothing else is quoted — see Remark A.295.
Throughout, \(\abs{\cdot}\) is the Euclidean length (Equation (6.12)) and, for a real matrix \(A\),
by the Cauchy–Schwarz inequality of Linear Algebra and Representation Theory applied row by row. Points of \(\R^{m}\times\R^{n}\) are written \((x,y)\) with \(x\in\R^{m}\), \(y\in\R^{n}\); for a map \(F\) into \(\R^{n}\) we write
the first an \(n\times m\) matrix, the second an \(n\times n\) matrix.
Statement
Let \(A\subseteq\R^{m}\times\R^{n}\) be open, let \(F:A\longrightarrow\R^{n}\) be of class \(C^{k}\) with \(k\ge1\), and let \((a,b)\in A\) satisfy
Then there are open sets \(W\subseteq\R^{m}\) with \(a\in W\) and \(V\subseteq\R^{n}\) with \(b\in V\), satisfying \(W\times V\subseteq A\), and a map \(h:W\longrightarrow V\) such that
The map \(h\) is unique: it is the only map \(W\longrightarrow V\) whose graph is the zero set. It is of class \(C^{k}\), and
If \(F\) is smooth, so is \(h\). Rests on Theorem 7.68, Proposition 7.72 and Definition 7.67.
The proof is in four movements: the equation is recast as a fixed-point problem whose iteration contracts (The contraction); the iteration is run and its limit shown to be the unique solution (Existence and uniqueness of the implicit map); the solution is shown Lipschitz and then differentiable, with Equation (A.525) (Continuity, differentiability and the derivative formula); and the smoothness is bootstrapped to \(C^{k}\) there as well. The two consequences the treatise uses draws the two corollaries the treatise uses.
The contraction
Under the hypotheses of Theorem A.286, put
Then there are \(r>0\) and \(\rho\in(0,r]\) such that, writing \(\overline{B}_{r}\) for closed balls, \(\overline{B}_{\rho}(a)\times\overline{B}_{r}(b)\subseteq A\) and:
-
for every \(y\in\overline{B}_{r}(b)\) and \(x\in\overline{B}_{\rho}(a)\), \(F(x,y)=0\) if and only if \(T(x,y)=y\);
-
\(\abs{T(x,y)-T(x,y')}\le\tfrac{1}{2}\abs{y-y'}\) for all \(y,y'\in\overline{B}_{r}(b)\) and \(x\in\overline{B}_{\rho}(a)\);
-
\(\abs{T(x,b)-b}\le r/4\) for every \(x\in\overline{B}_{\rho}(a)\);
-
\(T(x,\cdot)\) maps \(\overline{B}_{r}(b)\) into itself; and
-
\(\det\left[D_{y}F(x,y)\right]\ne0\) throughout \(\overline{B}_{\rho}(a)\times\overline{B}_{r}(b)\).
Rests on Equation (A.523), Theorem 7.35 and Definition 7.66.
Derives Lemma A.287. \(C\) exists by Equation (A.523). Assertion (1) is immediate from Equation (A.526): \(T(x,y)=y\) says \(C\,F(x,y)=0\), and \(C\) is invertible, so this says \(F(x,y)=0\).
Differentiating Equation (A.526) in \(y\),
which vanishes at \((a,b)\) by the definition of \(C\). The entries of \(D_{y}F\) are continuous, because \(F\) is \(C^{1}\) (Definition 7.66), hence so are those of \(D_{y}T\) and so is \(\det\left[D_{y}F\right]\). Choose \(r>0\) with \(\overline{B}_{r}(a)\times\overline{B}_{r}(b)\subseteq A\) and, by that continuity together with \(D_{y}T(a,b)=0\) and \(\det[D_{y}F(a,b)]\ne0\), small enough that
on the whole of \(\overline{B}_{r}(a)\times\overline{B}_{r}(b)\). That is assertion (5).
For (2), fix \(x\) and \(y,y'\) in the convex set \(\overline{B}_{r}(b)\). Applying Theorem 7.35 to \(\lambda\longmapsto T^{i}\left(x,y'+\lambda(y-y')\right)\) on \([0,1]\) gives \(\theta_{i}\in(0,1)\) with
so that \(\abs{T^{i}(x,y)-T^{i}(x,y')} \le\norm{D_{y}T(x,\xi_{i})}\,\abs{y-y'} \le\abs{y-y'}/(2\sqrt{n})\), the row length being at most the whole matrix norm Equation (A.521). Squaring and summing the \(n\) components gives (2).
For (3): \(T(a,b)=b\) because \(F(a,b)=0\), and \(x\longmapsto T(x,b)\) is continuous, so there is \(\rho\in(0,r]\) with \(\abs{T(x,b)-b}\le r/4\) for \(\abs{x-a}\le\rho\).
For (4), combine (2) and (3): for \(y\in\overline{B}_{r}(b)\) and \(\abs{x-a}\le\rho\),
Existence and uniqueness of the implicit map
With \(r,\rho\) as in Lemma A.287, fix \(x\in\overline{B}_{\rho}(a)\) and define
Then every \(y_{j}\) lies in \(\overline{B}_{r}(b)\), the sequence converges to a limit \(h(x)\) with
and \(h(x)\) is the unique point of \(\overline{B}_{r}(b)\) with \(F\bigl(x,h(x)\bigr)=0\). Rests on Lemma A.287, Theorem 7.8 and Proposition 7.46.
Derives Lemma A.288. The iterates stay in \(\overline{B}_{r}(b)\) by Equation (A.529) and induction. By assertion (2) of Lemma A.287,
the last step by assertion (3). For \(l>j\) the triangle inequality (Proposition 7.3) and the geometric series (Proposition 7.46) give
so each of the \(n\) coordinate sequences is a Cauchy sequence of reals and converges (Theorem 7.8); call the limit \(h(x)\), which lies in the closed set \(\overline{B}_{r}(b)\). Letting \(l\to\infty\) with \(j=0\) in the display gives \(\abs{h(x)-b}\le r/2\), which is Equation (A.531). Since \(T(x,\cdot)\) is continuous (indeed Lipschitz, by assertion (2)), passing to the limit in Equation (A.530) gives \(T(x,h(x))=h(x)\), i.e.\ \(F(x,h(x))=0\) by assertion (1).
Uniqueness in \(\overline{B}_{r}(b)\): if \(y\) and \(y'\) both satisfy \(F(x,\cdot)=0\) there, both are fixed points of \(T(x,\cdot)\), and assertion (2) gives \(\abs{y-y'}\le\tfrac12\abs{y-y'}\), hence \(y=y'\).
∎Put \(W=B_{\rho}(a)\) and \(V=B_{r}(b)\), the open balls. Then \(W\times V\subseteq A\), \(h(W)\subseteq V\), \(h(a)=b\), and Equation (A.524) holds; moreover \(h\) is the only map \(W\longrightarrow V\) whose graph is the zero set of \(F\) in \(W\times V\). Rests on Lemma A.288.
Derives Corollary A.289. \(h(x)\in V\) by Equation (A.531), since \(r/2<r\); and \(h(a)=b\) because \(y\equiv b\) is then a fixed point, unique by Lemma A.288. If \((x,y)\in W\times V\) has \(F(x,y)=0\) then \(y\in\overline{B}_{r}(b)\), so \(y=h(x)\) by the uniqueness clause; conversely \(F(x,h(x))=0\) for every \(x\in W\). That is Equation (A.524). Any other map \(g:W\longrightarrow V\) with the same graph would satisfy \(F(x,g(x))=0\) with \(g(x)\in\overline{B}_{r}(b)\), hence \(g=h\).
∎Continuity, differentiability and the derivative formula
There is \(K\ge0\) with
In particular \(h\) is continuous. Rests on Lemma A.288 and Theorem 7.35.
Derives Lemma A.290. The entries of \(D_{x}F\) are continuous, hence bounded on the compact set \(\overline{B}_{\rho}(a)\times\overline{B}_{r}(b)\) (Theorem 7.24); let \(\mu\) bound \(\norm{D_{x}F}\) there. Exactly as in step (2) of Lemma A.287, the mean value theorem along the segment from \(x'\) to \(x\) — which lies in the convex ball \(\overline{B}_{\rho}(a)\) — gives
Now, using that \(h(x)\) and \(h(x')\) are fixed points of \(T(x,\cdot)\) and \(T(x',\cdot)\),
the last term because \(T(x,y)-T(x',y) = -C\left[F(x,y)-F(x',y)\right]\). Absorbing the first term on the left and inserting Equation (A.534) gives Equation (A.533) with \(K = 2\sqrt{n}\,\mu\,\norm{C}\).
∎\(h\) is differentiable at every \(x\in W\) and its differential is Equation (A.525); the right-hand side of Equation (A.525) is continuous on \(W\), so \(h\) is \(C^{1}\). Rests on Lemma A.290, Theorem 7.68 and Definition 7.67.
Derives Lemma A.291. Fix \(x\in W\), write \(y=h(x)\), and let \(s\in\R^{m}\) be small enough that \(x+s\in W\). Put
by Equation (A.533). Since \(F\) is \(C^{1}\) it is differentiable at \((x,y)\) (Theorem 7.68), and its differential is the pair of blocks Equation (A.522) (Definition 7.70); so by Equation (7.45) there is a function \(\vect{\varepsilon}\) with \(\abs{\vect{\varepsilon}(s,\Delta)} /\left(\abs{s}+\abs{\Delta}\right)\longrightarrow0\) and
both values of \(F\) vanishing by Equation (A.524). Because \(\abs{\Delta}\le K\abs{s}\), the remainder obeys \(\abs{\vect{\varepsilon}}\le(1+K)\abs{s}\cdot \abs{\vect{\varepsilon}}/(\abs{s}+\abs{\Delta})\), so it is \(o\!\left(\abs{s}\right)\). The matrix \(D_{y}F(x,y)\) is invertible by assertion (5) of Lemma A.287, so solving Equation (A.535) for \(\Delta\),
The second term is \(o\!\left(\abs{s}\right)\) as well, the matrix being a fixed one. Hence \(h(x+s)-h(x) - \left(-\left[D_{y}F\right]^{-1}D_{x}F\right)s = o\!\left(\abs{s}\right)\), which by Definition 7.67 says that \(h\) is differentiable at \(x\) with the differential Equation (A.525); the differential being unique, no other candidate exists.
Continuity: the entries of \(\left[D_{y}F\right]^{-1}\) are, by Cramer's rule (Linear Algebra and Representation Theory), quotients of polynomials in the entries of \(D_{y}F\) by the determinant, which does not vanish on \(\overline{B}_{\rho}(a)\times\overline{B}_{r}(b)\); those entries are continuous, and so is \(h\) (Lemma A.290), so the composite \(x\longmapsto -\left[D_{y}F(x,h(x))\right]^{-1} D_{x}F(x,h(x))\) is continuous. By Definition 7.66, \(h\) is \(C^{1}\).
∎Proof of Theorem A.286. Derives Theorem A.286. Corollary A.289 supplies \(W\), \(V\), \(h\), Equation (A.524) and the uniqueness; Lemma A.291 supplies Equation (A.525) and the class \(C^{1}\). It remains to raise the smoothness. Write
If \(F\) is \(C^{k}\) then \(D_{x}F\) and \(D_{y}F\) are \(C^{k-1}\), and so is \(\left[D_{y}F\right]^{-1}\), its entries being quotients of polynomials in \(C^{k-1}\) functions with non-vanishing denominator; hence \(G\) is \(C^{k-1}\). We prove by induction on \(j\) that \(h\) is \(C^{j}\) for \(1\le j\le k\). The case \(j=1\) is Lemma A.291. Suppose \(h\) is \(C^{j}\) with \(j<k\). Then \(x\longmapsto\bigl(x,h(x)\bigr)\) is \(C^{j}\), and \(G\) is \(C^{k-1}\) with \(k-1\ge j\); a composition of a \(C^{j}\) map with a \(C^{j}\) map is \(C^{j}\), because by the chain rule Equation (7.52) each first partial derivative of the composite is a sum of products of first partial derivatives of the two factors, and an induction on \(j\) carries this to order \(j\). So \(Dh\) is \(C^{j}\) by Equation (A.537), that is, \(h\) is \(C^{j+1}\). After \(k-1\) steps \(h\) is \(C^{k}\). If \(F\) is smooth the argument applies for every \(k\), so \(h\) is smooth.
∎The two consequences the treatise uses
Let \(U\subseteq\R^{n}\) be open, \(f:U\longrightarrow\R^{n}\) of class \(C^{k}\) with \(k\ge1\), and \(\alpha\in U\) with \(\det\left[Df(\alpha)\right]\ne0\). Then there are open sets \(U_{0}\subseteq U\) with \(\alpha\in U_{0}\) and \(V_{0}\subseteq\R^{n}\) with \(f(\alpha)\in V_{0}\) such that \(f\) maps \(U_{0}\) bijectively onto \(V_{0}\), the inverse \(g:V_{0}\longrightarrow U_{0}\) is \(C^{k}\), and
Rests on Theorem A.286 and Proposition 7.72.
Derives Corollary A.292. Apply Theorem A.286 to
at the point \(\left(f(\alpha),\alpha\right)\), where \(\Phi=0\) and \(D_{u}\Phi = -Df(\alpha)\) is invertible. It supplies open sets \(V_{0}\ni f(\alpha)\) and \(V\ni\alpha\) and a \(C^{k}\) map \(g:V_{0}\longrightarrow V\) whose graph is the zero set of \(\Phi\) in \(V_{0}\times V\); that is,
Put \(U_{0} = V\cap f^{-1}(V_{0})\), open because \(f\) is continuous, and note \(\alpha\in U_{0}\). If \(u\in U_{0}\) then \((f(u),u)\in V_{0}\times V\), so \(u=g(f(u))\) by Equation (A.539): \(f\) is injective on \(U_{0}\). If \(v\in V_{0}\) then \(g(v)\in V\) and \(f(g(v))=v\in V_{0}\), so \(g(v)\in U_{0}\) and \(v\in f(U_{0})\): \(f\) maps \(U_{0}\) onto \(V_{0}\), with \(g\) as its two-sided inverse. Finally, differentiating \(g\circ f=\id\) on \(U_{0}\) by the chain rule Equation (7.52) gives \(Dg\bigl(f(u)\bigr)Df(u)=\identity\), which is Equation (A.538) at \(v=f(u)\).
∎Let \(A\subseteq\R^{n}\) be open, \(F:A\longrightarrow\R\) of class \(C^{k}\) with \(k\ge1\), and let \(P\in A\) satisfy \(F(P)=c\) and \(\pp F/\pp x^{n}(P)\neq0\). Write \(P=\left(P',P^{n}\right)\) with \(P'\in\R^{n-1}\). Then there are an open set \(W\subseteq\R^{n-1}\) containing \(P'\) and an open interval \(I\) containing \(P^{n}\), with \(W\times I\subseteq A\), and a \(C^{k}\) function \(\varsigma:W\longrightarrow I\) with \(\varsigma(P')=P^{n}\) such that, inside \(W\times I\),
and
Rests on Theorem A.286.
Derives Corollary A.293. Apply Theorem A.286 with \(m=n-1\), target dimension \(1\), and \(G(x',x^{n}) = F(x',x^{n})-c\) at the point \(\left(P',P^{n}\right)\): \(G=0\) there, and \(D_{y}G\) is the \(1\times1\) matrix \(\pp F/\pp x^{n}(P)\ne0\), whose determinant is itself. The theorem returns \(W\), an open ball \(V\subseteq\R\) — that is, an open interval \(I\) — and \(\varsigma\) of class \(C^{k}\) with Equation (A.524), which is Equation (A.540). Equation (A.525) reads, for \(1\times1\) blocks, \(D\varsigma = -\left(\pp F/\pp x^{n}\right)^{-1} \left(\pp F/\pp x^{1},\ldots,\pp F/\pp x^{n-1}\right)\), which is Equation (A.541).
∎Take \(n=3\) and \(F(x,y,z)=x^{2}+y^{2}+z^{2}\), \(c=R^{2}>0\), the level set of Example 13.61. At a point of the sphere with \(z\neq0\) one has \(\pp F/\pp z = 2z\ne0\), so Corollary A.293 applies with the whole of \(\R^{3}\) as \(A\); the function it produces is the one the chapter writes down by hand, \(\varsigma(x,y)=+\sqrt{R^{2}-x^{2}-y^{2}}\) on the northern hemisphere, and Equation (A.541) returns \(\pp\varsigma/\pp x = -x/z\) and \(\pp\varsigma/\pp y = -y/z\), which is what differentiating the explicit square root gives. The point of the corollary is that the same conclusion holds when no explicit solution can be written: at a point of the sphere with \(z=0\) one of the other two partial derivatives is non-zero, and the corollary — applied after relabelling the coordinates — produces a graph there too. If all three partial derivatives could vanish on the level set, no such chart would be guaranteed; that is exactly the regularity hypothesis of Theorem 13.59, and on the sphere it fails only at the origin, which is not on the level set. Rests on Corollary A.293 and Example 13.61.
Reading the result
This section imports no theorem that the treatise does not prove. In particular it does not use a fixed-point theorem: the Banach contraction principle is not in this book, and rather than quote it the iteration Equation (A.530) is run explicitly and its convergence read off from the geometric estimate Equation (A.532) and the Cauchy criterion Theorem 7.8 — the same device, and the same estimate, by which Theorem 9.8 constructs the solution of an initial value problem. The only other analytic inputs are the mean value theorem Theorem 7.35, used twice to convert a bound on a derivative into a Lipschitz bound; the theorem that a \(C^{1}\) map is differentiable (Theorem 7.68), used once to expand \(F\) to first order in Equation (A.535); and the chain rule Proposition 7.72. From linear algebra, only the invertibility criterion \(\det\neq0\) and Cramer's rule are used.
The invertibility of \(D_{y}F(a,b)\) cannot be dropped. On \(\R\times\R\), \(F(x,y)=y^{2}-x\) has \(F(0,0)=0\) and \(\pp F/\pp y(0,0)=0\); the zero set near the origin is the parabola \(x=y^{2}\), which is not the graph of any function of \(x\) — two solutions for \(x>0\), none for \(x<0\). The conclusion is also irreducibly local: \(F(x,y)=x^{2}+y^{2}-1\) on \(\R\times\R\) satisfies the hypotheses at \((0,1)\), and the implicit function \(\varsigma(x)=\sqrt{1-x^{2}}\) exists only for \(\abs{x}<1\) and cannot be continued past \(x=\pm1\), where \(\pp F/\pp y\) vanishes. Finally, the smoothness of \(h\) can be no better than that of \(F\): taking \(m=n=1\) and \(F(x,y)=y-\phi(x)\) with \(\phi\) of class \(C^{k}\) but not \(C^{k+1}\) makes \(h=\phi\). Rests on Theorem A.286.
The Implicit Function Theorem discharges the derivation owed at Remark 13.60 of Differentiable Manifolds, Tensors, and Curvature. The statement used there is Corollary A.293: in the proof of the regular value theorem Theorem 13.59 a chart is chosen in which \(\pp F/\pp x^{n}\ne0\) at the point, and Equation (A.540) is exactly the assertion Equation (13.146) that the level set meets the coordinate box in the graph of a smooth \(h\) — from which the chapter reads off the chart \(u\longmapsto(u,h(u))\) of the hypersurface. The regular value theorem in turn underlies Example 13.61 and Definition 13.58 and every later construction of a submanifold by one equation. The inverse function theorem Corollary A.292, proved here as a companion, is what licenses a change of chart whenever a Jacobian determinant is non-zero, and is used in that role in The General Stokes Theorem for Differential Forms.
The General Stokes Theorem for Differential Forms
This appendix proves the theorem that Equation (13.231) of Differentiable Manifolds, Tensors, and Curvature is an instance of: for a compact oriented \(n\)-manifold \(M\) with boundary and a smooth \((n-1)\)-form \(\omega\) on it,
It is the single statement of which Green's theorem (Theorem 7.99), the classical Stokes theorem (Theorem 7.100), Gauss's theorem (Theorem 7.101) and the antisymmetric-tensor form Equation (13.231) are the low-dimensional faces, and Remark 7.102 names it as the formulation the treatise was still owed.
There is no circularity to create. Real Analysis proves its three classical theorems directly from the fundamental theorem of calculus over simple regions and does not lean on the present section; the present section does not use them. What it does use from Real Analysis is the multiple integral (Definition 7.93), the three properties of it collected as declared inputs in Remark 7.96 — integrability of a continuous function, additivity, and reduction to an iterated integral — and the fundamental theorem of calculus (Theorem 7.43). From Differentiable Manifolds, Tensors, and Curvature it uses the whole exterior calculus: Definition 13.98, Proposition 13.100, Definition 13.102, Definition 13.103 and Equation (13.246). From this appendix it uses the pullback of a \(k\)-form (Definition A.70 and Lemma A.71) and the inverse function theorem (Corollary A.292). One input is quoted, the change-of-variables formula for multiple integrals; it is stated precisely as Theorem A.303 and discussed in Remark A.316.
Manifolds with boundary
The closed half-space is
carrying the subspace topology of \(\R^{n}\) (Definition 6.1). A map defined on a subset \(S\subseteq\mathbb{H}^{n}\) that is open in \(\mathbb{H}^{n}\) is called smooth if it extends to a smooth map on some open subset of \(\R^{n}\) containing \(S\); all partial derivatives are then defined at points of \(\pp\mathbb{H}^{n}\) as well, and take there the values of the one-sided limits, so that every identity of the differential calculus of Section 7.10.2 continues to hold on \(S\). Rests on Definitions 6.1 and 7.66.
An \(n\)-manifold with boundary is defined exactly as in Definition 13.48, with one change: the charts \(\varphi:U\longrightarrow\varphi(U)\) of the atlas (Definition 13.44) are homeomorphisms onto subsets of \(\mathbb{H}^{n}\) open in \(\mathbb{H}^{n}\), rather than onto open subsets of \(\R^{n}\), and the transition maps are smooth in the sense of Definition A.298. A point \(P\in M\) is a boundary point if \(\varphi(P)\in\pp\mathbb{H}^{n}\) for some chart around it, and an interior point otherwise; the set of boundary points is written \(\pp M\). Rests on Definitions 13.44, 13.48 and A.298.
No point is a boundary point in one chart and an interior point in another. Moreover \(\pp M\), with the charts obtained by dropping the last coordinate of a boundary chart, is an \((n-1)\)-manifold without boundary. Rests on Definition A.299 and Corollary A.292.
Derives Lemma A.300. Suppose \(P\) had charts \(\varphi\) on \(U\) and \(\psi\) on \(V\) with \(z_{0}=\varphi(P)\) satisfying \(z_{0}^{n}<0\) and \(\psi(P)\in\pp\mathbb{H}^{n}\). Because \(z_{0}^{n}<0\), the set \(\varphi(U\cap V)\) contains an open subset \(O\subseteq\R^{n}\) around \(z_{0}\). Write \(\tau=\psi\circ\varphi^{-1}\) on \(O\) and let \(\varsigma\) be a smooth extension to an open subset of \(\R^{n}\), as in Definition A.298, of the transition map \(\varphi\circ\psi^{-1}\) near \(\psi(P)\). Shrinking \(O\) so that \(\tau(O)\) lies in the domain of \(\varphi\circ\psi^{-1}\), we have \(\varsigma\circ\tau=\id\) on \(O\), so the chain rule Equation (7.52) gives \(D\varsigma\bigl(\tau(z_{0})\bigr)D\tau(z_{0})=\identity\) and \(\det\left[D\tau(z_{0})\right]\ne0\). By the inverse function theorem Corollary A.292, \(\tau\) maps some neighbourhood of \(z_{0}\) onto an open subset of \(\R^{n}\) containing \(\tau(z_{0})=\psi(P)\in\pp\mathbb{H}^{n}\). Every open subset of \(\R^{n}\) containing a point of \(\pp\mathbb{H}^{n}\) contains points with \(x^{n}>0\), which do not lie in \(\mathbb{H}^{n}\) — contradicting that \(\psi\) takes values in \(\mathbb{H}^{n}\).
For the second claim, let \(\varphi=(x^{1},\ldots,x^{n})\) be a chart with \(\varphi(U)\cap\pp\mathbb{H}^{n}\neq\varnothing\). By the first claim, \(\varphi(U\cap\pp M) = \varphi(U)\cap\pp\mathbb{H}^{n}\), which is an open subset of \(\R^{n-1}=\pp\mathbb{H}^{n}\), and \(\varphi'=(x^{1},\ldots,x^{n-1})\) is a homeomorphism of \(U\cap\pp M\) onto it. Two such charts are smoothly related, because a transition map of \(M\) carries \(\pp\mathbb{H}^{n}\) into \(\pp\mathbb{H}^{n}\) — by the first claim again — so its first \(n-1\) components restricted to \(x^{n}=0\) are the transition map of the boundary charts, and they are smooth. Hence Definition 13.44 is met with model space \(\R^{n-1}\).
∎Orientation and the integral of an $n$-form
Let \(\varphi=(x^{1},\ldots,x^{n})\) and \(\psi=(y^{1},\ldots,y^{n})\) be two charts on a common domain. Then
Consequently, if an \(n\)-form reads \(f\,\dd x^{1}\wedge\cdots\wedge\dd x^{n}\) and \(g\,\dd y^{1}\wedge\cdots\wedge\dd y^{n}\) in the two charts, then \(g = f\,\det\left[\pp x/\pp y\right]\). Rests on Definition 13.102, Equation (13.241) and Proposition 13.100.
Derives Lemma A.301. By the chain rule, \(\dd x^{i} = \left(\pp x^{i}/\pp y^{j}\right)\dd y^{j}\). Substituting into the left-hand side and expanding the wedge, every term in which two of the \(y\)-indices coincide vanishes (Equation (13.241)), so only the \(n!\) terms indexed by a permutation \(\rho\) of \((1,\ldots,n)\) survive, and reordering \(\dd y^{\rho(1)}\wedge\cdots\wedge\dd y^{\rho(n)}\) into increasing order costs \(\sgn\rho\):
and the bracket is the Leibniz expansion of the determinant (Linear Algebra and Representation Theory). The last sentence follows because \(\Lambda^{n}\) is one-dimensional (Equation (13.235)).
∎An atlas of \(M\) is oriented if every transition map between two of its charts has everywhere positive Jacobian determinant, \(\det\left[\pp x/\pp y\right]>0\). \(M\) is orientable if it admits such an atlas, and an oriented manifold is a manifold together with a maximal oriented atlas; a chart belongs to it, and is called positively oriented, if its transitions with every chart of the atlas have positive Jacobian determinant. Rests on Definition 13.44 and Lemma A.301.
Let \(\tau: O\longrightarrow O'\) be a diffeomorphism between open subsets of \(\R^{n}\) and let \(u\) be continuous with compact support in \(O'\). Then \(u\circ\tau\,\abs{\det D\tau}\) has compact support in \(O\) and
both integrals being multiple integrals in the sense of Definition 7.93. Rests on Definition 7.93 and Remark 7.96.
Let \(M\) be an oriented \(n\)-manifold with boundary, let \(\varphi=(x^{1},\ldots,x^{n})\) be a positively oriented chart on \(U\), and let \(\alpha\) be a continuous \(n\)-form on \(M\) whose support is a compact subset of \(U\). Writing \(\alpha = f\,\dd x^{1}\wedge\cdots\wedge\dd x^{n}\) on \(U\), set
Rests on Definition A.302, Definition 7.93 and Proposition 13.100.
The value Equation (A.546) does not depend on which positively oriented chart containing the support is used. Rests on Definition A.304, Theorem A.303 and Lemma A.301.
Derives Lemma A.305. Let \(\psi=(y^{1},\ldots,y^{n})\) be a second positively oriented chart containing \(\operatorname{supp}\alpha\), and let \(\tau=\varphi\circ\psi^{-1}\), a bijection between the images of the common domain.
Reduction to open sets. By Lemma A.300, \(\tau\) carries interior points to interior points and boundary points to boundary points, so it restricts to a diffeomorphism between the two open subsets of \(\R^{n}\) obtained by deleting the face \(x^{n}=0\); and deleting that face changes neither multiple integral, because a bounded portion of a hyperplane contributes nothing to an upper–lower Darboux gap — the argument of Remark 7.96 for the graph of a continuous function, applied to the constant function \(0\). Both sides of the claimed equality may therefore be computed over the open sets, where Theorem A.303 applies as stated.
By Lemma A.301 the component functions are related by \(g = f\,\det\left[\pp x/\pp y\right] = f\,\det D\tau\), and \(\det D\tau>0\) by Definition A.302, so \(\abs{\det D\tau}=\det D\tau\). Then Equation (A.545) with \(u=f\circ\varphi^{-1}\) gives
since \(f\circ\varphi^{-1}\circ\tau = f\circ\psi^{-1}\).
∎Bump functions and partitions of unity
The construction below is elementary and is carried out in full, because the treatise does not have it: nothing here is quoted.
The function
is smooth on \(\R\). Consequently, for \(0<a<b\),
is smooth, takes values in \([0,1]\), equals \(1\) for \(t\le a\) and \(0\) for \(t\ge b\); and for \(0<r_{1}<r_{2}\) the function
is smooth on \(\R^{n}\), equals \(1\) on \(\overline{B}_{r_{1}}(0)\), and vanishes off \(B_{r_{2}}(0)\). Rests on Definition 7.53, Theorem 7.37 and Proposition 7.72.
Derives Lemma A.306. Smoothness of \(\lambda\). On \(t>0\) an induction gives \(\lambda^{(k)}(t)=P_{k}(1/t)\,\ee^{-1/t}\) with \(P_{k}\) a polynomial: this holds for \(k=0\) with \(P_{0}=1\), and differentiating \(P_{k}(1/t)\ee^{-1/t}\) gives \(\left[-t^{-2}P_{k}'(1/t)+t^{-2}P_{k}(1/t)\right]\ee^{-1/t}\), of the same shape with \(P_{k+1}(s)=s^{2}\left(P_{k}(s)-P_{k}'(s)\right)\). On \(t<0\) every derivative is \(0\). At \(t=0\) we show by induction that \(\lambda^{(k)}(0)\) exists and is \(0\). Assume it for \(k\); the left difference quotient is \(0\), and the right one is \(P_{k}(1/t)\ee^{-1/t}/t\). Put \(s=1/t\to+\infty\); the quotient is \(s\,P_{k}(s)\,\ee^{-s}\), and for every \(N\) the exponential series (Equation (7.31)) gives \(\ee^{s}>s^{N}/N!\), i.e.\ \(\ee^{-s}<N!\,s^{-N}\), so choosing \(N\) larger than \(1+\deg P_{k}\) makes \(s\,P_{k}(s)\ee^{-s}\to0\). Hence \(\lambda^{(k+1)}(0)=0\). Each \(\lambda^{(k)}\) is then continuous at \(0\) by the same estimate, so \(\lambda\) is smooth.
\(\beta\). The denominator of Equation (A.548) never vanishes: for \(t<b\) the first term is positive, and for \(t\ge b>a\) the second is. So \(\beta\) is smooth, and it lies in \([0,1]\) because both terms are non-negative. For \(t\le a\) the second term vanishes and \(\beta=1\); for \(t\ge b\) the first vanishes and \(\beta=0\).
\(\chi\). \(x\longmapsto\abs{x}^{2}\) is a polynomial, hence smooth, so \(\chi\) is smooth by the chain rule Proposition 7.72; \(\abs{x}\le r_{1}\) gives \(\abs{x}^{2}\le a\) and \(\chi=1\), and \(\abs{x}\ge r_{2}\) gives \(\chi=0\).
∎Let \(M\) be a compact \(n\)-manifold with boundary and \(\set{U_{a}}\) an open cover of \(M\). Then there are finitely many smooth functions \(\rho_{1},\ldots,\rho_{N}:M\longrightarrow[0,1]\) such that each \(\operatorname{supp}\rho_{i}\) is a compact set contained in the domain of a single chart which is itself contained in some \(U_{a}\), and
Rests on Lemma A.306, Definition 6.9 and Definition A.299.
Derives Lemma A.307. Let \(P\in M\). Choose a chart \(\varphi_{P}\) on a domain \(V_{P}\ni P\) with \(V_{P}\) contained in some \(U_{a}\) and \(\varphi_{P}(P)=0\) (translate the chart). Since \(\varphi_{P}(V_{P})\) is open in \(\mathbb{H}^{n}\) there is \(r_{P}>0\) with \(\overline{B}_{r_{P}}(0)\cap\mathbb{H}^{n}\subseteq\varphi_{P}(V_{P})\). Let \(\chi_{P}\) be the function Equation (A.549) with \(r_{1}=r_{P}/2\) and \(r_{2}=r_{P}\), and set
This is smooth on \(M\): it is smooth on \(V_{P}\), and it vanishes identically on the open set \(M\setminus\varphi_{P}^{-1}\left(\overline{B}_{r_{P}}(0)\cap \mathbb{H}^{n}\right)\), whose complement in \(M\) is the compact set \(\varphi_{P}^{-1}\left(\overline{B}_{r_{P}}(0)\cap\mathbb{H}^{n}\right) \subseteq V_{P}\); the two open sets cover \(M\) and the definitions agree on the overlap. The same compact set contains \(\operatorname{supp}\widetilde{\chi}_{P}\).
Put \(W_{P}=\varphi_{P}^{-1}\left(B_{r_{P}/2}(0)\cap\mathbb{H}^{n}\right)\), an open neighbourhood of \(P\) on which \(\widetilde{\chi}_{P}=1\). The family \(\set{W_{P}}\) is an open cover of the compact \(M\) (Definition 6.9), so finitely many \(W_{P_{1}},\ldots,W_{P_{N}}\) cover it. Then \(S=\sum_{i}\widetilde{\chi}_{P_{i}}\) is smooth and satisfies \(S\ge1\) everywhere, since each point lies in some \(W_{P_{i}}\) where the corresponding term is \(1\) and no term is negative. Setting \(\rho_{i}=\widetilde{\chi}_{P_{i}}/S\) gives smooth functions with the stated supports and Equation (A.550).
∎Let \(M\) be a compact oriented \(n\)-manifold with boundary and \(\alpha\) a continuous \(n\)-form on \(M\). Choose a partition of unity \(\set{\rho_{i}}\) subordinate to a cover by domains of positively oriented charts (Lemma A.307) and set
each summand being a chart integral (Definition A.304). The value does not depend on the partition: if \(\set{\sigma_{j}}\) is another one, then by linearity of the multiple integral and Lemma A.305 — each \(\rho_{i}\sigma_{j}\alpha\) being supported in one chart —
Let \(M\) be oriented and \(\varphi=(x^{1},\ldots,x^{n})\) a positively oriented boundary chart. The induced orientation of \(\pp M\) is the one in which the restricted chart \(\varphi'=(x^{1},\ldots,x^{n-1})\) of Lemma A.300 is positively oriented when \(n\) is odd and negatively oriented when \(n\) is even; equivalently, the chart \(\left((-1)^{n-1}x^{1},x^{2},\ldots,x^{n-1}\right)\) is positively oriented. This is consistent: the transition map between two boundary charts has, at \(x^{n}=0\), Jacobian determinant equal to \(\left(\pp x^{n}/\pp y^{n}\right)\) times that of the restricted transition, and \(\pp x^{n}/\pp y^{n}>0\) because both charts map into \(\mathbb{H}^{n}\) and both send \(\pp M\) to \(x^{n}=0\); hence the restricted transitions inherit positive determinants, and multiplying every one of them by the same sign \((-1)^{n-1}\) leaves them positive. Rests on Definition A.302 and Lemma A.300.
The sign \((-1)^{n-1}\) is not a convention chosen to make Equation (A.542) come out right; it is the outward-normal-first rule counted honestly. In a boundary chart the outward direction is that of increasing \(x^{n}\), since \(\mathbb{H}^{n}=\set{x^{n}\le0}\), so the rule declares \((v_{1},\ldots,v_{n-1})\) positively oriented in \(T_{\pp M}\) exactly when \((\pp_{n},v_{1},\ldots,v_{n-1})\) is positively oriented in \(T_{M}\). Moving \(\pp_{n}\) from the front of \((\pp_{n},\pp_{1},\ldots,\pp_{n-1})\) to the back costs \(n-1\) transpositions, so \((\pp_{1},\ldots,\pp_{n-1})\) is positively oriented precisely when \((-1)^{n-1}=+1\), which is Definition A.309. For \(n=3\) — a surface in space, the case of Equation (13.231) — the sign is \(+1\) and the rule reduces to the familiar right-hand relation between the normal of \(S\) and the sense of circulation on \(\pp S\).
Statement and the half-space computation
Let \(M\) be a compact oriented \(n\)-manifold with boundary, \(n\ge1\), let \(\pp M\) carry the induced orientation, and let \(\omega\) be a smooth \((n-1)\)-form on \(M\). Then
where \(\iota:\pp M\longrightarrow M\) is the inclusion. If \(\pp M=\varnothing\) the right-hand side is \(0\). Rests on Definitions 13.103, A.308 and A.309.
Let \(\omega\) be a smooth \((n-1)\)-form on \(\mathbb{H}^{n}\) with compact support. Write
the hat marking the omitted factor; this is possible, with smooth \(f_{i}\), because those \(n\) products form a basis of \(\Lambda^{n-1}\) (Equation (13.235)). Then
the last integral taken with the induced orientation Definition A.309. If the support of \(\omega\) lies in the open half-space \(\set{x^{n}<0}\), both sides vanish. Rests on Definition 13.103, Theorem 7.43 and Remark 7.96.
Derives Lemma A.312. The exterior derivative. Write \(\theta_{i}=\dd x^{1}\wedge\cdots\wedge\widehat{\dd x^{i}} \wedge\cdots\wedge\dd x^{n}\). Carrying \(\dd x^{i}\) leftwards past the \(i-1\) factors \(\dd x^{1},\ldots,\dd x^{i-1}\) using Equation (13.241),
By the Leibniz rule Equation (13.246) and \(\dd(\dd x^{j})=0\),
and \(\dd x^{j}\wedge\theta_{i}=0\) unless \(j=i\), since otherwise \(\dd x^{j}\) is repeated. With Equation (A.555),
The integral. Choose \(R>0\) so large that \(\operatorname{supp}\omega\) is contained in the box \(Q=[-R,R]^{n-1}\times[-R,0]\), and note that each \(f_{i}\) vanishes identically outside it. By Equation (A.556), Definition A.304 and additivity, the left-hand side of Equation (A.554) is \(\sum_{i}\int_{Q}\pp_{i}f_{i}\). Reduce each term to an iterated integral, integrating first in \(x^{i}\) (Remark 7.96) and applying the fundamental theorem of calculus Theorem 7.43. For \(i<n\),
both endpoint values lying outside \(Q\); hence those \(n-1\) terms vanish. For \(i=n\) the \(x^{n}\)-integration runs only up to \(0\):
which gives the first equality in Equation (A.554). If \(\operatorname{supp}\omega\subseteq\set{x^{n}<0}\) then \(f_{n}(\cdot,0)=0\) and the expression vanishes.
The boundary integral. The inclusion is \(\iota(x^{1},\ldots,x^{n-1})=(x^{1},\ldots,x^{n-1},0)\), so by Equation (A.190) \(\iota^{\ast}\dd x^{j}=\dd x^{j}\) for \(j<n\) while \(\iota^{\ast}\dd x^{n}=\dd\left(x^{n}\circ\iota\right)=0\). By Equation (A.187) the pullback distributes over the wedge, so every \(\theta_{i}\) with \(i<n\) — each of which still contains the factor \(\dd x^{n}\) — pulls back to zero, and
The induced orientation is \((-1)^{n-1}\) times the one in which \((x^{1},\ldots,x^{n-1})\) is positively oriented (Definition A.309), and reversing an orientation reverses the sign of the integral Equation (A.546): replacing \(x^{1}\) by \(-x^{1}\) multiplies the component function by \(-1\) (Lemma A.301, the Jacobian determinant of the reflection being \(-1\)), while the multiple integral over the reflected region is unchanged by Equation (A.545), whose factor \(\abs{\det D\tau}\) is \(1\) for a reflection. Hence
which is the second equality in Equation (A.554).
∎Proof of the theorem
Proof of Theorem A.311. Derives Theorem A.311. Cover \(M\) by domains of positively oriented charts and let \(\set{\rho_{1},\ldots,\rho_{N}}\) be a subordinate partition of unity (Lemma A.307). Since \(\sum_{i}\rho_{i}=1\) is constant, \(\sum_{i}\dd\rho_{i}=\dd(1)=0\), so by the Leibniz rule Equation (13.246)
Each \(\rho_{i}\omega\) is a smooth \((n-1)\)-form whose support is a compact subset of the domain \(U_{i}\) of a positively oriented chart \(\varphi_{i}\); transporting it by \(\varphi_{i}\) and extending by zero gives a smooth compactly supported \((n-1)\)-form on the whole of \(\mathbb{H}^{n}\), smooth because the two definitions — the transported one on the open set \(\varphi_{i}(U_{i})\), and zero on the complement of the compact support — agree on the overlap of two open sets covering \(\mathbb{H}^{n}\). The transport carries the two sides of Equation (A.559) to the two sides of Equation (A.554): the chart integral Equation (A.546) is by definition the multiple integral of the chart representative, and pullback commutes with \(\dd\) (Equation (A.188)), so the chart representative of \(\dd(\rho_{i}\omega)\) is the exterior derivative of the chart representative of \(\rho_{i}\omega\). Hence Lemma A.312 applies and gives
where the right-hand side is \(0\) when \(U_{i}\) misses \(\pp M\), that being the last clause of Lemma A.312. Summing Equation (A.559) over \(i\), using Equation (A.558) on the left and, on the right, that \(\set{\iota^{\ast}\rho_{i}}\) restricts to a partition of unity on \(\pp M\) — \(\sum_{i}\rho_{i}=1\) there too, and each \(\iota^{\ast}\rho_{i}\) is supported in a boundary chart — gives Equation (A.552) by Definition A.308. If \(\pp M=\varnothing\), every chart misses the boundary and every term on the right vanishes.
∎The instance stated in the chapter
Let \(S\subset E_{3}\) be a compact oriented two-dimensional submanifold with boundary \(\pp S\), carrying the induced orientation, and let \(H_{i}\) be a smooth covector field defined near \(S\), with rotor tensor
Then, with \(\dd f^{ij}\) the antisymmetric surface element of \(S\),
which is Equation (13.231). Rests on Theorem A.311, Equation (13.232) and Definition 7.94.
Derives Corollary A.313. Let \(H=H_{i}\,\dd x^{i}\) be the one-form with the given components. By Definition 13.103,
the middle step by antisymmetrizing the coefficient against \(\dd x^{j}\wedge\dd x^{i}=-\dd x^{i}\wedge\dd x^{j}\) (Equation (13.241)). Theorem A.311 applied to \(M=S\) and \(\omega=\iota_{S}^{\ast}H\), the restriction of \(H\) to \(S\), gives
where pullback and \(\dd\) were exchanged by Equation (A.188).
It remains to recognize the two sides. Let \(\vect{r}:D\longrightarrow S\) be a positively oriented parametrization of a piece of \(S\), with parameters \((u,v)\). The two-form \(\dd x^{i}\wedge\dd x^{j}\) has components \(\delta^{i}_{\ a}\delta^{j}_{\ b}-\delta^{i}_{\ b}\delta^{j}_{\ a}\) (Definition 13.102), so Equation (A.186) gives
That bracket is the antisymmetric surface element. Indeed, Remark 13.97 writes \(\dd f^{ij}=\varepsilon^{ijk}\dd f_{k}\) with \(\dd f_{k}\) the oriented surface element Equation (7.108) of Definition 7.94, whose components are \(\dd f_{k}=\varepsilon_{klm}\,\pp_{u}r^{l}\,\pp_{v}r^{m}\,\dd u\,\dd v\); contracting the two Levi-Civita symbols (Linear Algebra and Representation Theory),
which is the bracket of Equation (A.564) times \(\dd u\,\dd v\). Hence, by Definition A.304 and Equation (A.562),
and, parametrizing \(\pp S\) by arc parameter \(t\), \(\iota_{\pp S}^{\ast}H = H_{i}(\vect{r}(t))\,\dv{r^{i}}{t}\,\dd t\), whose integral is the line integral Equation (7.106) of \(H_{i}\dd x^{i}\) over \(\pp S\). Substituting both into Equation (A.563) gives Equation (A.561).
∎Remark 13.97 records that Equation (13.231) holds for the rotor tensor Equation (13.232) and is false if \(H_{i}\) is instead read as the dual pseudo-vector \(\tfrac12\varepsilon_{ijk}H_{jk}\) of Equation (13.224). The derivation above respects that, and shows where the hypothesis enters: the only step in which the relation between \(H_{ij}\) and \(H_{i}\) is used is Equation (A.562), which asserts that the two-form \(\tfrac12 H_{ij}\dd x^{i}\wedge\dd x^{j}\) is exact, with the one-form \(H=H_{i}\dd x^{i}\) as its primitive. Under the dual reading there is no such relation: \(\tfrac12 H_{ij}\dd x^{i}\wedge\dd x^{j}\) would be an arbitrary two-form and \(H_{i}\dd x^{i}\) an unrelated one-form, so Theorem A.311 says nothing connecting their integrals, and the left-hand side of Equation (13.231) would be a flux and the right-hand side a circulation of the same vector field — two quantities with no reason to agree. Writing \(\dd f^{ij}=\varepsilon^{ijk}\dd f_{k}\) turns Equation (A.561) into \(\int_{S}\left(\nabla\times\vect{H}\right)\cdot\dd\vect{f} =\oint_{\pp S}\vect{H}\cdot\dd\vect{l}\), the classical statement Theorem 7.100 — which Real Analysis proves directly and independently, so the two derivations check each other rather than one resting on the other.
The theorem is an identity between integrals of forms and is unit-homogeneous by construction: if \(H_{i}\) carries an SI unit \(\left[H\right]\) then \(H_{i}\dd x^{i}\) carries \(\left[H\right]\cdot\mathrm{m}\), \(H_{ij}\) carries \(\left[H\right]/\mathrm{m}\), and \(\dd f^{ij}\) carries \(\mathrm{m}^{2}\), so both sides of Equation (A.561) carry \(\left[H\right]\cdot\mathrm{m}\). The physical instance the treatise meets first is the magnetic vector potential: with \(H_{i}=A_{i}\) in \(\mathrm{T}\,\mathrm{m}\) and \(H_{ij}=\pp_{i}A_{j}-\pp_{j}A_{i}\) in \(\mathrm{T}\), the identity reads
both sides in \(\mathrm{Wb}=\mathrm{T}\,\mathrm{m}^{2}\): the magnetic flux through a surface equals the circulation of the potential round its edge.
What is quoted, and what is not
Exactly one theorem is assumed without proof: Theorem A.303, the change-of-variables formula for multiple integrals. It is used once, in Lemma A.305, to show that the chart integral Equation (A.546) does not depend on the chart; every later step is built on that one fact about the integral. This treatise does not prove it, and the omission is not new here: Remark 7.102 of Real Analysis already records the same formula as a quoted input, needed there for the independence of a flux integral from its parametrization, alongside the three properties of the multiple integral collected in Remark 7.96. A citation for it is owed and no entry of the bibliography currently carries it, so it is stated here rather than referenced. Nothing else is quoted: the partition of unity, which is the input a reader would most expect to be imported, is constructed outright in Lemmas A.306 and A.307, and the topology of the boundary is settled in Lemma A.300 from the inverse function theorem proved in The Implicit Function Theorem.
Compactness enters once, in Lemma A.307, to make the partition of unity finite — so that the sum Equation (A.558) has finitely many terms and no convergence question arises. It may be traded for the weaker hypothesis that \(\omega\) have compact support, at the cost of building a locally finite partition of unity. Orientability enters twice, both times essentially: in Lemma A.305, where \(\det D\tau>0\) is what removes the absolute value from Equation (A.545) and lets the chart integrals agree rather than differ in sign; and in Definition A.309, where the induced orientation is what makes the sign \((-1)^{n-1}\) produced by the half-space computation cancel against the sign in Equation (A.557). On a non-orientable manifold no consistent \(\int_{M}\) of an \(n\)-form exists at all, and Equation (A.552) has no meaning.
The General Stokes Theorem for Differential Forms discharges the derivation owed at Equation (13.231) of Differentiable Manifolds, Tensors, and Curvature, and with it the general formulation that Remark 7.102 of Real Analysis names as the common source of Theorems 7.99, 7.100 and 7.101. The chapter's statement is recovered as Corollary A.313 under the hypothesis that Remark 13.97 makes explicit, and the general theorem Theorem A.311 is what Section 13.6 was building towards: the exterior derivative and the boundary operator are adjoint, so that \(\dd^{2}=0\) (Equation (13.245)) and \(\pp\pp=\varnothing\) are the same statement seen from the two sides of Equation (A.552).
The Poincaré–Birkhoff–Witt Theorem
This appendix proves Theorem 14.7 of Lie Groups, Lie Algebras, and Fibre Bundles: the ordered monomials in a basis of a Lie algebra \(\mathfrak{g}\) form a vector-space basis of its universal enveloping algebra \(U(\mathfrak{g})\) of Definition 14.5. Only the defining relations Equation (14.12) and the Jacobi identity are assumed. That the ordered monomials span \(U(\mathfrak{g})\) is a reordering argument and is done first (The filtration, and spanning); the content of the theorem is their linear independence, and the whole of it lies in the construction of an action of \(U(\mathfrak{g})\) on the polynomial algebra in \(\dim\mathfrak{g}\) commuting variables (The action on the polynomial algebra). The Jacobi identity is exactly what makes that action well defined: it is the statement that the two ways of reordering a triple product agree. The consequences drawn in Consequences: the symbol and the symmetrization map — the embedding of \(\mathfrak{g}\) in \(U(\mathfrak{g})\), and the symmetrization isomorphism between \(U(\mathfrak{g})\) and the polynomial algebra on \(\mathfrak{g}\) — are what Racah's Theorem on the Number of Casimir Operators uses to count the Casimir operators.
Throughout, \(\mathbb{K}\) is the field of scalars (\(\R\) or \(\C\) in every application in this book; the symbol \(k\) is reserved for the invariant tensors of Definition 14.14), \(\mathfrak{g}\) is a finite-dimensional Lie algebra over \(\mathbb{K}\) with basis \(T_{1},\ldots,T_{M}\), \(M = \dim\mathfrak{g}\), and structure constants fixed as in Equation (14.11),
with the summation convention in force. Nothing below uses the characteristic of \(\mathbb{K}\) except where the symmetrization map is introduced, which needs \(\Q \subseteq \mathbb{K}\) and is flagged there.
Statement
The ordered monomials
form a basis of \(U(\mathfrak{g})\) as a \(\mathbb{K}\)-vector space. In particular the linear map \(\mathfrak{g} \to U(\mathfrak{g})\) is injective, so that \(\mathfrak{g}\) may be regarded as a subspace of \(U(\mathfrak{g})\), and no relation holds among the generators beyond those forced by Equation (14.12). Rests on Definition 14.5 and Equation (14.12).
The filtration, and spanning
Let \(U_{n} \subseteq U(\mathfrak{g})\) be the span of the products \(T_{a_{1}}T_{a_{2}}\cdots T_{a_{m}}\) with \(m \le n\), the empty product being the unit. Then
the last equality because the \(T_{a}\) generate \(U(\mathfrak{g})\) as an algebra with unit. Rests on Definition 14.5.
\(U_{n}\) is spanned by the ordered monomials Equation (A.568) of total degree \(n_{1}+\cdots+n_{M} \le n\). Consequently the ordered monomials span \(U(\mathfrak{g})\). Rests on Definition A.320 and Equation (14.12).
Derives Lemma A.321. Induct on \(n\). For \(n = 0\) there is nothing to prove. Let \(n \ge 1\) and assume the claim for \(n-1\). It suffices to treat a single word \(w = T_{a_{1}}T_{a_{2}}\cdots T_{a_{n}}\) of length exactly \(n\), and we induct on its number of inversions, the number of pairs \(i < j\) with \(a_{i} > a_{j}\). If \(w\) has no inversion its indices are already nondecreasing and \(w\) is an ordered monomial. If it has at least one, some adjacent pair is out of order, \(a_{i} > a_{i+1}\); otherwise the sequence would be nondecreasing. Apply Equation (14.12) to that pair alone:
The first word on the right has length \(n\) and exactly one inversion fewer, so it is handled by the inner induction; the second has length \(n-1\) and lies in \(U_{n-1}\), handled by the outer induction. Hence \(w\) is a combination of ordered monomials of degree at most \(n\).
∎Everything that follows exists to show that this spanning set is free. The danger is real and must be stated plainly: nothing seen so far excludes the possibility that repeated use of Equation (A.570) along two different routes returns two different combinations of ordered monomials, which would force a relation among them and could in principle collapse \(U(\mathfrak{g})\) — in the extreme case to zero. What rules this out is the existence of one representation in which the ordered monomials visibly act independently.
The action on the polynomial algebra
Let \(P = \mathbb{K}[x_{1},\ldots,x_{M}]\) be the polynomial algebra in \(M\) commuting variables. A multi-index is a nondecreasing finite sequence \(A = (a_{1} \le a_{2} \le \cdots \le a_{m})\) of indices from \(\set{1,\ldots,M}\); we write \(x_{A} = x_{a_{1}}x_{a_{2}}\cdots x_{a_{m}}\), \(\abs{A} = m\), and \(x_{A} = 1\) for the empty sequence. The monomials \(x_{A}\) are a basis of \(P\). For an index \(a\) we write
that is, if \(a\) does not exceed any index occurring in \(A\); the condition is vacuous, hence true, for the empty multi-index. Finally \(P_{n}\) denotes the span of the \(x_{A}\) with \(\abs{A} \le n\), so that \(P_{0} = \mathbb{K}\) and \(P_{n}P_{m} \subseteq P_{n+m}\).
There exist linear maps \(\theta_{a} : P \to P\), \(a = 1,\ldots,M\), such that for every multi-index \(A\) and all indices \(a,b\)
The assignment \(T_{a} \mapsto \theta_{a}\) is therefore a representation of \(\mathfrak{g}\) on \(P\). Rests on Equation (A.567) and Notation A.322.
Derives Proposition A.323. The maps are constructed on \(P_{n}\) by induction on \(n\). Write \((A_{n})\), \((B_{n})\), \((C_{n})\) for the three conditions restricted as follows: \((A_{n})\) and \((B_{n})\) are Equations (A.572) and (A.573) for \(\abs{A} \le n\), and \((C_{n})\) is Equation (A.574) applied to \(x_{A}\) with \(\abs{A} \le n-1\) — the largest range in which both sides are defined once \(\theta\) is known on \(P_{n}\), since \(\theta_{b}x_{A} \in P_{\abs{A}+1} \subseteq P_{n}\) by \((B_{n-1})\).
Base. \(P_{0} = \mathbb{K}\); set \(\theta_{a}1 = x_{a}\). Then \((A_{0})\) holds because \(a \le A\) is vacuous for the empty multi-index, \((B_{0})\) holds because the difference is \(0\), and \((C_{0})\) is empty.
Extension to \(P_{n}\). Assume \(\theta\) defined on \(P_{n-1}\) with \((A_{n-1})\), \((B_{n-1})\), \((C_{n-1})\). Let \(\abs{A} = n\).
-
If \(a \le A\), set \(\theta_{a}x_{A} = x_{a}x_{A}\), as Equation (A.572) demands.
-
If not, let \(b\) be the smallest index occurring in \(A\), so that \(b < a\), and write \(A = (b,B)\) with \(b \le B\) and \(\abs{B} = n-1\). Then \(x_{A} = x_{b}x_{B} = \theta_{b}x_{B}\) by \((A_{n-1})\). Using \((B_{n-1})\) write
\begin{equation}\tag{A.575} \theta_{a}x_{B} = x_{a}x_{B} + w\ec\qquad w \in P_{n-1}\ec \end{equation}and define
\begin{equation}\tag{A.576} \theta_{a}x_{A} := x_{a}x_{A} + \theta_{b}w + C^{c}{}_{ab}\,\theta_{c}x_{B}\ep \end{equation}Every term on the right is already defined: \(w\) and \(x_{B}\) lie in \(P_{n-1}\).
Condition \((A_{n})\) holds by construction. For \((B_{n})\): in case 1 the difference is \(0\); in case 2 it is \(\theta_{b}w + C^{c}{}_{ab}\theta_{c}x_{B}\), and \((B_{n-1})\) gives \(\theta_{b}w \in P_{n}\) and \(\theta_{c}x_{B} \in P_{n}\).
Before proving \((C_{n})\), record what Equation (A.576) says. With \(b < a\) and \(b \le B\), the monomial \(x_{a}x_{B}\) is \(x_{A'}\) for the multi-index \(A'\) obtained by inserting \(a\) into \(B\); its smallest index is still \(b\), so \(b \le A'\) and \(\abs{A'} = n\), and case 1 gives \(\theta_{b}(x_{a}x_{B}) = x_{b}x_{a}x_{B} = x_{a}x_{A}\), the variables commuting. Hence, by Equation (A.575), \(\theta_{b}\theta_{a}x_{B} = x_{a}x_{A} + \theta_{b}w\), and Equation (A.576) reads
using \(\theta_{b}x_{B} = x_{A}\) once more. So the commutation relation holds by fiat in the ordered case; the work is to show it then holds in general.
Proof of \((C_{n})\). Let \(\abs{A} = n-1\); the cases \(\abs{A} < n-1\) are \((C_{n-1})\). The claim is antisymmetric in \(a,b\) and trivial for \(a = b\), since \(C^{c}{}_{aa} = 0\), so assume \(a > b\).
Case 1: \(b \le A\). This is Equation (A.577) with \(B = A\).
Case 2: the relation \(b \le A\) fails. Let \(c_{0}\) be the smallest index occurring in \(A\), so \(c_{0} < b < a\), and write \(A = (c_{0},B)\) with \(c_{0} \le B\) and \(\abs{B} = n-2\). We first establish the auxiliary relation
Indeed, by \((B_{n-1})\) write \(y = x_{b}x_{B} + w'\) with \(w' \in P_{n-2}\). The monomial \(x_{b}x_{B}\) is \(x_{A''}\) for the multi-index \(A''\) obtained by inserting \(b\) into \(B\); its smallest index is \(c_{0}\), because \(c_{0} \le B\) and \(c_{0} < b\), and \(\abs{A''} = n-1\). So Case 1 — just proved, for multi-indices of length \(n-1\) — applies to \(x_{A''}\) with the pair \((a,c_{0})\), and \((C_{n-1})\) applies to \(w' \in P_{n-2}\); adding the two gives Equation (A.578).
Now compute, using \(x_{A} = \theta_{c_{0}}x_{B}\):
the second line by \((C_{n-1})\) applied to \(x_{B}\), \(\abs{B} = n-2\), and the third by Equation (A.578) applied to \(y = \theta_{b}x_{B}\). Exchange \(a\) and \(b\) in Equation (A.579) and subtract:
where the four mixed terms have been paired so that each parenthesis is a commutator. Every commutator in Equation (A.580) acts on \(x_{B}\) with \(\abs{B} = n-2\), so \((C_{n-1})\) evaluates all three:
The quantity to be matched is, again by \((C_{n-1})\) on \(x_{B}\),
The first terms of Equations (A.581) and (A.582) agree, so \((C_{n})\) holds if and only if
after renaming the summation index in Equation (A.582). Written with brackets, Equation (A.583) is
which is the Jacobi identity: expanding the right-hand side, \(\comm{\comm{T_{a}}{T_{b}}}{T_{c_{0}}} = \comm{T_{a}}{\comm{T_{b}}{T_{c_{0}}}} - \comm{T_{b}}{\comm{T_{a}}{T_{c_{0}}}}\), and the second term equals \(+\comm{\comm{T_{a}}{T_{c_{0}}}}{T_{b}}\) by antisymmetry. This completes the induction, and with it the construction of \(\theta\) on all of \(P = \bigcup_{n}P_{n}\).
∎The bookkeeping of Proposition A.323 is elaborate, but the mathematical content sits entirely in Equation (A.584). Conditions Equation (A.572) and Equation (A.573) merely say that \(\theta_{a}\) is “multiplication by \(x_{a}\), up to lower degree”, and the definition Equation (A.576) is forced on us: it is the only value compatible with Equation (A.574). Whether that forced value is consistent — whether reordering a triple by two different routes gives the same answer — is precisely Equation (A.583). An antisymmetric bracket violating Jacobi would therefore not merely fail to be a Lie algebra: its enveloping algebra would have fewer ordered monomials than expected, because the two routes would impose a relation between them.
Independence, and proof of the theorem
Proof of Theorem A.319. Derives Theorem A.319. By Proposition A.323 the map \(T_{a} \mapsto \theta_{a}\) is a representation of \(\mathfrak{g}\) on the vector space \(P\), and by Remark 14.6 it extends uniquely to a homomorphism of associative algebras with unit
(The extension of Remark 14.6 is stated there for a representation on a vector space; that \(P\) is infinite-dimensional plays no part, since no topology or trace is involved.)
Evaluate \(\Phi\) of an ordered monomial on the constant polynomial \(1\). Reading Equation (A.568) from right to left, the letter \(T_{M}\) acts \(n_{M}\) times on multi-indices all of whose entries are \(M\), then \(T_{M-1}\) acts on multi-indices whose entries are all \(\ge M-1\), and so on: at every step the acting index is \(\le\) every index already present, so Equation (A.572) applies at every step and
Distinct ordered monomials of \(U(\mathfrak{g})\) thus have distinct monomials of \(P\) as images, and the monomials of a polynomial algebra are linearly independent. Hence a vanishing linear combination of ordered monomials of \(U(\mathfrak{g})\) maps to a vanishing linear combination of distinct monomials of \(P\), forcing every coefficient to vanish: the ordered monomials are linearly independent. With Lemma A.321 they are a basis.
The two consequences follow at once. The monomials of degree one are the \(T_{a}\) themselves, part of a basis, so they are linearly independent in \(U(\mathfrak{g})\) and the map \(\mathfrak{g} \to U(\mathfrak{g})\) is injective. And if some element of the free associative algebra on the symbols \(T_{a}\) were annihilated in \(U(\mathfrak{g})\) beyond what Equation (14.12) forces, its reduction to ordered form — which uses those relations only — would be a nontrivial vanishing combination of basis elements.
∎Consequences: the symbol and the symmetrization map
The next two results are what Racah's Theorem on the Number of Casimir Operators needs. From here on \(\Q \subseteq \mathbb{K}\), so that \(n!\) may be inverted.
Let \(S(\mathfrak{g}) = \mathbb{K}[T_{1},\ldots,T_{M}]\) be the polynomial algebra on the same generators taken as commuting symbols, graded by degree, \(S(\mathfrak{g}) = \bigoplus_{n \ge 0}S^{n}(\mathfrak{g})\); invariantly, \(S^{n}(\mathfrak{g})\) is the \(n\)-th symmetric power of the vector space \(\mathfrak{g}\). By Theorem A.319 the ordered monomials of degree exactly \(n\) are a basis of a complement of \(U_{n-1}\) in \(U_{n}\), so
\(n_{1}+\cdots+n_{M} = n\), is a linear isomorphism. The element \(\sigma_{n}(u + U_{n-1})\) is the symbol of \(u \in U_{n}\), and \(\sigma = \bigoplus_{n}\sigma_{n}\) is an isomorphism of graded algebras from \(\operatorname{gr}U(\mathfrak{g}) = \bigoplus_{n}U_{n}/U_{n-1}\) onto \(S(\mathfrak{g})\): it is multiplicative because two ordered monomials multiply, modulo \(U_{n+m-1}\), by concatenation and reordering, and Equation (A.570) shows the reordering costs only terms of lower degree. Rests on Theorem A.319 and Definition A.320.
The symmetrization map is the linear map \(\omega : S(\mathfrak{g}) \to U(\mathfrak{g})\) determined by
the sum running over the \(n!\) permutations \(\pi\) of \(\set{1,\ldots,n}\), with \(\omega(1) = 1\). The right-hand side is symmetric in \(X_{1},\ldots,X_{n}\) and multilinear, so it descends from \(\mathfrak{g}^{n}\) to \(S^{n}(\mathfrak{g})\) and \(\omega\) is well defined. Rests on Definitions 14.5 and A.325.
Let \(\Q \subseteq \mathbb{K}\). Then:
-
\(\omega\) maps \(S^{n}(\mathfrak{g})\) into \(U_{n}\) and \(\sigma_{n}\left(\omega(p) + U_{n-1}\right) = p\) for \(p \in S^{n}(\mathfrak{g})\); consequently \(\omega\) is a linear isomorphism of \(S(\mathfrak{g})\) onto \(U(\mathfrak{g})\).
-
Let \(\mathfrak{g}\) act on \(U(\mathfrak{g})\) by \(\ad_{X}u = Xu - uX\) and on \(S(\mathfrak{g})\) by the derivation extending \(\ad_{X}\) on \(\mathfrak{g}\). Then \(\omega\) intertwines the two actions, \(\omega \circ \ad_{X} = \ad_{X} \circ\, \omega\).
Rests on Definition A.326, Theorem A.319 and Definition A.325.
Derives Proposition A.327. (1) Each word on the right of Equation (A.588) is a product of \(n\) elements of \(\mathfrak{g}\), hence lies in \(U_{n}\). Take \(p = T_{1}^{n_{1}}\cdots T_{M}^{n_{M}}\) of degree \(n\). Every one of the \(n!\) words in Equation (A.588) is a rearrangement of the same letters, so by Equation (A.570) each equals the ordered monomial \(T_{1}^{n_{1}}\cdots T_{M}^{n_{M}}\) of \(U(\mathfrak{g})\) modulo \(U_{n-1}\); averaging, \(\omega(p) - T_{1}^{n_{1}}\cdots T_{M}^{n_{M}} \in U_{n-1}\), which is the assertion about \(\sigma_{n}\). Thus \(\omega\) maps a basis of \(S(\mathfrak{g})\) to a family of elements that is “triangular” with respect to the filtration, with the ordered monomials as leading terms. Such a family is a basis: a nontrivial vanishing combination would have a highest degree \(n\) occurring in it, and its image under \(\sigma_{n}\) would be a nontrivial vanishing combination of distinct monomials of \(S^{n}(\mathfrak{g})\). Surjectivity follows because the leading terms exhaust the basis of Theorem A.319.
(2) On \(U(\mathfrak{g})\) the map \(\ad_{X}\) is a derivation of the associative product, \(\ad_{X}(uv) = (\ad_{X}u)v + u(\ad_{X}v)\), as one checks by expanding \(X(uv) - (uv)X\) and inserting \(\pm uXv\). Hence
using \(\ad_{X}Y = \comm{X}{Y}\) for \(Y \in \mathfrak{g}\). Average Equation (A.589) over \(\pi\) and group the resulting \(n!\,n\) words by the value \(i = \pi(j)\) of the slot that was hit: for fixed \(i\), the pairs \((\pi,j)\) with \(\pi(j) = i\) put \(\comm{X}{X_{i}}\) in position \(j\) and let the remaining letters occupy the remaining positions in every possible order, each exactly once — which is the definition of \(\omega\left(X_{1}\cdots\comm{X}{X_{i}}\cdots X_{n}\right)\) up to the common factor \(1/n!\). Therefore
the last equality because \(\ad_{X}\) acts on \(S(\mathfrak{g})\) as the derivation extending \(\ad_{X}\), which is exactly the Leibniz sum on the right.
∎Let \(k^{a_{1}\cdots a_{r}}\) be a nonzero totally symmetric array. Then the element \(C_{r} = k^{a_{1}\cdots a_{r}}T_{a_{1}}\cdots T_{a_{r}}\) of Equation (14.19) is nonzero in \(U(\mathfrak{g})\), and its symbol is the nonzero polynomial \(k^{a_{1}\cdots a_{r}}T_{a_{1}}\cdots T_{a_{r}} \in S^{r}(\mathfrak{g})\). In particular the quadratic Casimir Equation (14.20) of a semisimple algebra is a nonzero element of \(U(\mathfrak{g})\). Rests on Theorem A.319, Definition A.325 and Theorem 14.16.
Derives Corollary A.328. Write \(p = k^{a_{1}\cdots a_{r}}T_{a_{1}}\cdots T_{a_{r}}\) for the corresponding element of \(S^{r}(\mathfrak{g})\), the same expression read in commuting symbols. Then \(C_{r} = \omega(p)\): by linearity \(\omega(p) = k^{a_{1}\cdots a_{r}}\,\frac{1}{r!}\sum_{\pi} T_{a_{\pi(1)}}\cdots T_{a_{\pi(r)}}\), and renaming the summation indices in each of the \(r!\) terms turns it into \(k^{a_{1}\cdots a_{r}}T_{a_{1}}\cdots T_{a_{r}}\), because \(k^{a_{1}\cdots a_{r}}\) is totally symmetric. Now \(p \neq 0\): the coefficient of the commuting monomial \(T_{1}^{n_{1}}\cdots T_{M}^{n_{M}}\) in \(p\) is the multinomial coefficient \(r!/(n_{1}!\cdots n_{M}!)\) times the corresponding component of \(k^{a_{1}\cdots a_{r}}\), and in characteristic zero that factor cannot vanish. Since \(\omega\) is injective by Proposition A.327, \(C_{r} = \omega(p) \neq 0\). For Equation (14.20) the array is \(\kappa^{ab}\), which is invertible, hence nonzero.
∎Theorem A.319 is a statement about \(U(\mathfrak{g})\) as a vector space: the ordered monomials are a basis. It is not a statement that \(U(\mathfrak{g})\) is commutative, and \(\omega\) of Proposition A.327 is emphatically not an algebra homomorphism — \(\omega(XY) - \omega(X)\omega(Y) = -\tfrac{1}{2}\comm{X}{Y}\) for \(X,Y \in \mathfrak{g}\), by Equation (A.588) and Equation (14.12). What survives at the level of algebras is the statement of Definition A.325, that the associated graded algebra of \(U(\mathfrak{g})\) is the commutative algebra \(S(\mathfrak{g})\); noncommutativity is entirely a phenomenon of lower order in the filtration. That is the exact sense in which \(U(\mathfrak{g})\) is a deformation of the polynomial algebra on \(\mathfrak{g}\), and it is the form in which the theorem is used in Racah's Theorem on the Number of Casimir Operators.
The Poincaré–Birkhoff–Witt Theorem discharges the proof obligation of Theorem 14.7. Its immediate consequence is the one Section 14.1 needs: the generators satisfy no hidden relations, so a Casimir element built as in Theorem 14.16 is a genuine, nonzero element of \(U(\mathfrak{g})\) and not an elaborate way of writing zero (Corollary A.328). The deeper consequence, Proposition A.327, is the first link of the chain that counts those Casimirs in Racah's Theorem on the Number of Casimir Operators.
Cartan's Criterion for Semisimplicity
This appendix proves Theorem 14.13 of Lie Groups, Lie Algebras, and Fibre Bundles: a finite-dimensional Lie algebra over a field of characteristic zero contains no nonzero abelian ideal if and only if its Killing form Equation (14.14) is nondegenerate. The chapter uses the criterion twice — to know when \(\kappa_{ab}\) may be inverted, which is what makes the quadratic Casimir Equation (14.20) available, and in Proposition 14.77 on central extensions, whose proof in Whitehead's Lemmas and the Rigidity of Semisimple Algebras rests on this one throughout.
Nothing is quoted. The easy half is a two-line index computation (The easy half: an abelian ideal lies in the radical of the Killing form). The converse needs Cartan's criterion for solvability, which in turn needs the Jordan decomposition of an endomorphism and Engel's theorem; both are proved here (The Jordan decomposition and Engel's theorem), the trace argument that is the heart of the matter is The trace lemma, and the pieces are assembled in Proof of the criterion. Consequences, and the case of $\mathfrak{so}(p,q)$ draws the structural corollaries that Whitehead's Lemmas and the Rigidity of Semisimple Algebras needs and verifies the criterion on the family this treatise actually uses, the algebras \(\mathfrak{so}(p,q)\).
Throughout, \(\mathbb{K}\) is a field of characteristic zero with algebraic closure \(\overline{\mathbb{K}}\), and \(\mathfrak{g}\) is a finite-dimensional Lie algebra over \(\mathbb{K}\) with basis \(\set{T_{a}}\) and structure constants \(\comm{T_{a}}{T_{b}} = C^{c}{}_{ab}T_{c}\) as in Equation (14.5). For a finite-dimensional vector space \(\mathbb{V}\) we write \(\mathfrak{gl}(\mathbb{V})\) for the Lie algebra of all linear maps of \(\mathbb{V}\) with the commutator bracket. Recall that \(\kappa(X,Y) = \tr(\ad_{X}\ad_{Y})\) and that \(\kappa\) is invariant, Equation (14.16).
Statement
Let \(\mathfrak{g}\) be a finite-dimensional Lie algebra over a field of characteristic zero. Then \(\mathfrak{g}\) contains no nonzero abelian ideal if and only if its Killing form is nondegenerate, that is, if and only if
Rests on Definition 14.10 and Lemma 14.11.
Solvability, and the language of the proof
For subspaces \(\mathfrak{a},\mathfrak{b} \subseteq \mathfrak{g}\) let \(\comm{\mathfrak{a}}{\mathfrak{b}}\) denote the span of all \(\comm{X}{Y}\) with \(X \in \mathfrak{a}\), \(Y \in \mathfrak{b}\). A subspace \(\mathfrak{a}\) is a subalgebra if \(\comm{\mathfrak{a}}{\mathfrak{a}} \subseteq \mathfrak{a}\) and an ideal if \(\comm{\mathfrak{g}}{\mathfrak{a}} \subseteq \mathfrak{a}\). The derived series of a subalgebra \(\mathfrak{a}\) is
a decreasing chain of subalgebras, and \(\mathfrak{a}\) is solvable if \(\mathfrak{a}^{(m)} = 0\) for some \(m\). An abelian subalgebra is solvable with \(m = 1\).
Let \(\mathfrak{a}\) be a subalgebra of \(\mathfrak{g}\).
-
Subalgebras and homomorphic images of a solvable algebra are solvable.
-
If \(\mathfrak{b} \subseteq \mathfrak{a}\) is an ideal of \(\mathfrak{a}\) with \(\mathfrak{b}\) and \(\mathfrak{a}/\mathfrak{b}\) solvable, then \(\mathfrak{a}\) is solvable.
-
If \(\mathfrak{a}\) is an ideal of \(\mathfrak{g}\), then every \(\mathfrak{a}^{(j)}\) is an ideal of \(\mathfrak{g}\).
-
A nonzero solvable ideal of \(\mathfrak{g}\) contains a nonzero abelian ideal of \(\mathfrak{g}\).
Rests on Definition A.331.
Derives Lemma A.332. (1) If \(\mathfrak{b} \subseteq \mathfrak{a}\) then \(\mathfrak{b}^{(j)} \subseteq \mathfrak{a}^{(j)}\) by induction; if \(\phi\) is a homomorphism then \(\phi\left(\mathfrak{a}^{(j)}\right) = \phi(\mathfrak{a})^{(j)}\), again by induction, since \(\phi\) preserves brackets.
(2) If \(\left(\mathfrak{a}/\mathfrak{b}\right)^{(m)} = 0\) then \(\mathfrak{a}^{(m)} \subseteq \mathfrak{b}\) by (1) applied to the quotient map, and if \(\mathfrak{b}^{(l)} = 0\) then \(\mathfrak{a}^{(m+l)} \subseteq \mathfrak{b}^{(l)} = 0\).
(3) By induction on \(j\); the case \(j = 0\) is the hypothesis. Let \(\mathfrak{b} = \mathfrak{a}^{(j)}\) be an ideal. For \(Z \in \mathfrak{g}\) and \(X,Y \in \mathfrak{b}\) the Jacobi identity gives
and both terms on the right lie in \(\comm{\mathfrak{b}}{\mathfrak{b}} = \mathfrak{a}^{(j+1)}\) because \(\comm{Z}{X}\) and \(\comm{Z}{Y}\) lie in \(\mathfrak{b}\).
(4) Let \(\mathfrak{s} \neq 0\) be a solvable ideal and let \(m\) be largest with \(\mathfrak{s}^{(m)} \neq 0\). Then \(\comm{\mathfrak{s}^{(m)}}{\mathfrak{s}^{(m)}} = \mathfrak{s}^{(m+1)} = 0\), so \(\mathfrak{s}^{(m)}\) is abelian, and it is an ideal of \(\mathfrak{g}\) by (3).
∎Theorem 14.13 defines semisimple as “contains no nonzero abelian ideal”. The literature usually says “contains no nonzero solvable ideal”, equivalently “the radical — the largest solvable ideal — is zero”. Part (4) of Lemma A.332 is exactly the statement that the two conditions coincide: an abelian ideal is solvable, and a nonzero solvable ideal produces a nonzero abelian one. We use whichever is convenient and say “semisimple” for both.
The easy half: an abelian ideal lies in the radical of the Killing form
Let \(\mathfrak{a} \subseteq \mathfrak{g}\) be an abelian ideal. Then \(\mathfrak{a} \subseteq \mathfrak{g}^{\perp}\). Consequently, if \(\kappa\) is nondegenerate then \(\mathfrak{g}\) has no nonzero abelian ideal. Rests on Definitions 14.10 and A.331.
Derives Proposition A.334. Let \(X \in \mathfrak{a}\) and \(Y \in \mathfrak{g}\), and put \(N = \ad_{X}\ad_{Y}\). Since \(\mathfrak{a}\) is an ideal, \(\ad_{X}\mathfrak{g} = \comm{X}{\mathfrak{g}} \subseteq \mathfrak{a}\), so \(N\mathfrak{g} \subseteq \mathfrak{a}\); and since \(\mathfrak{a}\) is an ideal, \(\ad_{Y}\mathfrak{a} \subseteq \mathfrak{a}\), while \(\ad_{X}\mathfrak{a} = \comm{\mathfrak{a}}{\mathfrak{a}} = 0\) because \(\mathfrak{a}\) is abelian. Hence
so \(N^{2} = 0\). A nilpotent endomorphism has all eigenvalues zero and therefore vanishing trace, so \(\kappa(X,Y) = \tr N = 0\) for every \(Y\), that is, \(X \in \mathfrak{g}^{\perp}\).
The same computation in components, with the index conventions of Equation (14.15), is worth displaying because it is the form in which the statement is usually met. Choose a basis adapted to \(\mathfrak{a}\): indices \(i,j,k\) label a basis of \(\mathfrak{a}\) and indices \(\mu,\nu\) complete it. That \(\mathfrak{a}\) is an ideal says \(C^{\mu}{}_{a i} = 0\) for every \(a\); that it is abelian says \(C^{c}{}_{ij} = 0\). Then, from Equation (14.15),
the second equality because \(C^{c}{}_{id}\) vanishes unless the upper index lies in \(\mathfrak{a}\), the third because \(C^{d}{}_{bj}\) vanishes unless \(d\) does, and the last because \(C^{j}{}_{ik} = 0\).
∎The Jordan decomposition
Let \(\mathbb{V}\) be a finite-dimensional vector space over an algebraically closed field \(\mathbb{F}\) and let \(x \in \operatorname{End}(\mathbb{V})\). There are unique \(x_{s}, x_{n} \in \operatorname{End}(\mathbb{V})\) with
Moreover \(x_{s}\) and \(x_{n}\) are polynomials in \(x\) with zero constant term. Rests on Definition 5.35.
Derives Theorem A.335. Let \(a_{1},\ldots,a_{r}\) be the distinct eigenvalues of \(x\) and \(\prod_{i}(t-a_{i})^{m_{i}}\) its characteristic polynomial, which splits because the field is algebraically closed. Put \(\mathbb{V}_{i} = \ker\left(x-a_{i}\right)^{m_{i}}\). The polynomials \((t-a_{i})^{m_{i}}\) are pairwise coprime, so by the Chinese remainder theorem in \(\mathbb{F}[t]\) — applied modulo the characteristic polynomial, which annihilates \(x\) by Cayley–Hamilton — the \(\mathbb{V}_{i}\) are the images of projections that are polynomials in \(x\), and \(\mathbb{V} = \bigoplus_{i}\mathbb{V}_{i}\).
By the same theorem choose \(p \in \mathbb{F}[t]\) with
the last congruence being imposed only when \(0\) is not among the \(a_{i}\) — when it is, say \(a_{1} = 0\), the congruence \(p \equiv 0 \pmod{t^{m_{1}}}\) already forces \(p(0) = 0\), and the moduli stay pairwise coprime in either case. Set \(x_{s} = p(x)\) and \(x_{n} = x - x_{s}\). On \(\mathbb{V}_{i}\) the first congruence gives \(x_{s} = a_{i}\identity\), so \(x_{s}\) is diagonalizable, and \(x_{n} = x - a_{i}\identity\) there, which is nilpotent on \(\mathbb{V}_{i}\) by definition of \(\mathbb{V}_{i}\); being nilpotent on each summand of a finite direct sum it is nilpotent. Both are polynomials in \(x\) with \(p(0) = 0\), hence commute with \(x\) and with each other.
Uniqueness. Let \(x = s' + n'\) be another such decomposition. Since \(s'\) and \(n'\) commute with each other they commute with \(x\), hence with every polynomial in \(x\), hence with \(x_{s}\) and \(x_{n}\). Then \(x_{s} - s' = n' - x_{n}\) is at once diagonalizable — a difference of two commuting diagonalizable maps, which are simultaneously diagonalizable — and nilpotent, being a difference of two commuting nilpotent maps. A map that is both is zero.
∎With the notation of Theorem A.335, viewing \(\ad_{x}(y) = \comm{x}{y}\) as an endomorphism of \(\operatorname{End}(\mathbb{V})\),
In particular \(\ad_{x_{s}}\) is a polynomial in \(\ad_{x}\) with zero constant term. Rests on Theorem A.335.
Derives Lemma A.336. Let \(e_{1},\ldots,e_{m}\) be a basis of eigenvectors of \(x_{s}\), \(x_{s}e_{i} = a_{i}e_{i}\), and let \(E_{ij}\) be the corresponding matrix units. Then \(\ad_{x_{s}}E_{ij} = (a_{i}-a_{j})E_{ij}\), so \(\ad_{x_{s}}\) is diagonalizable. Writing \(L(y) = x_{n}y\) and \(R(y) = yx_{n}\), the maps \(L\) and \(R\) commute and are nilpotent, so \(\ad_{x_{n}} = L - R\) is nilpotent. And \(\comm{\ad_{x_{s}}}{\ad_{x_{n}}} = \ad_{\comm{x_{s}}{x_{n}}} = 0\). Hence \(\ad_{x} = \ad_{x_{s}} + \ad_{x_{n}}\) satisfies the three conditions of Equation (A.596), and uniqueness identifies the two summands. The last sentence is the final statement of Theorem A.335 applied to the endomorphism \(\ad_{x}\) of \(\operatorname{End}(\mathbb{V})\).
∎Engel's theorem
If \(x \in \operatorname{End}(\mathbb{V})\) is nilpotent then \(\ad_{x}\) is nilpotent. Rests on Definition 5.35.
Derives Lemma A.337. With \(L(y) = xy\) and \(R(y) = yx\) as above, \(L\) and \(R\) commute and each is nilpotent, since \(x^{N} = 0\) gives \(L^{N} = R^{N} = 0\). The binomial theorem for commuting maps gives \(\left(L-R\right)^{2N} = 0\).
∎Let \(\mathbb{V} \neq 0\) be finite-dimensional and let \(\mathfrak{h} \subseteq \mathfrak{gl}(\mathbb{V})\) be a Lie subalgebra all of whose elements are nilpotent endomorphisms. Then there is a nonzero \(v \in \mathbb{V}\) with \(\mathfrak{h}v = 0\); consequently there is a basis of \(\mathbb{V}\) in which every element of \(\mathfrak{h}\) is strictly upper triangular, and \(\mathfrak{h}\) is solvable. Rests on Lemma A.337 and Definition A.331.
Derives Theorem A.338. Induct on \(\dim\mathfrak{h}\). If \(\mathfrak{h} = 0\) any nonzero \(v\) will do. Let \(\dim\mathfrak{h} \ge 1\) and choose a maximal proper subalgebra \(\mathfrak{m} \subsetneq \mathfrak{h}\), which exists because \(0\) is a proper subalgebra and dimensions are finite.
\(\mathfrak{m}\) is an ideal of codimension one. The adjoint action of \(\mathfrak{m}\) on \(\mathfrak{h}\) preserves \(\mathfrak{m}\), so it descends to an action on \(\mathfrak{h}/\mathfrak{m} \neq 0\); each \(\ad_{X}\), \(X \in \mathfrak{m}\), is nilpotent by Lemma A.337, hence so is the induced map. The image of \(\mathfrak{m}\) in \(\mathfrak{gl}\left(\mathfrak{h}/\mathfrak{m}\right)\) is a subalgebra of dimension at most \(\dim\mathfrak{m} < \dim\mathfrak{h}\) consisting of nilpotent maps, so the induction hypothesis supplies a nonzero \(z + \mathfrak{m}\) annihilated by it: there is \(z \in \mathfrak{h} \setminus \mathfrak{m}\) with \(\comm{\mathfrak{m}}{z} \subseteq \mathfrak{m}\). Then \(\mathfrak{m} + \mathbb{K}z\) is a subalgebra strictly containing \(\mathfrak{m}\), so it is \(\mathfrak{h}\) by maximality, and \(\mathfrak{m}\) is an ideal of codimension one.
The common kernel. Let \(\mathbb{W} = \set{v \in \mathbb{V} \mid \mathfrak{m}v = 0}\), which is nonzero by the induction hypothesis applied to \(\mathfrak{m}\). It is stable under \(\mathfrak{h}\): for \(X \in \mathfrak{m}\), \(Y \in \mathfrak{h}\) and \(v \in \mathbb{W}\),
since \(\comm{X}{Y} \in \mathfrak{m}\). In particular \(z\) maps \(\mathbb{W}\) to itself, and \(z\) is nilpotent, so it has a nonzero kernel vector \(v \in \mathbb{W}\). That \(v\) is annihilated by \(\mathfrak{m}\) and by \(z\), hence by \(\mathfrak{h} = \mathfrak{m} + \mathbb{K}z\).
The flag. Induct on \(\dim\mathbb{V}\): take \(v\) as above as the first basis vector and apply the statement to the induced action on \(\mathbb{V}/\mathbb{K}v\), whose elements are again nilpotent. In the resulting basis every \(X \in \mathfrak{h}\) is strictly upper triangular.
Solvability. Let \(\mathfrak{n}_{r}\) be the space of matrices whose entries vanish except at least \(r\) places above the diagonal, so that \(\mathfrak{h} \subseteq \mathfrak{n}_{1}\) and \(\comm{\mathfrak{n}_{r}}{\mathfrak{n}_{s}} \subseteq \mathfrak{n}_{r+s}\) by matrix multiplication. Then \(\mathfrak{h}^{(j)} \subseteq \mathfrak{n}_{2^{j}}\) by induction, and \(\mathfrak{n}_{r} = 0\) once \(r \ge \dim\mathbb{V}\).
∎The trace lemma
Everything so far has been preparation. The following lemma is the mechanism of Cartan's criterion, and the only place where characteristic zero is used in an essential way — through the rational numbers, of all things, inside a field that may be as large as \(\C\).
Let \(\mathbb{V}\) be finite-dimensional over an algebraically closed field \(\mathbb{F}\) of characteristic zero, let \(A \subseteq B\) be subspaces of \(\mathfrak{gl}(\mathbb{V})\), and put
If \(x \in M\) satisfies \(\tr(xy) = 0\) for every \(y \in M\), then \(x\) is nilpotent. Rests on Theorem A.335 and Lemma A.336.
Derives Lemma A.339. Let \(x = x_{s} + x_{n}\) be the Jordan decomposition (Theorem A.335), let \(e_{1},\ldots,e_{m}\) be a basis of eigenvectors of \(x_{s}\) with \(x_{s}e_{i} = a_{i}e_{i}\), and let \(E \subseteq \mathbb{F}\) be the \(\Q\)-linear span of \(a_{1},\ldots,a_{m}\), a finite-dimensional vector space over \(\Q\). The claim is that \(E = 0\); then every \(a_{i} = 0\), so \(x_{s} = 0\) and \(x = x_{n}\) is nilpotent.
Since \(E\) is finite-dimensional over \(\Q\), it is enough to show that every \(\Q\)-linear functional \(f : E \to \Q\) vanishes. Fix such an \(f\) and let \(y \in \operatorname{End}(\mathbb{V})\) be defined by \(ye_{i} = f(a_{i})e_{i}\).
Step 1: \(\ad_{y}\) is a polynomial in \(\ad_{x}\) with zero constant term. In the basis of matrix units \(E_{ij}\) built from the \(e_{i}\),
Choose \(r \in \mathbb{F}[t]\) with \(r(0) = 0\) and \(r\left(a_{i}-a_{j}\right) = f(a_{i})-f(a_{j})\) for all \(i,j\). This is a consistent finite set of interpolation conditions: if \(a_{i}-a_{j} = a_{p}-a_{q}\) then, applying the \(\Q\)-linear \(f\), \(f(a_{i})-f(a_{j}) = f(a_{p})-f(a_{q})\); and if \(a_{i}-a_{j} = 0\) the prescribed value is \(0\), which is what \(r(0) = 0\) demands. Lagrange interpolation over the finitely many distinct values then produces \(r\). By Equation (A.601), \(\ad_{y} = r\left(\ad_{x_{s}}\right)\), and by Lemma A.336 \(\ad_{x_{s}} = q\left(\ad_{x} \right)\) for some \(q\) with \(q(0) = 0\). Hence \(\ad_{y} = r\left(q\left(\ad_{x}\right)\right)\), a polynomial in \(\ad_{x}\) without constant term.
Step 2: \(y \in M\). Since \(x \in M\) we have \(\ad_{x}(B) \subseteq A \subseteq B\), so \(\left(\ad_{x}\right)^{j}(B) \subseteq A\) for every \(j \ge 1\). A polynomial in \(\ad_{x}\) with zero constant term is a combination of such powers, so \(\ad_{y}(B) \subseteq A\), that is, \(y \in M\).
Step 3: the trace. By hypothesis \(\tr(xy) = 0\). Compute it. Since \(x_{n}\) commutes with \(x_{s}\) it preserves each eigenspace of \(x_{s}\), and inside each eigenspace a basis may be chosen making the nilpotent \(x_{n}\) strictly upper triangular; taking the \(e_{i}\) so adapted, \(x\) is upper triangular with diagonal entries \(a_{1},\ldots,a_{m}\) while \(y\) is diagonal with entries \(f(a_{1}),\ldots,f(a_{m})\). Then \(xy\) is upper triangular with diagonal entries \(a_{i}f(a_{i})\), so
The left-hand side lies in \(E\), since each \(f(a_{i}) \in \Q\) and \(E\) is a \(\Q\)-subspace containing the \(a_{i}\). Apply \(f\) to Equation (A.602) and use \(\Q\)-linearity:
A sum of squares of rational numbers vanishes only if every term does, so \(f(a_{i}) = 0\) for all \(i\); as the \(a_{i}\) span \(E\) over \(\Q\), \(f = 0\). Every \(\Q\)-linear functional on \(E\) vanishes, so \(E = 0\).
∎Let \(\mathbb{V}\) be finite-dimensional over a field \(\mathbb{K}\) of characteristic zero and let \(\mathfrak{h} \subseteq \mathfrak{gl}(\mathbb{V})\) be a Lie subalgebra with
Then \(\mathfrak{h}\) is solvable. Rests on Lemma A.339 and Theorem A.338.
Derives Theorem A.340. Assume first that \(\mathbb{K}\) is algebraically closed. Apply Lemma A.339 with \(A = \comm{\mathfrak{h}}{\mathfrak{h}}\), \(B = \mathfrak{h}\) and \(M\) as in Equation (A.600); note \(\mathfrak{h} \subseteq M\). Let \(x \in \comm{\mathfrak{h}}{\mathfrak{h}}\) and \(y \in M\). By bilinearity it suffices to treat \(x = \comm{u}{v}\) with \(u,v \in \mathfrak{h}\), and then
using cyclicity of the trace twice. Now \(y \in M\) gives \(\comm{v}{y} \in \comm{\mathfrak{h}}{\mathfrak{h}}\), and \(u \in \mathfrak{h}\), so the right-hand side vanishes by Equation (A.604). Lemma A.339 therefore makes every element of \(\comm{\mathfrak{h}}{\mathfrak{h}}\) a nilpotent endomorphism, and Theorem A.338 makes \(\comm{\mathfrak{h}}{\mathfrak{h}}\) solvable. Since \(\mathfrak{h}/\comm{\mathfrak{h}}{\mathfrak{h}}\) is abelian, Lemma A.332(2) makes \(\mathfrak{h}\) solvable.
For general \(\mathbb{K}\), extend scalars to \(\overline{\mathbb{K}}\). The hypothesis Equation (A.604) is bilinear and \(\comm{\mathfrak{h}}{\mathfrak{h}}\) spans \(\comm{\mathfrak{h}_{\overline{\mathbb{K}}}} {\mathfrak{h}_{\overline{\mathbb{K}}}}\), so the hypothesis holds for \(\mathfrak{h}_{\overline{\mathbb{K}}} \subseteq \mathfrak{gl}\left(\mathbb{V}_{\overline{\mathbb{K}}}\right)\), which is therefore solvable. Since \(\left(\mathfrak{h}^{(j)}\right)_{\overline{\mathbb{K}}} = \left(\mathfrak{h}_{\overline{\mathbb{K}}}\right)^{(j)}\) — brackets of spanning sets span — and a subspace vanishes if and only if its extension does, \(\mathfrak{h}\) is solvable.
∎Proof of the criterion
Proof of Theorem A.330. Derives Theorem A.330. One direction is Proposition A.334. For the other, assume \(\mathfrak{g}\) has no nonzero abelian ideal and let \(\mathfrak{s} = \mathfrak{g}^{\perp}\) be the radical of the Killing form, Equation (A.591).
\(\mathfrak{s}\) is an ideal. Let \(X \in \mathfrak{s}\) and \(Z,Y \in \mathfrak{g}\). By the invariance of the Killing form, Equation (14.16),
so \(\comm{Z}{X} \in \mathfrak{s}\).
\(\ad(\mathfrak{s})\) is solvable. Consider \(\ad(\mathfrak{s}) \subseteq \mathfrak{gl}(\mathfrak{g})\), a Lie subalgebra because \(\ad\) is a homomorphism. Its derived algebra is \(\ad\left(\comm{\mathfrak{s}}{\mathfrak{s}}\right)\). Take \(x = \ad_{W}\) with \(W \in \comm{\mathfrak{s}}{\mathfrak{s}} \subseteq \mathfrak{s}\) and \(y = \ad_{Y}\) with \(Y \in \mathfrak{s}\); then
because \(W \in \mathfrak{s} = \mathfrak{g}^{\perp}\). So Theorem A.340 applies and \(\ad(\mathfrak{s})\) is solvable.
\(\mathfrak{s}\) is solvable. The kernel of \(\ad\) restricted to \(\mathfrak{s}\) is \(\mathfrak{s} \cap Z(\mathfrak{g})\), where \(Z(\mathfrak{g}) = \set{X \in \mathfrak{g} \mid \comm{X}{\mathfrak{g}} = 0}\) is the centre; it is abelian, hence solvable; the image \(\ad(\mathfrak{s})\) is solvable; so Lemma A.332(2) applies to the ideal \(\ker\left(\ad|_{\mathfrak{s}}\right)\) of \(\mathfrak{s}\).
Conclusion. If \(\mathfrak{s} \neq 0\) then, being a nonzero solvable ideal of \(\mathfrak{g}\), it would contain a nonzero abelian ideal of \(\mathfrak{g}\) by Lemma A.332(4), contrary to hypothesis. Hence \(\mathfrak{s} = 0\) and \(\kappa\) is nondegenerate.
∎Semisimplicity is unchanged by enlarging the field, and the criterion is what makes this obvious. In a basis of \(\mathfrak{g}\) over \(\mathbb{K}\) the Killing form has a Gram matrix \(\kappa_{ab}\) of scalars, and the same basis is a basis of \(\mathfrak{g}_{K} = \mathfrak{g}\otimes_{\mathbb{K}}K\) over any extension \(K\) with the same structure constants, hence the same \(\kappa_{ab}\) by Equation (14.15). Nondegeneracy is the nonvanishing of \(\det\left(\kappa_{ab}\right)\), a condition on that matrix alone. So \(\mathfrak{g}\) is semisimple if and only if \(\mathfrak{g}_{K}\) is — a statement that is not obvious from the definition by ideals, since \(\mathfrak{g}_{K}\) has many more subspaces than \(\mathfrak{g}\).
Consequences, and the case of $\mathfrak{so}(p,q)$
Let \(\mathfrak{g}\) be semisimple and \(\mathfrak{a} \subseteq \mathfrak{g}\) an ideal, and let \(\mathfrak{a}^{\perp} = \set{X \mid \kappa(X,\mathfrak{a}) = 0}\). Then
both summands are ideals, both are semisimple, and \(\mathfrak{g}/\mathfrak{a} \cong \mathfrak{a}^{\perp}\). Consequently \(\mathfrak{g}\) is a direct sum of simple ideals, \(Z(\mathfrak{g}) = 0\), and \(\comm{\mathfrak{g}}{\mathfrak{g}} = \mathfrak{g}\). Rests on Theorems A.330 and A.340.
Derives Corollary A.342. \(\mathfrak{a}^{\perp}\) is an ideal by the computation Equation (A.606). Put \(\mathfrak{b} = \mathfrak{a} \cap \mathfrak{a}^{\perp}\), an ideal. For \(W \in \comm{\mathfrak{b}}{\mathfrak{b}} \subseteq \mathfrak{a}\) and \(Y \in \mathfrak{b} \subseteq \mathfrak{a}^{\perp}\) we have \(\tr\left(\ad_{W}\ad_{Y}\right) = \kappa(W,Y) = 0\), so Theorem A.340 makes \(\ad(\mathfrak{b})\) solvable and, exactly as in the proof of Theorem A.330, \(\mathfrak{b}\) solvable. A semisimple algebra has no nonzero solvable ideal (Remark A.333), so \(\mathfrak{b} = 0\). Nondegeneracy of \(\kappa\) gives \(\dim\mathfrak{a}^{\perp} = \dim\mathfrak{g} - \dim\mathfrak{a}\), whence the direct sum. Then \(\comm{\mathfrak{a}}{\mathfrak{a}^{\perp}} \subseteq \mathfrak{a}\cap\mathfrak{a}^{\perp} = 0\), both being ideals.
An abelian ideal of \(\mathfrak{a}\) is an ideal of \(\mathfrak{g}\), since \(\mathfrak{a}^{\perp}\) brackets to zero with it; so it vanishes, and \(\mathfrak{a}\) is semisimple, as is \(\mathfrak{a}^{\perp}\). The projection \(\mathfrak{g} \to \mathfrak{a}^{\perp}\) along \(\mathfrak{a}\) is a homomorphism with kernel \(\mathfrak{a}\).
Iterating the splitting on a proper nonzero ideal, and stopping when none exists, writes \(\mathfrak{g} = \mathfrak{g}_{1}\oplus\cdots\oplus \mathfrak{g}_{s}\) with each \(\mathfrak{g}_{j}\) having no proper nonzero ideal; each is nonabelian, since an abelian \(\mathfrak{g}_{j}\) would be an abelian ideal of \(\mathfrak{g}\), so each is simple. The centre of \(\mathfrak{g}\) is an abelian ideal, hence zero. Finally \(\comm{\mathfrak{g}_{j}}{\mathfrak{g}_{j}}\) is a nonzero ideal of the simple \(\mathfrak{g}_{j}\) — nonzero because \(\mathfrak{g}_{j}\) is not abelian — so it is all of \(\mathfrak{g}_{j}\), and summing over \(j\) gives \(\comm{\mathfrak{g}}{\mathfrak{g}} = \mathfrak{g}\).
∎In the decomposition \(\mathfrak{g} = \bigoplus_{j}\mathfrak{g}_{j}\) of Corollary A.342, \(\kappa\left(\mathfrak{g}_{i}, \mathfrak{g}_{j}\right) = 0\) for \(i \neq j\), and the restriction of \(\kappa\) to \(\mathfrak{g}_{j}\) is the Killing form of \(\mathfrak{g}_{j}\) computed in \(\mathfrak{g}_{j}\) itself. Rests on Corollary A.342 and Definition 14.10.
Derives Corollary A.343. Let \(X \in \mathfrak{g}_{i}\), \(Y \in \mathfrak{g}_{j}\), \(i \neq j\). Then \(\ad_{Y}\) annihilates every summand but \(\mathfrak{g}_{j}\) and maps \(\mathfrak{g}_{j}\) into itself, while \(\ad_{X}\) annihilates \(\mathfrak{g}_{j}\); so \(\ad_{X}\ad_{Y} = 0\) and \(\kappa(X,Y) = 0\). For \(X,Y \in \mathfrak{g}_{j}\) the map \(\ad_{X}\ad_{Y}\) annihilates the other summands and preserves \(\mathfrak{g}_{j}\), so its trace over \(\mathfrak{g}\) equals its trace over \(\mathfrak{g}_{j}\).
∎Let \(D = p+q \ge 3\) and let \(\mathfrak{so}(p,q)\) have the generators \(J_{AB}\) and brackets Equation (14.90). In the basis \(\set{J_{AB}}_{A<B}\) the Killing form is diagonal,
with \(\kappa(J_{AB},J_{AB}) = -2(D-2)\,\eta_{AA}\eta_{BB} \neq 0\) for \(A \neq B\) and no summation. Hence \(\kappa\) is nondegenerate and \(\mathfrak{so}(p,q)\) is semisimple for every signature with \(D \ge 3\); in particular the Lorentz algebra \(\mathfrak{so}(D-1,1)\) is, and \(\mathfrak{so}(2)\) — where Equation (A.609) vanishes identically — is not, being abelian. Rests on Theorem A.330, Equation (14.90) and Proposition 14.59.
Derives Corollary A.344. That the \(J_{AB}\) with \(A<B\) are a basis is part (ii) of Proposition 14.59, and \(\eta\) is diagonal by Notation 14.1.
Off-diagonal entries vanish. Fix an index \(A_{0}\) and let \(\varsigma\) be the reflection \(x^{A_{0}} \mapsto -x^{A_{0}}\) of \(\R^{D}\), the other coordinates unchanged. It preserves \(\eta\), so its pushforward on vector fields maps Killing fields to Killing fields and, being induced by a diffeomorphism, preserves Lie brackets: it is an automorphism of \(\mathfrak{so}(p,q)\). On the generators Equation (14.86) it acts by \(\varsigma_{*}J_{AB} = \epsilon_{A}\epsilon_{B}J_{AB}\), where \(\epsilon_{A} = -1\) if \(A = A_{0}\) and \(+1\) otherwise, because both \(x_{A}\) and \(\pp_{B}\) pick up their own sign. The Killing form is invariant under any automorphism \(\phi\), since \(\ad_{\phi X} = \phi\,\ad_{X}\phi^{-1}\) and the trace is conjugation invariant; hence
If \(\set{A,B} \neq \set{C,D}\), choose \(A_{0}\) in the symmetric difference of the two pairs; then exactly one of the four factors is \(-1\) — the pairs have distinct entries — so the sign is \(-1\) and \(\kappa\left(J_{AB},J_{CD}\right) = 0\).
Diagonal entries. Take \(H = J_{12}\); the general case is the same computation with the indices renamed. From Equation (14.90), \(\comm{J_{12}}{J_{CD}} = \eta_{1C}J_{D2} + \eta_{2D}J_{C1} + \eta_{1D}J_{2C} + \eta_{2C}J_{1D}\), and since \(\eta\) is diagonal only the terms whose two indices coincide survive. For \(k \ge 3\),
while \(\ad_{H}J_{12} = 0\) and \(\ad_{H}J_{kl} = 0\) for \(k,l \ge 3\). Hence \(\ad_{H}^{2}\) is diagonal in the basis, acting as \(-\eta_{11}\eta_{22}\) on each of the \(2(D-2)\) generators \(J_{1k}\), \(J_{2k}\) with \(k \ge 3\), and as zero on the rest, so
which is Equation (A.609) for \(A=C=1\), \(B=D=2\); the general case of Equation (A.609) follows because both sides vanish off the diagonal and both are antisymmetric in \(AB\) and in \(CD\). Since \(\eta_{AA}\eta_{BB} = \pm1\), the diagonal entries are nonzero exactly when \(D \neq 2\), and Theorem A.330 converts nondegeneracy into semisimplicity.
∎Three things follow immediately and are used without further comment. First, \(\kappa_{ab}\) may be inverted for a semisimple algebra, which is what Corollary 14.17 needs to write \(C_{2} = \kappa^{ab}T_{a}T_{b}\) and what Remark 14.15 needs to move indices between the two forms of invariance. Second, by Corollary A.342 a semisimple algebra equals its own derived algebra; this is why a semisimple algebra has no nonzero homomorphism to an abelian one, a fact used repeatedly in Whitehead's Lemmas and the Rigidity of Semisimple Algebras. Third, Remark 14.23 is now precise rather than plausible: the Poincaré algebra Equation (14.99) contains the translations as an abelian ideal, so Proposition A.334 puts every translation generator in the radical of its Killing form, the form is degenerate, and no statement proved for semisimple algebras may be applied to it.
Cartan's Criterion for Semisimplicity discharges the proof obligation of Theorem 14.13. It is used at Corollary 14.17 to invert the Killing form, at Proposition 14.22 and Theorem 14.21 wherever semisimplicity is assumed, and — through Corollaries A.342 and A.344 — as the foundation of Whitehead's Lemmas and the Rigidity of Semisimple Algebras, which proves Proposition 14.77 on central extensions.
Racah's Theorem on the Number of Casimir Operators
This appendix proves Theorem 14.21 of Lie Groups, Lie Algebras, and Fibre Bundles: for a finite-dimensional semisimple Lie algebra of rank \(\ell\) (Definition 14.20), the centre \(Z\left(U(\mathfrak{g})\right)\) of the universal enveloping algebra is a polynomial algebra on \(\ell\) algebraically independent Casimir elements, so that an irreducible representation carries exactly \(\ell\) invariant labels. The chapter needs the count in two places: to know that \(\mathfrak{su}(2)\) has one Casimir and \(\mathfrak{su}(3)\) two (Remark 14.19), and to convert the rank of \(\mathfrak{so}(p,q)\) computed in Proposition 14.22 into the number of Lorentz invariants.
The proof is a chain of three identifications:
where \(\mathfrak{h}\) is a Cartan subalgebra, \(W\) its Weyl group, and “invariant” means annihilated by every \(\ad_{X}\). The first link is proved here in full (The centre of the enveloping algebra and the invariant polynomials), and rests on the Poincaré–Birkhoff–Witt theorem of The Poincaré–Birkhoff–Witt Theorem. The second is the Chevalley restriction theorem and the third is Chevalley's theorem on finite reflection groups; both belong to the structure theory of semisimple Lie algebras, which this treatise does not build, and both are stated precisely as quoted inputs in The two quoted theorems and named again in Remark A.353. The assembly and the count are The count, and the answer is checked against the chapter's own worked algebras in The count checked against the chapter's algebras.
Throughout, \(\mathfrak{g}\) is a finite-dimensional semisimple Lie algebra over a field \(\mathbb{K}\) of characteristic zero, with basis \(\set{T_{a}}\), Killing form \(\kappa\) — nondegenerate by Theorem A.330 — and enveloping algebra \(U(\mathfrak{g})\) of Definition 14.5, filtered by \(U_{n}\) as in Definition A.320. We write \(S(\mathfrak{g})\) for the symmetric algebra of Definition A.325 and \(\omega\) for the symmetrization map Equation (A.588).
Statement
Let \(\mathfrak{g}\) be a finite-dimensional semisimple Lie algebra of rank \(\ell\) over an algebraically closed field of characteristic zero. Then there are \(\ell\) Casimir elements \(C^{(1)},\ldots,C^{(\ell)} \in Z\left(U(\mathfrak{g})\right)\), algebraically independent, such that every element of the centre is a polynomial in them with scalar coefficients,
and no set of fewer than \(\ell\) elements of the centre generates it. By Corollary 14.18 each \(C^{(i)}\) acts on a finite-dimensional irreducible representation as a scalar, so such a representation carries \(\ell\) invariant labels and no more; that these \(\ell\) numbers also separate the finite-dimensional irreducible representations is the quoted Theorem A.354. Rests on Definition 14.20, Definition 14.8 and Theorem A.330.
The centre of the enveloping algebra and the invariant polynomials
For \(u \in U(\mathfrak{g})\) the following are equivalent: \(u\) is central; \(\comm{u}{X} = 0\) for every \(X \in \mathfrak{g}\); and \(u\) is invariant under the adjoint action \(\ad_{X}u = Xu - uX\) of \(\mathfrak{g}\) on \(U(\mathfrak{g})\). In particular the Casimir elements of Definition 14.8 are exactly the \(\ad\)-invariant elements of \(U(\mathfrak{g})\). Rests on Definitions 14.5 and 14.8.
Derives Lemma A.347. The second and third conditions are the same statement written twice. A central element certainly commutes with every \(X \in \mathfrak{g}\). Conversely, the set of elements commuting with a fixed \(u\) is a subalgebra of \(U(\mathfrak{g})\) containing the unit — it is closed under sums and, by the Leibniz rule \(\comm{vw}{u} = v\comm{w}{u} + \comm{v}{u}w\), under products. If it contains \(\mathfrak{g}\) it therefore contains the algebra generated by \(\mathfrak{g}\) and the unit, which is \(U(\mathfrak{g})\) by Definition 14.5.
∎Let \(S(\mathfrak{g})^{\mathfrak{g}}\) denote the \(\ad\)-invariant elements of \(S(\mathfrak{g})\). Then:
-
\(\omega\) restricts to a linear isomorphism \(S(\mathfrak{g})^{\mathfrak{g}} \to Z\left(U(\mathfrak{g})\right)\);
-
\(S(\mathfrak{g})^{\mathfrak{g}}\) is a graded subalgebra of \(S(\mathfrak{g})\), and the symbol map Equation (A.587) identifies the associated graded algebra of the centre with it,
\begin{equation}\tag{A.615} \bigoplus_{n \ge 0} \frac{\left(Z\left(U(\mathfrak{g})\right) \cap U_{n}\right) + U_{n-1}}{U_{n-1}} \;\cong\; S(\mathfrak{g})^{\mathfrak{g}} \end{equation}as graded algebras.
Rests on Proposition A.327, Definition A.325 and Lemma A.347.
Derives Proposition A.348. (1) By Proposition A.327 the map \(\omega\) is a linear isomorphism intertwining the two adjoint actions, so it carries the invariants of \(S(\mathfrak{g})\) bijectively onto the invariants of \(U(\mathfrak{g})\), and the latter are the centre by Lemma A.347.
(2) The action of \(\ad_{X}\) on \(S(\mathfrak{g})\) is by a derivation that preserves degree, so an invariant decomposes into invariant homogeneous pieces and \(S(\mathfrak{g})^{\mathfrak{g}}\) is graded; it is a subalgebra because a derivation annihilating two elements annihilates their product.
For Equation (A.615), note first that \(\ad_{X}\) maps \(U_{n}\) into \(U_{n}\) — by the Leibniz rule it replaces one letter of a word of length \(n\) by \(\comm{X}{\cdot}\) of it, which is again a letter — and that the induced map on \(U_{n}/U_{n-1} \cong S^{n}(\mathfrak{g})\) is exactly the derivation \(\ad_{X}\) of \(S^{n}(\mathfrak{g})\). Hence the symbol of a central element of \(U_{n}\) is an invariant of \(S^{n}(\mathfrak{g})\), which gives the inclusion “\(\subseteq\)” in Equation (A.615). Conversely, if \(p \in S^{n}(\mathfrak{g})^{\mathfrak{g}}\) then \(\omega(p)\) lies in \(U_{n}\), is central by part (1), and has symbol \(p\) by Proposition A.327(1); so the inclusion is an equality. The identification is multiplicative because the symbol map is (Definition A.325).
∎Suppose \(S(\mathfrak{g})^{\mathfrak{g}}\) is a free polynomial algebra on homogeneous elements \(p_{1},\ldots,p_{\ell}\) of degrees \(d_{1},\ldots,d_{\ell}\). Put \(C^{(i)} = \omega(p_{i}) \in Z\left(U(\mathfrak{g})\right) \cap U_{d_{i}}\). Then the \(C^{(i)}\) are algebraically independent and generate \(Z\left(U(\mathfrak{g})\right)\) as an algebra, so Equation (A.614) holds. Rests on Proposition A.348 and Definition A.325.
Derives Lemma A.349. Give the monomial \(\left(C^{(1)}\right)^{\alpha_{1}}\cdots \left(C^{(\ell)}\right)^{\alpha_{\ell}}\) the weight \(\sum_{i}\alpha_{i}d_{i}\); by Equation (A.569) a monomial of weight \(N\) lies in \(U_{N}\), and by Proposition A.348 its symbol in \(S^{N}(\mathfrak{g})\) is the corresponding monomial \(p^{\alpha} = p_{1}^{\alpha_{1}}\cdots p_{\ell}^{\alpha_{\ell}}\).
Independence. Let \(F\) be a nonzero polynomial in \(\ell\) variables and let \(N\) be the largest weight carried by a monomial occurring in it. Then \(F\left(C^{(1)},\ldots,C^{(\ell)}\right) \in U_{N}\) and its symbol is \(\sum_{\text{weight }\alpha = N}\lambda_{\alpha}p^{\alpha}\), a nonzero element of \(S^{N}(\mathfrak{g})\) because the \(p_{i}\) are algebraically independent. A nonzero symbol means the element is not in \(U_{N-1}\); in particular it is not zero.
Generation. Let \(z\) be central and induct on the smallest \(n\) with \(z \in U_{n}\). Its symbol is a homogeneous invariant of degree \(n\), hence a weighted-homogeneous polynomial \(Q\) of weight \(n\) in the \(p_{i}\); then \(z - Q\left(C^{(1)},\ldots,C^{(\ell)}\right)\) is central and lies in \(U_{n-1}\), and the induction hypothesis applies. The base \(n = 0\) is \(z \in \mathbb{K}\).
∎The Killing form, being nondegenerate and invariant, is an isomorphism \(\mathfrak{g} \to \mathfrak{g}^{*}\) of \(\mathfrak{g}\)-modules, and hence induces an isomorphism of graded algebras
the right-hand side being the algebra of polynomial functions on \(\mathfrak{g}\) annihilated by the natural action of every \(\ad_{X}\). In components, an element of \(S^{r}(\mathfrak{g})^{\mathfrak{g}}\) is exactly a totally symmetric invariant tensor \(k^{a_{1}\cdots a_{r}}\) in the sense of Definition 14.14, and the corresponding Casimir element is Equation (14.19). Rests on Definition 14.10, Definition 14.14 and Theorem A.330.
Derives Lemma A.350. The map \(X \mapsto \kappa(X,\cdot)\) is injective because \(\kappa\) is nondegenerate (Theorem A.330) and hence bijective by dimension; it intertwines the adjoint and coadjoint actions precisely because \(\kappa\) is invariant, Equation (14.16). An isomorphism of modules induces one of symmetric algebras, and it carries invariants to invariants. The component statement is Remark 14.15: writing an element of \(S^{r}(\mathfrak{g})\) as \(k^{a_{1}\cdots a_{r}}T_{a_{1}}\cdots T_{a_{r}}\) with \(k\) totally symmetric, the vanishing of \(\ad_{T_{c}}\) on it is Equation (14.18) term by term, and Corollary A.328 identifies \(\omega\) of that element with the Casimir Equation (14.19).
∎So the first link of Equation (A.613) is established, in the strong form of Lemma A.349: the enveloping algebra contributes nothing of its own to the count. Counting Casimir operators is counting the algebraically independent invariant polynomials on \(\mathfrak{g}\), and that is a question about the adjoint action alone.
The two quoted theorems
Let \(\mathfrak{g}\) be a semisimple Lie algebra over an algebraically closed field of characteristic zero, \(\mathfrak{h} \subseteq \mathfrak{g}\) a Cartan subalgebra (Definition 14.20), and \(W\) the Weyl group of \(\mathfrak{g}\) acting on \(\mathfrak{h}\). Then restriction of functions from \(\mathfrak{g}\) to \(\mathfrak{h}\),
is an isomorphism of graded algebras. Rests on Definition 14.20 and Lemma A.350.
Let \(\mathbb{V}\) be a vector space of dimension \(\ell\) over a field of characteristic zero and let \(W\) be a finite group of linear transformations of \(\mathbb{V}\) generated by reflections — elements fixing a hyperplane pointwise. Then the algebra of \(W\)-invariant polynomial functions on \(\mathbb{V}\) is a free polynomial algebra,
on exactly \(\ell\) algebraically independent homogeneous generators \(f_{1},\ldots,f_{\ell}\), whose degrees are determined by \(W\). Moreover, for \(\mathfrak{g}\) semisimple with Cartan subalgebra \(\mathfrak{h}\) of dimension \(\ell\), the Weyl group \(W\) is a finite group acting faithfully on \(\mathfrak{h}\) and generated by the reflections in the root hyperplanes, so Equation (A.618) applies to it with \(\mathbb{V} = \mathfrak{h}\). Rests on Definition 14.20.
The free-polynomial statement is [Chevalley:1955]. That the Weyl group is a finite reflection group of rank \(\ell\) acting faithfully on \(\mathfrak{h}\) is standard structure theory and is quoted with it; no entry of this bibliography carries that half.
Two theorems are assumed and not proved, and it is worth being exact about what they are and why they are not proved here. Theorem A.351 and the last sentence of Theorem A.352 both belong to the root-space theory of semisimple Lie algebras: the decomposition \(\mathfrak{g} = \mathfrak{h} \oplus \bigoplus_{\alpha} \mathfrak{g}_{\alpha}\) into eigenspaces of \(\mathfrak{h}\), the root system it produces, the conjugacy of Cartan subalgebras, and the Weyl group generated by root reflections. That apparatus is the content of several chapters of a book on Lie algebras and this treatise does not build it: Lie Groups, Lie Algebras, and Fibre Bundles constructs the algebras it needs by hand and computes their ranks directly, as Proposition 14.22 does for \(\mathfrak{so}(p,q)\). Equation (A.618) itself is a theorem of invariant theory, not of Lie theory, and is likewise quoted.
What is not quoted is the first link of Equation (A.613), which is the part that could plausibly hide a circularity — it is proved in The centre of the enveloping algebra and the invariant polynomials from the Poincaré–Birkhoff–Witt theorem of The Poincaré–Birkhoff–Witt Theorem, itself proved from nothing but the Jacobi identity; and it is the part that fails without semisimplicity, which is why Remark 14.23 is careful to exclude the Poincaré algebra. Nor is the arithmetic quoted: the count in The count and the check in The count checked against the chapter's algebras are carried out here.
The count
Proof of Theorem A.346. Derives Theorem A.346. Chain the identifications. By Lemma A.350, \(S(\mathfrak{g})^{\mathfrak{g}} \cong S\left(\mathfrak{g}^{*}\right)^{\mathfrak{g}}\) as graded algebras; by Theorem A.351 the latter is \(S\left(\mathfrak{h}^{*}\right)^{W}\); and by Theorem A.352 that is a free polynomial algebra on \(\ell = \dim\mathfrak{h}\) homogeneous generators \(f_{1},\ldots,f_{\ell}\), where \(\ell\) is the rank of \(\mathfrak{g}\) by Definition 14.20. Transporting the \(f_{i}\) back gives homogeneous \(p_{1},\ldots,p_{\ell} \in S(\mathfrak{g})^{\mathfrak{g}}\) that are algebraically independent and generate, and Lemma A.349 converts them into Casimir elements \(C^{(i)} = \omega(p_{i})\) with the properties claimed in Equation (A.614).
No fewer than \(\ell\). A commutative algebra generated by \(m\) elements has transcendence degree at most \(m\) over \(\mathbb{K}\), since every element is a polynomial in the generators and any \(m+1\) elements of such an algebra are algebraically dependent. By Equation (A.614) the centre has transcendence degree exactly \(\ell\), the \(C^{(i)}\) being algebraically independent. Hence no \(m < \ell\) elements of the centre generate it, and no \(m < \ell\) Casimir eigenvalues can carry the information that all of them do.
∎Let \(\mathfrak{g}\) be semisimple over an algebraically closed field of characteristic zero. Two finite-dimensional irreducible representations of \(\mathfrak{g}\) on which every element of \(Z\left(U(\mathfrak{g})\right)\) takes the same scalar value are equivalent. Equivalently, the \(\ell\) numbers
determine the equivalence class of a finite-dimensional irreducible representation. Rests on Definition 14.8 and Corollary 14.18.
Theorem A.346 says how many independent invariants there are; Theorem A.354 says that their values are enough to tell two irreducible representations apart. The two are different assertions and the second does not follow from the first, which is why the chapter's phrase “a complete set of invariant labels” has been split here. Remark 14.19 makes the same point from the other side: one Casimir separates the irreducible representations of \(\mathfrak{su}(2)\) and does not separate those of \(\mathfrak{su}(3)\), and the reason is not a failure of Theorem A.346 but the fact that \(\mathfrak{su}(3)\) has rank \(2\) — one label is simply not the whole set.
Theorem A.346 is stated over an algebraically closed field, but the algebras of Lie Groups, Lie Algebras, and Fibre Bundles are real. The transition is harmless and worth stating once. For a real semisimple \(\mathfrak{g}\) one has \(U\left(\mathfrak{g}\right)\otimes_{\R}\C = U\left(\mathfrak{g}\otimes_{\R}\C\right)\), because both sides are generated by \(\mathfrak{g}\otimes\C\) with the same relations Equation (14.12), and taking the centre commutes with this extension: an element is central if and only if its real and imaginary parts are. So the number of independent Casimir elements of a real form equals that of its complexification, and the rank meant throughout is that of the complexification. This is what Proposition 14.22 in fact computes: its proof passes to \(\mathfrak{h}_{\C}\) at once and counts the eigenvalue pairs \(\pm\alpha_{i}\) there, which is the only sensible reading, since the adjoint action of \(J_{12}\) on the compact algebra \(\mathfrak{so}(3)\) has eigenvalues \(\pm\ii\) and is not diagonalizable over \(\R\) at all.
The count checked against the chapter's algebras
Any single generator spans a maximal abelian subalgebra of \(\mathfrak{su}(2) \cong \mathfrak{so}(3)\), whose complexification is \(\mathfrak{sl}(2,\C)\); the rank is \(\ell = 1\), in agreement with Equation (14.22) at \(D = 3\). Theorem A.346 then predicts a single Casimir generator, and the invariant found in Definition 14.40 is \(\vect{J}^{2}\), of degree two — the free generator being \(f_{1}\) of degree \(2\). That it generates everything is visible in Theorem 14.43: the eigenvalue \(j(j+1)\) is injective in \(j \ge 0\), so by Theorem A.354 nothing further is needed, and by Theorem A.346 nothing further exists. Rests on Theorems 14.43 and A.346.
The diagonal traceless matrices give a maximal abelian subalgebra of dimension \(2\) (Proposition 14.50), so \(\ell = 2\) and Theorem A.346 predicts exactly two independent Casimir elements. They are found in Proposition 14.54: one quadratic, built from \(\kappa^{ab}\), and one cubic, built from the totally symmetric tensor \(d^{abc}\) of Equation (14.70) — degrees \(2\) and \(3\), which are the degrees Chevalley's theorem assigns to the invariants of the Weyl group of \(\mathfrak{su}(3)\). The count is not idle: by Proposition 14.55 the fundamental and antifundamental representations share the value of the quadratic invariant and are told apart only by the cubic one. One label would have failed; three would have been one too many. Rests on Theorem A.346, Proposition 14.54 and Proposition 14.55.
By Corollary A.344 the algebra \(\mathfrak{so}(3,1)\) is semisimple, and by Equation (14.22) its rank is \(\lfloor 4/2 \rfloor = 2\). Theorem A.346 therefore gives exactly two independent Casimir operators for the Lorentz algebra of the observed spacetime, and they are the two familiar quadratics,
the second existing only because \(D = 4\) makes the Levi-Civita symbol carry exactly four indices — the same accident of dimension that Remark 14.67 records for the Pauli–Lubanski vector. Both are invariant by Lemma 14.62 and Corollary 14.12, and by Theorem A.346 there is no third. Rests on Theorem A.346, Corollary A.344 and Equation (14.22).
Theorem 14.65 finds two invariants for the Poincaré algebra in \(3+1\) dimensions, \(P^{2}\) and \(W^{2}\), and it is tempting to read that as Theorem A.346 applied to a rank-two algebra. It is not, and the temptation must be resisted for a reason that is now provable rather than asserted: the translations form an abelian ideal, so by Proposition A.334 the Killing form of the Poincaré algebra is degenerate, the algebra is not semisimple, and every step of The centre of the enveloping algebra and the invariant polynomials — which inverts \(\kappa\) in Lemma A.350 — fails at once. That the two counts agree at \(D = 4\) is the coincidence recorded in Remark 14.23, and Equation (14.107) shows the agreement breaking in odd dimension, where \(\lceil D/2 \rceil \neq \lfloor D/2 \rfloor\). The Galilei algebra of Example 14.80 makes the point sharper still: its invariants include the central mass, which is not an invariant of any semisimple algebra at all.
Where Theorem A.346 is used for an inhomogeneous algebra, it is applied to a semisimple subalgebra and not to the whole: the count Equation (14.107) applies it to the little algebra \(\mathfrak{so}(D-1)\) of Lemma 14.68, which is semisimple for \(D-1 \ge 3\) by Corollary A.344, and adds the mass by hand.
Racah's Theorem on the Number of Casimir Operators discharges the proof obligation of Theorem 14.21, modulo the two theorems of structure theory named in Remark A.353 and the separation theorem Theorem A.354. It is used at Proposition 14.22 to turn the rank \(\lfloor D/2 \rfloor\) into a number of Lorentz invariants, at Remark 14.67 to count the labels of a massive representation in general dimension, and throughout Section 14.2 whenever a representation is labelled by its Casimir eigenvalues. Its physical content is stated in Remark 14.24: the labels are pure numbers, and the measured quantity is the label multiplied by the power of \(\hbar\) that carries its SI dimension.
Covering Spaces and the Fundamental Group of the Rotation Group
This appendix proves Corollary 14.38 of Lie Groups, Lie Algebras, and Fibre Bundles: that \(\pi_{1}\left(\SO(3,\R)\right) \cong \Z_{2}\), and more generally that when a simply connected topological group covers another one, the fundamental group of the base is the kernel of the covering homomorphism. Theorem 14.37 already supplies everything algebraic — that \(\Phi : \SU(2) \rightarrow \SO(3,\R)\) is a surjective homomorphism with kernel \(\set{\identity,-\identity}\) — and Proposition 14.33 supplies the topology of the total space, that \(\SU(2) \cong S^{3}\) is compact, connected and simply connected. What is missing, and is built here, is the bridge between the two: the covering-space machinery that converts a discrete kernel upstairs into a fundamental group downstairs.
Only the material of Topological and Metric Spaces is assumed: open sets, continuity, connectedness, compactness (Theorem 6.11), and simple connectedness (Definition 6.17). That chapter already carries the technique in its simplest instance. The continuous argument along a path of Lemma 6.19 is precisely the unique lifting of a path through the covering \(\R \rightarrow \C\setminus\set{0}\), \(\theta \mapsto \ee^{\ii\theta}\) composed with a radius, and the local-constancy argument of Proposition 6.23 is the homotopy invariance of the lift in that instance. The two lemmas below are those two arguments written for an arbitrary covering; the reader who has followed Section 6.1.5 has seen the whole idea already, and what is added here is generality, not a new device.
Statement
A continuous surjection \(p : E \longrightarrow B\) of topological spaces is a covering map if every \(b \in B\) has an open neighbourhood \(V\) that is evenly covered: \(p^{-1}(V)\) is a union of pairwise disjoint open sets \(\set{W_{\alpha}}\) — the sheets over \(V\) — each of which \(p\) maps homeomorphically onto \(V\). The set \(p^{-1}(b)\) is the fibre over \(b\); a lift of a continuous map \(g : X \longrightarrow B\) is a continuous \(\tilde{g} : X \longrightarrow E\) with \(p \circ \tilde{g} = g\). Rests on Definitions 6.2 and 6.6.
Let \(X\) be a topological space and \(x_{0} \in X\). A loop at \(x_{0}\) is a continuous \(\gamma : [0,1] \longrightarrow X\) with \(\gamma(0) = \gamma(1) = x_{0}\). Two loops \(\gamma_{0}, \gamma_{1}\) at \(x_{0}\) are path homotopic, written \(\gamma_{0} \simeq \gamma_{1}\), if there is a continuous \(H : [0,1]\times[0,1] \longrightarrow X\) with
The concatenation \(\gamma \cdot \delta\) is the loop equal to \(\gamma(2s)\) for \(s \le 1/2\) and to \(\delta(2s-1)\) for \(s \ge 1/2\), and the reverse is \(\bar{\gamma}(s) = \gamma(1-s)\). The set of path-homotopy classes of loops at \(x_{0}\) with the product \([\gamma][\delta] = [\gamma\cdot\delta]\) is the fundamental group \(\pi_{1}(X,x_{0})\). A path-connected \(X\) is simply connected in the sense of Definition 6.17 exactly when \(\pi_{1}(X,x_{0})\) is trivial. Rests on Definitions 6.15 and 6.17.
Let \(p : E \longrightarrow B\) be a covering map with \(E\) path-connected and simply connected, \(B\) path-connected, and fix \(e_{0} \in E\), \(b_{0} = p(e_{0})\). Then the monodromy map
where \(\tilde{\gamma}\) is the unique lift of \(\gamma\) with \(\tilde{\gamma}(0) = e_{0}\), is a well-defined bijection. Rests on Definitions A.361 and A.362.
Let \(E\) and \(B\) be topological groups, \(E\) path-connected and simply connected, and let \(p : E \longrightarrow B\) be a surjective continuous homomorphism that is also a covering map, with kernel \(N = p^{-1}(1_{B})\). Then \(N\) is a discrete central subgroup of \(E\), the monodromy map Equation (A.622) based at \(e_{0} = 1_{E}\) is a group isomorphism
and the group of deck transformations of \(p\) — the homeomorphisms \(\varphi\) of \(E\) with \(p\circ\varphi = p\) — consists exactly of the left translations by elements of \(N\), so that it too is isomorphic to \(N\) and hence to \(\pi_{1}(B,1_{B})\). Rests on Theorem A.363 and Definition A.361.
The proof occupies the rest of this section: two lifting lemmas (Two lifting lemmas), the monodromy theorem (The monodromy theorem), the group case (Covering homomorphisms and deck transformations), and the application to \(\SU(2) \rightarrow \SO(3,\R)\) (The rotation group).
Two lifting lemmas
Every lifting argument in covering-space theory is the same argument: chop the domain into pieces small enough to sit inside an evenly covered neighbourhood, lift piece by piece using the sheet that contains the value already fixed, and check that the pieces agree. The chopping is supplied by the following standard consequence of compactness, which Lemma 6.19 used in the disguised form of a partition of mesh smaller than a modulus of uniform continuity.
Let \(K \subset \R^{N}\) be closed and bounded and let \(\set{U_{i}}\) be an open cover of \(K\). There is \(\delta > 0\) such that every subset of \(K\) of diameter smaller than \(\delta\) is contained in a single \(U_{i}\). Rests on Theorem 6.11 and Definition 6.2.
Derives Lemma A.365. For each \(x \in K\) choose \(i(x)\) with \(x \in U_{i(x)}\) and then \(r(x) > 0\) with \(B_{2r(x)}(x) \subseteq U_{i(x)}\), which is possible because \(U_{i(x)}\) is open. The balls \(B_{r(x)}(x)\) cover \(K\), which is compact by Theorem 6.11, so finitely many of them suffice, say those centred at \(x_{1},\ldots,x_{k}\); put \(\delta = \min_{j} r(x_{j}) > 0\). Let \(S \subseteq K\) have diameter smaller than \(\delta\) and pick \(y \in S\). Then \(y \in B_{r(x_{j})}(x_{j})\) for some \(j\), and every \(z \in S\) satisfies
so \(S \subseteq B_{2r(x_{j})}(x_{j}) \subseteq U_{i(x_{j})}\).
∎Let \(p : E \longrightarrow B\) be a covering map, let \(\gamma : [0,1] \longrightarrow B\) be continuous, and let \(e \in E\) satisfy \(p(e) = \gamma(0)\). There is exactly one continuous \(\tilde{\gamma} : [0,1] \longrightarrow E\) with \(p\circ\tilde{\gamma} = \gamma\) and \(\tilde{\gamma}(0) = e\). Rests on Definition A.361 and Lemma A.365.
Derives Lemma A.366. Existence. Cover \(B\) by evenly covered open sets \(V\) and pull them back: the sets \(\gamma^{-1}(V)\) form an open cover of \([0,1]\), which is closed and bounded, so Lemma A.365 supplies a \(\delta > 0\) and with it a partition \(0 = t_{0} < t_{1} < \cdots < t_{M} = 1\) of mesh smaller than \(\delta\) such that each \(\gamma\left([t_{j-1},t_{j}]\right)\) lies in a single evenly covered \(V_{j}\).
Define \(\tilde{\gamma}\) inductively. Set \(\tilde{\gamma}(t_{0}) = e\). Suppose \(\tilde{\gamma}\) has been defined and is continuous on \([0,t_{j-1}]\) with \(p\circ\tilde{\gamma} = \gamma\) there. The point \(\tilde{\gamma}(t_{j-1})\) lies in \(p^{-1}(V_{j})\), hence in exactly one sheet \(W\) over \(V_{j}\), and \(p|_{W} : W \longrightarrow V_{j}\) is a homeomorphism. Put
This is continuous on \([t_{j-1},t_{j}]\), satisfies \(p\circ\tilde{\gamma} = \gamma\) there, and agrees at \(t_{j-1}\) with the value already assigned, because \(\tilde{\gamma}(t_{j-1}) \in W\) and \(p|_{W}\) is injective. Two continuous functions on the closed intervals \([0,t_{j-1}]\) and \([t_{j-1},t_{j}]\) agreeing at the common endpoint patch to a continuous function on the union, so the induction proceeds and terminates at \(t_{M} = 1\).
Uniqueness. Let \(\tilde{\gamma}\) and \(\tilde{\gamma}'\) be two lifts with the same initial value, and let \(A = \set{t \in [0,1] \mid \tilde{\gamma}(t) = \tilde{\gamma}'(t)}\), which contains \(0\). Fix \(t \in [0,1]\), let \(V\) be an evenly covered neighbourhood of \(\gamma(t)\), and let \(W, W'\) be the sheets over \(V\) containing \(\tilde{\gamma}(t)\) and \(\tilde{\gamma}'(t)\). The set \(O = \tilde{\gamma}^{-1}(W) \cap (\tilde{\gamma}')^{-1}(W')\) is an open neighbourhood of \(t\) in \([0,1]\). If \(t \in A\) then \(W = W'\), the sheets being disjoint, and on \(O\) both lifts equal \(\left(p|_{W}\right)^{-1}\circ\gamma\), so \(O \subseteq A\): the set \(A\) is open. If \(t \notin A\) then \(W \neq W'\), since a common sheet would force \(\tilde{\gamma}(t) = \tilde{\gamma}'(t)\) by injectivity of \(p|_{W}\); the sheets being disjoint, \(O\) misses \(A\) entirely, so the complement of \(A\) is open too. Thus \(A\) is a nonempty subset of \([0,1]\) that is both open and closed, and \([0,1]\) is connected (Theorem 6.14), so \(A = [0,1]\).
∎Let \(p : E \longrightarrow B\) be a covering map, let \(H : [0,1]\times[0,1] \longrightarrow B\) be continuous, and let \(e \in E\) satisfy \(p(e) = H(0,0)\). There is exactly one continuous \(\tilde{H} : [0,1]\times[0,1] \longrightarrow E\) with \(p\circ\tilde{H} = H\) and \(\tilde{H}(0,0) = e\). If moreover \(H\) is a path homotopy in the sense of Equation (A.621) — so that \(H(0,t)\) and \(H(1,t)\) are constant in \(t\) — then \(\tilde{H}(0,t)\) and \(\tilde{H}(1,t)\) are constant in \(t\) as well. Rests on Lemmas A.365 and A.366.
Derives Lemma A.367. Existence. As before, the sets \(H^{-1}(V)\) with \(V\) evenly covered form an open cover of the square \([0,1]^{2}\), which is closed and bounded in \(\R^{2}\); Lemma A.365 supplies \(\delta > 0\), and a grid \(0 = s_{0} < \cdots < s_{M} = 1\), \(0 = t_{0} < \cdots < t_{M} = 1\) of mesh smaller than \(\delta/\sqrt{2}\) makes each closed cell \(R_{jk} = [s_{j-1},s_{j}]\times[t_{k-1},t_{k}]\) have diameter smaller than \(\delta\), hence \(H(R_{jk})\) contained in a single evenly covered \(V_{jk}\).
Order the cells lexicographically, bottom row left to right, then the next row, and lift them one at a time. At each stage the part of the boundary of the current cell \(R\) on which \(\tilde{H}\) has already been defined — call it \(C\) — is the union of at most the left edge and the bottom edge of \(R\), together with the corner joining them, and is therefore connected. (For the very first cell \(C\) is the single point \((s_{0},t_{0})\), where \(\tilde{H} = e\) is prescribed.) The set \(\tilde{H}(C)\) is connected and lies in \(p^{-1}(V)\) for the evenly covered \(V \supseteq H(R)\), which is the disjoint union of open sheets; a connected subset of a disjoint union of open sets lies in one of them, so \(\tilde{H}(C) \subseteq W\) for a single sheet \(W\). Define \(\tilde{H} = \left(p|_{W}\right)^{-1}\circ H\) on \(R\). This is continuous on \(R\), projects to \(H\), and agrees on \(C\) with what was there — again because \(p|_{W}\) is injective and both values lie in \(W\). The extended function is continuous, being continuous on each of finitely many closed cells and consistent on the overlaps. After the last cell \(\tilde{H}\) is defined on the whole square.
Uniqueness. Verbatim the argument of Lemma A.366: the agreement set of two lifts is open and closed, and the square is connected, being path-connected (Proposition 6.16).
The vertical edges. Suppose \(H(0,t) = b_{0}\) for every \(t\). Then \(t \mapsto \tilde{H}(0,t)\) is a lift of the constant path at \(b_{0}\); so is the constant map \(t\mapsto \tilde{H}(0,0)\); the two agree at \(t = 0\), so by the uniqueness clause of Lemma A.366 they agree throughout. The same applies at \(s = 1\).
∎The monodromy theorem
The elementary bookkeeping behind Definition A.362 is recorded first, because it is what makes \(\pi_{1}\) a group at all and because the injectivity of \(\Psi\) uses it.
Let \(\alpha,\beta,\eta\) be paths in a space \(X\) with \(\alpha(1) = \beta(0)\) and \(\beta(1) = \eta(0)\), let \(c_{x}\) denote the constant path at \(x\), and write \(\simeq\) for homotopy with endpoints held fixed. Then
Consequently \(\pi_{1}(X,x_{0})\) of Definition A.362 is a group, with \([\gamma]^{-1} = [\bar{\gamma}]\); and a continuous \(p : E \longrightarrow B\) with \(p(e_{0}) = b_{0}\) induces a group homomorphism \(p_{*} : \pi_{1}(E,e_{0}) \longrightarrow \pi_{1}(B,b_{0})\), \(p_{*}[\sigma] = [p\circ\sigma]\). Rests on Definition A.362.
Derives Lemma A.368. All three relations of Equation (A.625) are reparametrizations. If \(\phi : [0,1] \rightarrow [0,1]\) is continuous with \(\phi(0) = 0\) and \(\phi(1) = 1\), then for any path \(\mu\) the map \(H(s,t) = \mu\left((1-t)\phi(s) + ts\right)\) is continuous, fixes the endpoints, and joins \(\mu\circ\phi\) to \(\mu\); so \(\mu\circ\phi \simeq \mu\). The two sides of the first relation are the same path \(\alpha\cdot\left(\beta\cdot\eta\right)\) composed with such a \(\phi\) — the piecewise-linear map sending \(1/4\) to \(1/2\) and \(1/2\) to \(3/4\) — and the two sides of the second are likewise related by the piecewise-linear \(\phi\) sending \(1/2\) to \(0\), respectively to \(1\). For the third, put
which is continuous (the two formulas agree at \(s = 1/2\)), equals \(\alpha\cdot\bar{\alpha}\) at \(t = 0\) and \(c_{\alpha(0)}\) at \(t = 1\), and is constant on \(s = 0\) and on \(s = 1\). That the product on classes is well defined — homotopies of the two factors concatenate to a homotopy of the product — is immediate from the definition of concatenation. Finally \(p\circ\left(\alpha\cdot\beta\right) = \left(p\circ\alpha\right)\cdot\left(p\circ\beta\right)\) by inspection, and \(p\circ H\) is a path homotopy whenever \(H\) is, so \(p_{*}\) is a well-defined homomorphism.
∎Proof of Theorem A.363. Derives Theorem A.363. \(\Psi\) is well defined. The lift \(\tilde{\gamma}\) exists and is unique by Lemma A.366, and \(p\left(\tilde{\gamma}(1)\right) = \gamma(1) = b_{0}\), so \(\tilde{\gamma}(1) \in p^{-1}(b_{0})\). That it depends only on the class \([\gamma]\) is the content of Lemma A.367: let \(H\) be a path homotopy from \(\gamma_{0}\) to \(\gamma_{1}\) and \(\tilde{H}\) its lift with \(\tilde{H}(0,0) = e_{0}\). The left edge \(t \mapsto \tilde{H}(0,t)\) is constant, so \(\tilde{H}(\cdot,t)\) is for each \(t\) the lift of \(H(\cdot,t)\) starting at \(e_{0}\); the right edge \(t\mapsto\tilde{H}(1,t)\) is constant too, so the common endpoint \(\tilde{H}(1,0) = \tilde{H}(1,1)\), that is \(\tilde{\gamma}_{0}(1) = \tilde{\gamma}_{1}(1)\).
\(\Psi\) is surjective. Let \(e \in p^{-1}(b_{0})\). Since \(E\) is path-connected there is a path \(\sigma\) in \(E\) from \(e_{0}\) to \(e\); then \(\gamma = p\circ\sigma\) is a loop at \(b_{0}\), and \(\sigma\) is its lift starting at \(e_{0}\) by uniqueness, so \(\Psi\left([\gamma]\right) = \sigma(1) = e\).
\(\Psi\) is injective. Suppose \(\tilde{\gamma}_{0}(1) = \tilde{\gamma}_{1}(1)\) for loops \(\gamma_{0},\gamma_{1}\) at \(b_{0}\). Both lifts start at \(e_{0}\), so \(\tilde{\gamma}_{0}\) and \(\tilde{\gamma}_{1}\) are paths in \(E\) with the same two endpoints, and \(\sigma = \tilde{\gamma}_{0}\cdot\overline{\tilde{\gamma}_{1}}\) is a loop in \(E\) at \(e_{0}\). Since \(E\) is simply connected, \([\sigma]\) is the identity of \(\pi_{1}(E,e_{0})\). Apply the homomorphism \(p_{*}\) of Lemma A.368:
using \(p\circ\tilde{\gamma}_{k} = \gamma_{k}\) and \(\left[\bar{\gamma}_{1}\right] = \left[\gamma_{1}\right]^{-1}\). Hence \([\gamma_{0}] = [\gamma_{1}]\).
∎Covering homomorphisms and deck transformations
Proof of Theorem A.364. Derives Theorem A.364. \(N\) is discrete and central. Let \(V\) be an evenly covered neighbourhood of \(1_{B}\) and \(W\) the sheet over it containing \(1_{E}\). Then \(W \cap N = \set{1_{E}}\), because \(p\) is injective on \(W\) and already sends \(1_{E}\) to \(1_{B}\); translating, \(nW \cap N = \set{n}\) for every \(n \in N\), so \(N\) carries the discrete topology. For centrality, fix \(n \in N\) and consider \(f : E \longrightarrow E\), \(f(x) = xnx^{-1}\). It is continuous, takes values in \(N\) — because \(p\left(xnx^{-1}\right) = p(x)\,1_{B}\,p(x)^{-1} = 1_{B}\) — and \(E\) is connected, so its image is a connected subset of a discrete space, hence the single point \(f(1_{E}) = n\). Thus \(xn = nx\) for all \(x\).
\(\Psi\) is a homomorphism. Let \(\gamma,\delta\) be loops at \(1_{B}\) with lifts \(\tilde{\gamma},\tilde{\delta}\) starting at \(1_{E}\), and put \(n = \tilde{\gamma}(1) \in N\). The path \(t \mapsto n\,\tilde{\delta}(t)\) is continuous, starts at \(n\,1_{E} = n = \tilde{\gamma}(1)\), and projects to \(p(n)\,\delta(t) = \delta(t)\). Hence the concatenation \(\tilde{\gamma}\cdot\left(n\tilde{\delta}\right)\) is a lift of \(\gamma\cdot\delta\) starting at \(1_{E}\), and by uniqueness (Lemma A.366) it is the lift. Its endpoint is \(n\,\tilde{\delta}(1)\), so
Being a bijection by Theorem A.363, \(\Psi\) is an isomorphism, which is Equation (A.623).
Deck transformations. If \(n \in N\) then left translation \(L_{n}(x) = nx\) is a homeomorphism of \(E\) with \(p\left(L_{n}(x)\right) = p(n)p(x) = p(x)\), so \(L_{n}\) is a deck transformation, and \(n \mapsto L_{n}\) is an injective homomorphism (\(L_{n} = L_{n'}\) forces \(n = n'\) at \(x = 1_{E}\)). Conversely let \(\varphi\) be any deck transformation and set \(n = \varphi(1_{E})\), which lies in \(N\) because \(p(n) = p(1_{E}) = 1_{B}\). Both \(\varphi\) and \(L_{n}\) are lifts of the map \(p : E \longrightarrow B\) through \(p\) itself, and they agree at \(1_{E}\). The agreement set of two lifts of one map on a connected domain is open and closed — the argument is word for word the uniqueness half of Lemma A.366, with \([0,1]\) replaced by \(E\) — and \(E\) is connected, so \(\varphi = L_{n}\).
∎The rotation group
It remains to check that \(\Phi\) of Equation (14.46) is a covering map. That is the one place where the group is used concretely, and it is short.
The homomorphism \(\Phi : \SU(2) \longrightarrow \SO(3,\R)\) of Theorem 14.37 is an open map and a covering map, each fibre having exactly two points. Rests on Theorem 14.37, Proposition 14.33 and Definition A.361.
Derives Lemma A.369. Write \(K = \set{\identity,-\identity} = \ker\Phi\) (Equation (14.47)). Because \(\Phi\) is a homomorphism with kernel \(K\), its fibres are exactly the pairs \(\set{U,-U}\), which have two elements since \(U \neq -U\); and \(\Phi\) is surjective by Theorem 14.37.
\(\Phi\) is open. First, \(\Phi\) is a closed map: \(\SU(2)\) is compact (Proposition 14.33) and \(\SO(3,\R)\) is Hausdorff, being a subspace of the matrices \(\R^{3\times3}\), so a closed subset of \(\SU(2)\) is compact, its continuous image is compact, and a compact subset of a Hausdorff space is closed. A continuous closed surjection is a quotient map: a set \(S \subseteq \SO(3,\R)\) with \(\Phi^{-1}(S)\) open has open complement's preimage closed, hence \(\SO(3,\R)\setminus S = \Phi\left(\SU(2)\setminus\Phi^{-1}(S)\right)\) closed — the equality holding because \(\Phi\) is surjective and \(\Phi^{-1}(S)\) is a union of fibres — so \(S\) is open. Now let \(O \subseteq \SU(2)\) be open. Then \(-O\) is open, multiplication by \(-\identity\) being a homeomorphism, and \(\Phi^{-1}\left(\Phi(O)\right) = O \cup (-O)\) because the fibre through \(U\) is \(\set{U,-U}\). That union is open, so \(\Phi(O)\) is open.
Even covering. Let \(V_{0} = \set{U \in \SU(2) \mid \norm{U - \identity} < 1}\), the norm being any norm on the \(2\times2\) matrices with \(\norm{2\identity} = 2\); for instance \(\norm{M}^{2} = \tfrac{1}{2}\tr\left(M^{\dagger}M\right)\), for which \(\norm{\identity} = 1\). Then \(V_{0} \cap (-V_{0}) = \varnothing\): if \(U\) and \(-U\) both lay within distance \(1\) of \(\identity\) then
which is absurd. Fix \(U_{0} \in \SU(2)\) and put \(W = U_{0}V_{0}\), \(V = \Phi(W)\), which is open by the previous paragraph. The two sheets are \(W\) and \(-W\): they are disjoint, since \(U_{0}v = -U_{0}v'\) would give \(v = -v'\) with \(v,v' \in V_{0}\); their union is \(\Phi^{-1}(V)\), because a point of \(\Phi^{-1}(V)\) is \(\pm w\) for some \(w \in W\); and \(\Phi\) restricted to either is a continuous open bijection onto \(V\), hence a homeomorphism. Every point of \(\SO(3,\R)\) lies in such a \(V\), namely for \(U_{0}\) any preimage of it.
∎Proof of Corollary 14.38. Derives Corollary 14.38. By Theorem 14.37 the map \(\Phi\) is a surjective homomorphism with kernel \(\set{\identity,-\identity}\), so the induced map \(\SU(2)/\set{\identity,-\identity} \longrightarrow \SO(3,\R)\) is a group isomorphism, and it is a homeomorphism because \(\Phi\) is continuous, open and surjective (Lemma A.369). By Proposition 14.33 the group \(\SU(2)\) is path-connected and simply connected, and by Lemma A.369 the map \(\Phi\) is a covering map; \(\SO(3,\R)\) is path-connected, being a continuous image of a path-connected space. Theorem A.364 therefore applies and gives
which is Equation (14.48), and identifies the deck-transformation group of the covering with the same \(\Z_{2}\), generated by \(U \mapsto -U\).
The concrete description in Corollary 14.38 is now read off the isomorphism. The path \(\phi \mapsto U(\hat{n},\phi)\), \(0\le\phi\le2\pi\), runs in \(\SU(2)\) from \(\identity\) to \(-\identity\) by Equation (14.45); it is therefore the lift, starting at \(\identity\), of the loop \(\phi \mapsto \Phi\left(U(\hat{n},\phi)\right) = R(\hat{n},\phi)\) in \(\SO(3,\R)\), whose endpoint \(R(\hat{n},2\pi) = \identity\) makes it a loop indeed. Its monodromy image is \(-\identity \neq \identity\), so its class in \(\pi_{1}\) is the nontrivial element: the loop is not contractible. Doubling the parameter range to \(0\le\phi\le4\pi\) doubles the class, and \((-\identity)^{2} = \identity\), so that loop is contractible.
∎Nothing above is special to \(\SU(2)\) until Lemma A.369: the two lifting lemmas and Theorems A.363 and A.364 hold for any covering, and the reader who compares them with Lemma 6.19, Lemma 6.22 and Proposition 6.23 will find the same three steps — lift along a partition, patch, and use connectedness of the parameter interval to rule out ambiguity — carried out there for the single covering \(\theta \mapsto \ee^{\ii\theta}\) of the unit circle. The winding number of Definition 6.20 is the monodromy map Equation (A.622) for that covering, whose total space \(\R\) is simply connected and whose kernel is \(2\pi\Z\); that is why Equation (6.5) is an integer.
The purchase is Equation (14.48), and through it Remark 14.39: the rotation group is not simply connected, its universal cover is \(\SU(2)\), and a quantum system may therefore carry a representation of \(\SU(2)\) that is only a projective representation of \(\SO(3,\R)\). Theorem 14.43 and Proposition 14.45 say which ones those are — exactly the half-integer \(j\) — so the existence of half-integer angular momentum is a corollary of Equation (A.628) and not an independent postulate.
The Structure Constants of $\mathfrak{su}(3)$
This appendix completes Proposition 14.52 of Lie Groups, Lie Algebras, and Fibre Bundles: it computes every nonvanishing component of the totally antisymmetric array \(f_{abc}\) and the totally symmetric array \(d_{abc}\) of Equation (14.67) in the Gell-Mann basis of Definition 14.51, and exhibits the two complete tables. The chapter carries two entries as specimens, \(f_{458}\) and \(d_{247}\); the remaining twenty-three are obtained here, and — more to the point — so is the statement that there are no others. Nothing beyond Definition 14.51 and Equation (14.66) is used.
The naive route is \(8^{3} = 512\) traces, or \(\binom{8+2}{3} = 120\) after the symmetries are imposed. That is not what is done below. A single observation about which products of three matrix units have a nonzero trace cuts the work to three short families, and the classification of those families is what makes the tables exhaustive rather than merely checked.
One trace for both arrays
With \(T_{a} = \tfrac{1}{2}\lambda_{a}\) as in Definition 14.51,
so that \(d_{abc} = \tfrac{1}{2}\Re\tr(\lambda_{a}\lambda_{b}\lambda_{c})\) and \(f_{abc} = \tfrac{1}{2}\Im\tr(\lambda_{a}\lambda_{b}\lambda_{c})\); in particular both arrays are real. Rests on Equation (14.68), Definition 14.51 and Equation (14.66).
Derives Lemma A.371. Substituting \(T_{a} = \tfrac{1}{2}\lambda_{a}\) into Equation (14.68), each of the three factors contributes a \(\tfrac{1}{2}\), so
Now \(\lambda_{a}\lambda_{b} = \tfrac{1}{2}\acomm{\lambda_{a}}{\lambda_{b}} + \tfrac{1}{2}\comm{\lambda_{a}}{\lambda_{b}}\), so
which is Equation (A.629). Reality: the \(\lambda_{a}\) are Hermitian, so \(\overline{\tr(\lambda_{a}\lambda_{b}\lambda_{c})} = \tr\left((\lambda_{a}\lambda_{b}\lambda_{c})^{\dagger}\right) = \tr(\lambda_{c}\lambda_{b}\lambda_{a})\), and reversing the order of three factors under the trace exchanges \(a\) and \(c\) in Equation (A.629); by the total symmetry of \(d\) and total antisymmetry of \(f\) established in Proposition 14.52, that conjugates the right-hand side. Hence \(d\) and \(f\) are the real and imaginary parts of one complex number, and both are real numbers.
∎Proposition 14.52 proves that \(f_{abc}\) is totally antisymmetric and \(d_{abc}\) totally symmetric in \((a,b,c)\). Those two rules fix every component from the tables below, and it is worth saying exactly how many components each independent entry generates.
A nonzero \(f_{abc}\) necessarily has three distinct indices — \(f_{aab} = -f_{aab} = 0\) — and its six permutations give six nonzero components, three equal to \(+f_{abc}\) (the cyclic ones) and three equal to \(-f_{abc}\). With nine independent entries that is \(9\times6 = 54\) nonvanishing components of \(f\).
For \(d\) the count depends on the repetition pattern. An entry with three distinct indices generates six equal components; an entry of the shape \(d_{aab}\) with \(a\neq b\) generates three; and \(d_{aaa}\) generates one. Of the sixteen independent entries below, four have distinct indices, eleven have exactly one repetition, and one is \(d_{888}\); so \(d\) has \(4\times6 + 11\times3 + 1 = 58\) nonvanishing components.
Ordering within an entry is therefore free, and the tables list each independent entry once, in the index order used in Equations (14.69) and (14.70).
Which triples can be nonzero at all
Write \(E_{jk}\) for the \(3\times3\) matrix whose only nonzero entry is a \(1\) in row \(j\) and column \(k\), so that \(E_{jk}E_{lm} = \delta_{kl}E_{jm}\) and \(\tr E_{jk} = \delta_{jk}\). Reading off Equations (14.63), (14.64) and (14.65), each Gell-Mann matrix is supported on one of four index sets:
Read as a graph on the vertex set \(\set{1,2,3}\), the first row lives on the edge \(1\)–\(2\), the second on \(1\)–\(3\), the third on \(2\)–\(3\), and the fourth — \(\lambda_{3}\) and \(\lambda_{8}\), the diagonal ones — on loops. Call \(\mathrm{A} = \set{1,2}\), \(\mathrm{B} = \set{4,5}\), \(\mathrm{C} = \set{6,7}\) the three off-diagonal pairs of generator labels and \(\mathrm{D} = \set{3,8}\) the diagonal pair.
\(\tr\left(\lambda_{a}\lambda_{b}\lambda_{c}\right) = 0\) unless the unordered triple \(\set{a,b,c}\) falls into one of the three families
-
all three indices in \(\mathrm{D}\);
-
exactly one index in \(\mathrm{D}\) and the other two in a single off-diagonal pair;
-
one index in each of \(\mathrm{A}\), \(\mathrm{B}\) and \(\mathrm{C}\).
Derives Lemma A.373. Expanding each factor in the matrix units of Equation (A.631),
so a nonzero contribution needs a closed walk \(i \to j \to k \to i\) of length three in the index set \(\set{1,2,3}\), whose three steps are supplied by \(\lambda_{a}\), \(\lambda_{b}\), \(\lambda_{c}\) in that order. A \(\lambda\) with index in \(\mathrm{D}\) supplies only loops \(i \to i\); a \(\lambda\) in an off-diagonal pair supplies only the two steps along its own edge, and along no other.
Count the loops among the three steps. If all three steps are loops, all three indices lie in \(\mathrm{D}\): family (1). If exactly two are loops, the third is an edge step, which cannot close a walk that has returned to its starting vertex twice — a walk \(i\to i\to i\to j\) ends at \(j \neq i\) — so this case contributes nothing. If exactly one step is a loop, the other two are edge steps \(u \to v\) and \(v \to u\) on one and the same edge, since the walk must return; hence one index in \(\mathrm{D}\) and two in a single pair, which is family (2). If no step is a loop, the walk visits three vertices along three distinct edges of the triangle on \(\set{1,2,3}\) — two of the three steps on one edge would force the third to be a loop — so one index comes from each of \(\mathrm{A}\), \(\mathrm{B}\), \(\mathrm{C}\): family (3).
∎Everything is now a finite enumeration inside three small families, and the products needed are the ten diagonal matrices
each read straight off Equation (A.631): for instance \(\lambda_{6}\lambda_{7} = \left(E_{23}+E_{32}\right)\left(-\ii E_{23}+\ii E_{32}\right) = \ii E_{23}E_{32} - \ii E_{32}E_{23} = \ii E_{22} - \ii E_{33}\).
Family (1): three diagonal indices
Here \(\lambda_{3}\) and \(\lambda_{8}\) commute, so every commutator vanishes and \(f_{abc} = 0\) throughout the family. For \(d\) the four independent triples are \(333\), \(338\), \(388\), \(888\), and Equation (A.629) needs only a trace of a product of diagonal matrices:
Halving each, the family contributes
Family (2): one diagonal index and one off-diagonal pair
Put the diagonal generator last, so that the trace is that of \(\left(\lambda_{a}\lambda_{b}\right)\lambda_{\delta}\) with \(a,b\) in one pair and \(\delta \in \set{3,8}\); the first factor is one of the six diagonal matrices Equation (A.633). There are \(3\times3\times2 = 18\) independent triples — three pairs, three choices of \((a,b)\) up to the symmetries, two diagonal generators — and each is one line. Writing \(\lambda_{3} = \diag(1,-1,0)\) and \(\lambda_{8} = \tfrac{1}{\sqrt3}\diag(1,1,-2)\):
Pair \(\mathrm{A} = \set{1,2}\), on \(\diag(\cdot,\cdot,0)\).
The second line covers \(\lambda_{2}\lambda_{2}\) as well, since \(\lambda_{2}^{2} = \lambda_{1}^{2}\) by Equation (A.633).
Pair \(\mathrm{B} = \set{4,5}\), on \(\diag(\cdot,0,\cdot)\).
The last line reproduces the specimen computed in Proposition 14.52.
Pair \(\mathrm{C} = \set{6,7}\), on \(\diag(0,\cdot,\cdot)\).
Every entry follows from Equation (A.629) by halving the real part for \(d\) and the imaginary part for \(f\); the index order has been rearranged into the one used in Equations (14.69) and (14.70), which is legitimate by Remark A.372 — note that \(f_{367} = f_{673}\) and \(f_{345} = f_{453}\), both cyclic and therefore sign-preserving.
Family (3): one index from each pair
This family is eight triples \((a,b,c)\) with \(a \in \set{1,2}\), \(b \in \set{4,5}\), \(c \in \set{6,7}\), and all eight fall out of a single observation: in Equation (A.632) only one closed walk survives. The first factor lies on the edge \(1\)–\(2\), the second on \(1\)–\(3\), the third on \(2\)–\(3\); a walk \(i \to j \to k \to i\) must therefore start with a step of the first edge, and the only choice compatible with the second step lying on \(1\)–\(3\) is \(j = 1\), hence \(i = 2\), \(k = 3\). So
and Equation (A.631) supplies the three factors:
Multiplying out the eight products of Equation (A.636) and halving real and imaginary parts as in Equation (A.629):
the unlisted partner of each entry vanishing. The value \(d_{247} = -\tfrac{1}{2}\) reproduces the second specimen of Proposition 14.52.
The tables
Collecting Equation (A.634), the three blocks of Family (2): one diagonal index and one off-diagonal pair and Equation (A.637) gives Tables A.2 and A.3. By Lemma A.373 no triple outside the three families can contribute, and inside them the enumeration was exhaustive; the tables are therefore complete, and reproduce Equations (14.69) and (14.70).
| $abc$ | $f_{abc}$ | family | generating trace |
|---|---|---|---|
| $123$ | $1$ | (2), pair $\set{1,2}$ | $\tr(\lambda_{1}\lambda_{2}\lambda_{3}) = 2\ii$ |
| $147$ | $\tfrac{1}{2}$ | (3) | $\tr(\lambda_{1}\lambda_{4}\lambda_{7}) = \ii$ |
| $156$ | $-\tfrac{1}{2}$ | (3) | $\tr(\lambda_{1}\lambda_{5}\lambda_{6}) = -\ii$ |
| $246$ | $\tfrac{1}{2}$ | (3) | $\tr(\lambda_{2}\lambda_{4}\lambda_{6}) = \ii$ |
| $257$ | $\tfrac{1}{2}$ | (3) | $\tr(\lambda_{2}\lambda_{5}\lambda_{7}) = \ii$ |
| $345$ | $\tfrac{1}{2}$ | (2), pair $\set{4,5}$ | $\tr(\lambda_{4}\lambda_{5}\lambda_{3}) = \ii$ |
| $367$ | $-\tfrac{1}{2}$ | (2), pair $\set{6,7}$ | $\tr(\lambda_{6}\lambda_{7}\lambda_{3}) = -\ii$ |
| $458$ | $\tfrac{\sqrt3}{2}$ | (2), pair $\set{4,5}$ | $\tr(\lambda_{4}\lambda_{5}\lambda_{8}) = \ii\sqrt3$ |
| $678$ | $\tfrac{\sqrt3}{2}$ | (2), pair $\set{6,7}$ | $\tr(\lambda_{6}\lambda_{7}\lambda_{8}) = \ii\sqrt3$ |
| $abc$ | $d_{abc}$ | $abc$ | $d_{abc}$ |
|---|---|---|---|
| Family (1): indices in $\set{3,8}$ | |||
| $338$ | $\tfrac{1}{\sqrt3}$ | $888$ | $-\tfrac{1}{\sqrt3}$ |
| Family (2): one index in $\set{3,8}$, two in one pair | |||
| $118$ | $\tfrac{1}{\sqrt3}$ | $448$ | $-\tfrac{1}{2\sqrt3}$ |
| $228$ | $\tfrac{1}{\sqrt3}$ | $558$ | $-\tfrac{1}{2\sqrt3}$ |
| $344$ | $\tfrac{1}{2}$ | $668$ | $-\tfrac{1}{2\sqrt3}$ |
| $355$ | $\tfrac{1}{2}$ | $778$ | $-\tfrac{1}{2\sqrt3}$ |
| $366$ | $-\tfrac{1}{2}$ | $377$ | $-\tfrac{1}{2}$ |
| Family (3): one index from each pair | |||
| $146$ | $\tfrac{1}{2}$ | $247$ | $-\tfrac{1}{2}$ |
| $157$ | $\tfrac{1}{2}$ | $256$ | $\tfrac{1}{2}$ |
Two entries of Equation (14.70) deserve a word, because the chapter lists them in an order that hides where they come from. \(d_{338} = \tfrac{1}{\sqrt3}\) is a family-(1) entry, computed from three diagonal matrices, while \(d_{118}\) and \(d_{228}\) are family-(2) entries computed from \(\lambda_{1}^{2} = \lambda_{2}^{2} = \diag(1,1,0)\); the three come out equal because \(\lambda_{3}^{2}\) equals that same matrix. That coincidence is the reason the chapter can write \(d_{118} = d_{228} = d_{338}\) in one line.
A Jacobi check
The tables are a computation, and a computation deserves an independent test. The Jacobi identity for the brackets Equation (14.67) is such a test, and it is nontrivial on a triple that involves several different entries at once. Take \((T_{1},T_{4},T_{5})\):
The first inner bracket has two nonzero components, which is the trap: \(\comm{T_{4}}{T_{5}} = \ii f_{45c}T_{c} = \ii\left(\tfrac{1}{2}T_{3} + \tfrac{\sqrt3}{2}T_{8}\right)\), by \(f_{453} = f_{345} = \tfrac{1}{2}\) and \(f_{458} = \tfrac{\sqrt3}{2}\) from Table A.2. Since \(f_{18c} = 0\) for every \(c\) — no entry of Table A.2 contains both \(1\) and \(8\) — the \(T_{8}\) piece drops, and with \(f_{132} = -f_{123} = -1\),
For the second term, \(f_{516} = -f_{156} = \tfrac{1}{2}\) gives \(\comm{T_{5}}{T_{1}} = \tfrac{\ii}{2}T_{6}\), and \(f_{462} = f_{246} = \tfrac{1}{2}\) gives
For the third, \(f_{147} = \tfrac{1}{2}\) gives \(\comm{T_{1}}{T_{4}} = \tfrac{\ii}{2}T_{7}\), and \(f_{572} = f_{257} = \tfrac{1}{2}\) gives
The three add to \(\left(\tfrac{1}{2} - \tfrac{1}{4} - \tfrac{1}{4}\right)T_{2} = 0\), so Equation (A.638) holds. The check consumes \(f_{123}\), \(f_{147}\), \(f_{156}\), \(f_{246}\), \(f_{257}\), \(f_{345}\) and \(f_{458}\) — seven of the nine independent entries, drawn from both families that carry a nonzero \(f\) — and it fails if any one of them is altered. Had the \(f_{453}\) contribution to \(\comm{T_{4}}{T_{5}}\) been overlooked, leaving only the \(f_{458}\) piece that the chapter's specimen computes, the identity would have returned \(-\tfrac{1}{2}T_{2}\) instead of zero.
Tables A.2 and A.3 discharge the tabulation asserted in Proposition 14.52 and used throughout Lie Groups, Lie Algebras, and Fibre Bundles: in Proposition 14.53, where the entry \(\kappa_{33} = 3\) is checked against the \(f_{3cd}\) read off Table A.2; in Proposition 14.54, whose cubic Casimir \(C_{3} = d_{abc}T_{a}T_{b}T_{c}\) is built from Table A.3 and whose invariance identity Equation (14.76) couples the two tables; and in Proposition 14.55, where the adjoint generators are \(\left(T_{a}^{\mathrm{ad}}\right)_{bc} = -\ii f_{abc}\). The completeness established by Lemma A.373 is what licenses the sums over \(c\) and \(d\) in those places to be evaluated by listing the entries of the tables and stopping.
Simplicity of $\mathfrak{su}(3)$
This appendix proves what Lie Groups, Lie Algebras, and Fibre Bundles quotes in two places: that \(\mathfrak{su}(3)\) is a simple Lie algebra. The statement is used in Proposition 14.53, where the nondegeneracy of the Killing form is announced as the semisimplicity demanded by Theorem 14.13, and — more substantially — in Theorem 14.57, whose proof identifies the eight-dimensional summand of \(\vect{3}\otimes\bar{\vect{3}}\) with the adjoint representation and calls it irreducible because an invariant subspace would be an ideal of a simple Lie algebra. That last inference is what is discharged here.
The proof is the standard one and is short: complexify, split a putative ideal into eigenspaces of the two-dimensional Cartan subalgebra, and then walk around the root hexagon of Equation (14.80), using the fact that any two roots at \(120\) degrees add to a third. A single root vector inside the ideal therefore drags in all six, and with them the Cartan directions.
Two remarks on what is not used, because both matter for circularity. First, nothing below quotes the classification of simple Lie algebras; only Definition 14.51 and the multiplication table of the \(3\times3\) matrix units enter. Second, nothing below uses Proposition 14.53. That is deliberate: the chapter computes the Killing form of \(\mathfrak{su}(3)\) directly, precisely so as not to lean on the uniqueness of an invariant form on a simple algebra, and the present appendix repays the compliment by not leaning on the Killing form.
Statement
A Lie algebra \(\mathfrak{g}\) over a field is simple if it is not abelian and its only ideals are \(\set{0}\) and \(\mathfrak{g}\) itself; an ideal is a subspace \(\mathfrak{i}\) with \(\comm{\mathfrak{g}}{\mathfrak{i}} \subseteq \mathfrak{i}\). Rests on Definition 14.49.
The real Lie algebra \(\mathfrak{su}(3)\) of Definition 14.49 is simple, and so is its complexification \(\mathfrak{sl}(3,\C)\), the algebra of traceless complex \(3\times3\) matrices. Rests on Definition 14.49, Definition A.375 and Proposition 14.50.
The adjoint representation of \(\mathfrak{su}(3)\) on itself, and equally the representation \(M \mapsto UMU^{\dagger}\) of \(\SU(3)\) on the traceless complex \(3\times3\) matrices, has no invariant subspace other than \(\set{0}\) and the whole space. Rests on Theorem A.376 and Proposition 14.55.
Complexification
As a real vector space,
the sum being direct. If \(\mathfrak{i} \subseteq \mathfrak{su}(3)\) is an ideal of the real algebra, then \(\mathfrak{i}_{\C} = \mathfrak{i} + \ii\,\mathfrak{i}\) is a complex ideal of \(\mathfrak{sl}(3,\C)\) with \(\dim_{\C}\mathfrak{i}_{\C} = \dim_{\R}\mathfrak{i}\). Rests on Definition 14.49 and Proposition 14.50.
Derives Lemma A.378. The splitting. Let \(X\) be traceless. Put \(X_{-} = \tfrac{1}{2}\left(X - X^{\dagger}\right)\) and \(X_{+} = \tfrac{1}{2}\left(X + X^{\dagger}\right)\), so that \(X = X_{-} + X_{+}\). Both are traceless, since \(\tr X^{\dagger} = \overline{\tr X} = 0\); \(X_{-}\) is anti-Hermitian, hence in \(\mathfrak{su}(3)\) by Definition 14.49; and \(X_{+}\) is Hermitian, so \(-\ii X_{+}\) is anti-Hermitian and \(X_{+} = \ii\left(-\ii X_{+}\right) \in \ii\,\mathfrak{su}(3)\). The sum is direct: if \(Y \in \mathfrak{su}(3) \cap \ii\,\mathfrak{su}(3)\), write \(Y = \ii Z\) with \(Z\) anti-Hermitian; then \(Y^{\dagger} = -\ii Z^{\dagger} = \ii Z = Y\), so \(Y\) is Hermitian as well as anti-Hermitian, and \(Y = 0\).
Ideals. The bracket of \(\mathfrak{sl}(3,\C)\) is \(\C\)-bilinear, so for \(X = A + \ii B\) with \(A,B \in \mathfrak{su}(3)\) and \(Y = P + \ii Q\) with \(P,Q \in \mathfrak{i}\),
and all four brackets lie in \(\mathfrak{i}\) because \(\mathfrak{i}\) is an ideal of \(\mathfrak{su}(3)\). So \(\comm{X}{Y} \in \mathfrak{i}_{\C}\), and \(\mathfrak{i}_{\C}\) is closed under multiplication by \(\ii\) hence a complex subspace. Finally \(\mathfrak{i} \cap \ii\,\mathfrak{i} = 0\), being contained in \(\mathfrak{su}(3) \cap \ii\,\mathfrak{su}(3) = 0\), so \(\mathfrak{i}_{\C} = \mathfrak{i} \oplus \ii\,\mathfrak{i}\) has real dimension \(2\dim_{\R}\mathfrak{i}\) and complex dimension \(\dim_{\R}\mathfrak{i}\).
∎The root decomposition
Write \(E_{jk}\) for the matrix unit with a single \(1\) in row \(j\) and column \(k\), so that \(E_{jk}E_{lm} = \delta_{kl}E_{jm}\). Let
the traceless diagonal matrices, which is two-dimensional and is the complex span of the \(T_{3}\) and \(T_{8}\) of Definition 14.51; it is the Cartan subalgebra whose existence Proposition 14.50 established, of rank \(\ell = 2\). For \(H = \diag(a_{1},a_{2},a_{3})\) a one-line computation with the matrix units gives
so that \(\mathfrak{sl}(3,\C)\) decomposes as
a direct sum of \(2 + 6 = 8\) pieces, each an eigenspace of every \(\ad_{H}\), and each of the six off-diagonal pieces one-dimensional. The linear functional \(\alpha_{jk} \in \mathfrak{h}^{*}\), \(\alpha_{jk}(H) = a_{j}-a_{k}\), is the root carried by \(E_{jk}\).
Evaluating \(\alpha_{jk}\) on the two Cartan generators of Definition 14.51 gives the coordinates used in Remark 14.56. With \(T_{3} = \tfrac{1}{2}\diag(1,-1,0)\) and \(T_{8} = \tfrac{1}{2\sqrt3}\diag(1,1,-2)\),
and \(E_{21}, E_{31}, E_{32}\) carry the negatives. These are exactly the six vectors of Equation (14.80): unit vectors at \(60\) degrees to one another, the regular hexagon. What the proof below uses is the one geometric property that hexagon has, namely that the sum of two roots at \(120\) degrees is again a root — \(|\gamma|=|\delta|=1\) and \(\gamma\cdot\delta = -\tfrac{1}{2}\) give \(\abs{\gamma+\delta}^{2} = 1\), and the sum bisects the angle. In matrix units that property reads \(\comm{E_{jk}}{E_{kl}} = E_{jl}\) for \(j,k,l\) distinct, which is the form in which it will be applied.
Let \(\mathfrak{i}\) be an ideal of \(\mathfrak{sl}(3,\C)\) and let \(x = h + \sum_{j\neq k}c_{jk}E_{jk} \in \mathfrak{i}\) be the decomposition of an element according to Equation (A.642). Then \(h \in \mathfrak{i}\) and \(c_{jk}E_{jk} \in \mathfrak{i}\) for every pair; in particular each \(E_{jk}\) with \(c_{jk} \neq 0\) lies in \(\mathfrak{i}\). Rests on Equation (A.642) and Definition A.375.
Derives Lemma A.380. Take the particular Cartan element \(H_{0} = \diag(3,-1,-2) \in \mathfrak{h}\). Its six root values are
together with \(-4,-5,-1\): six distinct and nonzero numbers, which is the only property of \(H_{0}\) that is used. Enumerate the six ordered pairs as \(\mu = 1,\ldots,6\), write \(E_{\mu}\) for the corresponding matrix unit and \(\lambda_{\mu} \in \set{\pm1,\pm4,\pm5}\) for its root value, so that \(x = h + \sum_{\mu}c_{\mu}E_{\mu}\). Because \(\mathfrak{i}\) is an ideal and \(\comm{H_{0}}{h} = 0\), all six elements
lie in \(\mathfrak{i}\), by Equation (A.641). The \(6\times6\) coefficient matrix \(\left(\lambda_{\mu}^{n}\right)\) has determinant \(\left(\prod_{\mu}\lambda_{\mu}\right)\) times the Vandermonde determinant \(\prod_{\mu<\nu}(\lambda_{\nu}-\lambda_{\mu})\), which is nonzero because the \(\lambda_{\mu}\) are nonzero and pairwise distinct. Inverting it expresses each \(c_{\mu}E_{\mu}\) as a linear combination of the \(\ad_{H_{0}}^{n}x\), so \(c_{\mu}E_{\mu} \in \mathfrak{i}\); and then \(h = x - \sum_{\mu}c_{\mu}E_{\mu} \in \mathfrak{i}\). Since \(\C E_{\mu}\) is one-dimensional, \(c_{\mu} \neq 0\) forces \(E_{\mu} \in \mathfrak{i}\).
∎Walking round the hexagon
Let \(\mathfrak{i}\) be an ideal of \(\mathfrak{sl}(3,\C)\) containing \(E_{jk}\) for one pair \(j \neq k\). Then \(\mathfrak{i} = \mathfrak{sl}(3,\C)\). Rests on Equations (A.641) and (A.642).
Derives Lemma A.381. Let \(l\) be the remaining index, so that \(\set{j,k,l} = \set{1,2,3}\). Every bracket below is a bracket of an element of \(\mathfrak{sl}(3,\C)\) with an element of \(\mathfrak{i}\), and therefore lies in \(\mathfrak{i}\).
The opposite root vector. From \(E_{kj}E_{jk} = E_{kk}\) and \(E_{jk}E_{kj} = E_{jj}\),
and \(H_{jk}\) is the diagonal matrix with entries \(+1\) at \(j\), \(-1\) at \(k\) and \(0\) at \(l\), so Equation (A.641) gives \(\comm{H_{jk}}{E_{kj}} = (-1-1)E_{kj} = -2E_{kj}\). Hence \(E_{kj} \in \mathfrak{i}\). In the language of Remark A.379: the root \(\gamma\) and the coroot direction it generates carry the ideal to \(-\gamma\), the opposite vertex of the hexagon.
The two neighbouring roots. Because \(k \neq l\) and \(l \neq j\),
the second term of each commutator vanishing — \(E_{jk}E_{lj} = 0\) since \(k \neq l\), and \(E_{kl}E_{jk} = 0\) since \(l \neq j\). These are the two roots at \(60\) degrees from the starting one, reached by adding the roots of \(E_{lj}\) and of \(E_{kl}\), each at \(120\) degrees to it.
Closing. Apply the first step to \(E_{lk}\) and to \(E_{jl}\): it returns \(E_{kl}, H_{lk} \in \mathfrak{i}\) and \(E_{lj}, H_{jl} \in \mathfrak{i}\). All six matrix units are now in \(\mathfrak{i}\), which is the whole hexagon. The diagonal part follows from Equation (A.645): \(H_{jk} = E_{jj}-E_{kk}\) and \(H_{lk} = E_{ll}-E_{kk}\) are linearly independent elements of the two-dimensional \(\mathfrak{h}\) of Equation (A.640), so they span it. By Equation (A.642), \(\mathfrak{i} = \mathfrak{sl}(3,\C)\).
∎Proof of the theorem
Proof of Theorem A.376. Derives Theorem A.376. The complex algebra. \(\mathfrak{sl}(3,\C)\) is not abelian, since \(\comm{E_{12}}{E_{21}} = E_{11}-E_{22} \neq 0\). Let \(\mathfrak{i} \neq \set{0}\) be an ideal and pick \(x \in \mathfrak{i}\) nonzero. By Lemma A.380 either some \(E_{jk}\) lies in \(\mathfrak{i}\) — in which case \(\mathfrak{i} = \mathfrak{sl}(3,\C)\) by Lemma A.381 — or every \(c_{jk}\) vanishes for every element of \(\mathfrak{i}\), so that \(\mathfrak{i} \subseteq \mathfrak{h}\) and \(x = h \neq 0\) is a nonzero traceless diagonal matrix. In that second case Equation (A.641) gives \(\comm{h}{E_{jk}} = \alpha_{jk}(h)\,E_{jk} \in \mathfrak{i}\); the right-hand side lies in \(\C E_{jk}\), which meets \(\mathfrak{h}\) only in \(0\), so \(\alpha_{jk}(h) = 0\) for all six pairs. Writing \(h = \diag(a_{1},a_{2},a_{3})\) this reads \(a_{1} = a_{2} = a_{3}\), which with \(\sum a_{i} = 0\) gives \(h = 0\) — contradicting the choice of \(x\). (Equivalently: the roots Equation (A.643) span \(\mathfrak{h}^{*}\), since \((1,0)\) and \(\left(-\tfrac{1}{2},\tfrac{\sqrt3}{2}\right)\) are independent, so a Cartan element annihilated by all of them is zero.) The second case is therefore empty, and \(\mathfrak{sl}(3,\C)\) is simple.
The real algebra. \(\mathfrak{su}(3)\) is not abelian: the generators \(T_{1},T_{2},T_{3}\) of Definition 14.51 satisfy \(\comm{T_{1}}{T_{2}} = \ii T_{3} \neq 0\) by Equation (14.69), and \(\ii T_{a}\) is the anti-Hermitian element of \(\mathfrak{su}(3)\) attached to the Hermitian generator \(T_{a}\). Let \(\mathfrak{i} \neq \set{0}\) be an ideal of \(\mathfrak{su}(3)\). By Lemma A.378, \(\mathfrak{i}_{\C}\) is a nonzero complex ideal of \(\mathfrak{sl}(3,\C)\), hence all of it by the previous paragraph, so
the last equality being Proposition 14.50. A subspace of full dimension is the whole space, so \(\mathfrak{i} = \mathfrak{su}(3)\).
∎Proof of Corollary A.377. Derives Corollary A.377. An invariant subspace \(V\) of the adjoint representation of a Lie algebra \(\mathfrak{g}\) on itself satisfies \(\comm{\mathfrak{g}}{V} \subseteq V\) by definition, which is exactly the statement that \(V\) is an ideal (Definition A.375); so Theorem A.376 leaves only \(\set{0}\) and \(\mathfrak{su}(3)\).
For the group statement, note first that the space of traceless complex \(3\times3\) matrices is \(\mathfrak{sl}(3,\C)\) and that \(M \mapsto UMU^{\dagger} = UMU^{-1}\) is the adjoint action of \(\SU(3)\) on it, as Proposition 14.55 and the proof of Theorem 14.57 record. A subspace \(V\) invariant under the group is invariant under the algebra: for \(X \in \mathfrak{su}(3)\) and \(M \in V\), the curve \(t \mapsto \ee^{tX}M\ee^{-tX}\) lies in \(V\), which is a closed subspace of a finite-dimensional space, and its derivative at \(t=0\) is \(\comm{X}{M}\), so \(\comm{X}{M} \in V\). Complex-linearity extends this from \(\mathfrak{su}(3)\) to \(\mathfrak{sl}(3,\C)\) by Equation (A.639), and the first paragraph applies.
∎Theorem A.376 discharges the two quotations named at the head of this section. In Proposition 14.53 it justifies the phrase “which is simple and hence semisimple”: the Killing form \(\kappa_{ab} = 3\delta_{ab}\) computed there is nondegenerate, as Theorem 14.13 requires of a semisimple algebra, and the implication now runs in the direction the chapter states rather than resting on an unproved assertion. In Theorem 14.57 it supplies the irreducibility of the eight-dimensional summand of \(\vect{3}\otimes\bar{\vect{3}}\), so that the decomposition \(3\times3 = 1+8\) of Equation (14.81) is proved and not merely counted; Corollary A.377 is the form in which it is used there.
One limitation is worth stating plainly. What is proved is the simplicity of this one algebra, by an argument that inspects its root system directly. The general theorem — that a complex simple Lie algebra is determined by its root system, and that the possible root systems are the classified Dynkin diagrams — is not proved here and is not needed: the whole of Lie Groups, Lie Algebras, and Fibre Bundles uses \(\mathfrak{su}(2)\) and \(\mathfrak{su}(3)\) and nothing else, and both are handled by hand.
From the Little Algebra to the Labels of a Massive Representation
This appendix supplies the step left owed after Lemma 14.68 of Lie Groups, Lie Algebras, and Fibre Bundles: the passage from the algebra that fixes a timelike momentum to the count of labels of a massive irreducible representation of the inhomogeneous algebra, and with it the arithmetic of Equation (14.107),
What Lemma 14.68 establishes is that the subalgebra annihilating a timelike \(p\) is \(\mathfrak{so}(D-1)\), of rank \(\lfloor(D-1)/2\rfloor\). What is owed is the representation theory that converts a rank into a count: that a massive irreducible representation is determined by the mass together with an irreducible representation of that subalgebra, and that no further invariant of the enveloping algebra escapes those labels.
The construction is Wigner's method of induced representations [Wigner:1939]. Particles as Poincaré Representations carries it out in the observed four dimensions — Theorem 94.27 is the construction and Theorem 94.29 the result — and this appendix does the same in general \(D\), with two differences of emphasis: the geometric ingredients (transitivity of the Lorentz group on the mass shell, the standard boost, the invariant measure) are proved here rather than quoted, and the arithmetic of the label count is separated cleanly from the classification theorem it rests on. That classification theorem is the one genuine import, and Remark A.393 says exactly what it is.
Under rule 7 of this treatise the general-\(D\) statement is the natural mathematical one and is stated as such; every worked instantiation below is \(3+1\).
Conventions, and what is assumed
Indices \(A,B,\ldots\) run from \(1\) to \(D\) and \(\eta_{AB} = \diag(+1,\ldots,+1,-1)\) is the flat metric of signature \((D-1,1)\) in the ordering of Notation 14.1, so that the time direction is the last one and \(x^{D} = ct\), as in Remark 14.88. A vector \(p\) is timelike when \(\eta_{AB}p^{A}p^{B} < 0\), which is the convention of Lemma 14.68.
The symmetry group is the connected inhomogeneous group
whose Lie algebra is \(\mathfrak{iso}(D-1,1)\) of Equation (14.99), with brackets Equations (14.84), (14.90) and (14.97); the translation \(a\) carries the SI unit \(\mathrm{m}\). A quantum symmetry is a projective unitary representation, and by Theorem 94.3 — proved in Wigner's Theorem on Quantum Symmetries — each symmetry is implemented by a unitary or antiunitary operator unique up to a phase, while Corollary 94.4 rules out the antiunitary branch for a connected group. Bargmann's analysis then replaces the projective representations of \(G\) by the ordinary unitary representations of its universal cover \(\tilde{G}\) [Bargmann:1954], whose Lorentz factor is \(\Spin(D-1,1)\). Everything below is written for \(G\); for the half-integer labels every \(\Lambda\) is read as its lift to \(\Spin(D-1,1)\) and every little-group element as an element of \(\Spin(D-1)\), exactly as Theorem 94.27 does in four dimensions.
The generators are normalized as in Remark 14.66: with \(P_{A} = \pp_{A}\) from Equation (14.82) the momentum operator is \(\hat{P}_{A} = -\ii\hbar P_{A}\), of SI unit \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), and with \(J_{AB}\) from Equation (14.86) the angular-momentum operator is \(\hat{J}_{AB} = -\ii\hbar J_{AB}\), of unit \(\mathrm{J}\,\mathrm{s}\).
In the mostly-plus ordering fixed above, a timelike momentum satisfies \(\eta_{AB}p^{A}p^{B} = -m^{2}c^{2}\) and therefore \(\hat{P}^{2} = -m^{2}c^{2}\) on a momentum eigenstate. Equation (14.105) records the same content with the opposite overall sign, being written in the mostly-minus convention customary in Particles as Poincaré Representations; the two differ by \(\eta \mapsto -\eta\) and nothing else. Only \(\abs{\hat{P}^{2}} = m^{2}c^{2}\) is used below, and the ratio Equation (14.106) is unaffected by the choice, both invariants changing sign together.
Statement
For \(m > 0\) put
so that \(k \in \mathcal{M}_{m}\) is the rest momentum, of components \(\vect{p} = 0\) and \(p^{D} = mc\). The little group is the stabilizer \(G_{k} = \set{\Lambda \in \SO(D-1,1)^{\uparrow} \mid \Lambda k = k}\). Rests on Notation 14.1 and Lemma 14.68.
Let \(D \ge 4\) and \(m > 0\). The unitary equivalence classes of irreducible, strongly continuous, positive-energy unitary representations \(U\) of \(\tilde{G}\) with \(\hat{P}^{2} = -m^{2}c^{2}\,\identity\) are in bijection with the unitary equivalence classes of irreducible unitary representations \(\sigma\) of \(\Spin(D-1)\), the bijection being \(U \cong U^{(m,\sigma)}\) with \(U^{(m,\sigma)}\) the induced representation Equation (A.654) below. Consequently the class of \(U\) is fixed by the mass \(m\) together with the \(\lfloor(D-1)/2\rfloor\) invariant labels of \(\sigma\) furnished by Theorem 14.21, and by no fewer numbers; a massive irreducible representation therefore carries exactly \(\lceil D/2\rceil\) labels, which is Equation (A.647). Every central element of the enveloping algebra acts on \(U\) as a scalar that is a function of those labels. Rests on Lemma 14.68, Theorem 14.21 and Definition A.384.
The bijection itself — the classification — is the imported ingredient; the geometry that makes it meaningful, the construction of \(U^{(m,\sigma)}\), and the arithmetic of the count are proved here.
The geometry of the mass shell
\(G_{k} = \SO(D-1)\), acting on the first \(D-1\) coordinates, and its Lie algebra is the \(\mathfrak{so}(D-1)\) of Lemma 14.68. It is compact. Rests on Definition A.384, Equation (14.23) and Lemma 14.68.
Derives Lemma A.386. Let \(\Lambda k = k\), that is \(\Lambda e_{D} = e_{D}\) where \(e_{D}\) is the unit vector in the time direction. Since \(\Lambda\) preserves \(\eta\), it preserves the orthogonal complement \(e_{D}^{\perp} = \R^{D-1}\), on which \(\eta\) restricts to the Euclidean metric; so the restriction lies in \(\Ogrp(D-1)\). Its determinant equals \(\det\Lambda = 1\) because \(\Lambda\) acts as the identity on the remaining direction, whence \(\Lambda|_{\R^{D-1}} \in \SO(D-1)\); conversely every such rotation, extended by \(e_{D} \mapsto e_{D}\), preserves \(\eta\), has determinant \(1\) and is orthochronous. Compactness is that of a closed bounded subset of \(\R^{(D-1)\times(D-1)}\): the rows of an orthogonal matrix are unit vectors, and orthogonality is a closed condition. Differentiating a curve through the identity gives exactly the subalgebra of Lemma 14.68, namely the elements of \(\mathfrak{so}(D-1,1)\) annihilating \(k\).
∎For \(p \in \mathcal{M}_{m}\) write \(\vect{p} = (p^{1},\ldots,p^{D-1})\) and \(\abs{\vect{p}}\) for its Euclidean length, so that \(p^{D} = \sqrt{\abs{\vect{p}}^{2} + m^{2}c^{2}}\). The matrix
with \(i,j = 1,\ldots,D-1\), is a proper orthochronous Lorentz transformation, depends continuously on \(p\) with \(L(k) = \identity\), and satisfies \(L(p)k = p\). In particular \(\SO(D-1,1)^{\uparrow}\) acts transitively on \(\mathcal{M}_{m}\), and \(\mathcal{M}_{m}\) is a single orbit. Rests on Definition A.384 and Equation (14.23).
Derives Lemma A.387. Put \(\hat{n}^{i} = p^{i}/\abs{\vect{p}}\), \(\cosh\zeta = p^{D}/(mc)\) and \(\sinh\zeta = \abs{\vect{p}}/(mc)\); these are consistent because \(\cosh^{2}\zeta - \sinh^{2}\zeta = \left((p^{D})^{2} - \abs{\vect{p}}^{2}\right)/(mc)^{2} = 1\) by Equation (A.649), and \(\zeta \ge 0\) is determined. In these terms Equation (A.650) reads
\(\hat{n}\) being regarded as a vector in \(\R^{D}\) with vanishing last component. Reading off the entries confirms that this is Equation (A.650): the \(\hat{n}\hat{n}\transpose\) term gives \(L^{i}{}_{j}\), the two mixed terms give \(L^{i}{}_{D} = L^{D}{}_{i} = \sinh\zeta\,\hat{n}^{i} = p^{i}/(mc)\), and the \(e_{D}e_{D}\transpose\) term raises \(L^{D}{}_{D}\) from \(1\) to \(\cosh\zeta = p^{D}/(mc)\). It is the identity on the \((D-2)\)-dimensional subspace orthogonal to both \(\hat{n}\) and \(e_{D}\), and on the plane they span it is \(\begin{pmatrix}\cosh\zeta & \sinh\zeta\\ \sinh\zeta & \cosh\zeta\end{pmatrix}\) in the basis \((\hat{n},e_{D})\). Restricted to that plane \(\eta\) is \(\diag(+1,-1)\), and \(\cosh^{2}\zeta - \sinh^{2}\zeta = 1\) is exactly the statement that the matrix preserves it; on the complement \(L(p)\) is the identity and the two blocks are \(\eta\)-orthogonal. So Equation (14.23) holds at the group level and \(L(p) \in \Ogrp(D-1,1)\). Its determinant is \(\cosh^{2}\zeta - \sinh^{2}\zeta = 1\) and its \((D,D)\) entry is \(\cosh\zeta \ge 1 > 0\), so it is proper and orthochronous.
It maps \(k\) to \(p\). From Equation (A.650), \(\left(L(p)k\right)^{i} = L(p)^{i}{}_{D}\,mc = p^{i}\) and \(\left(L(p)k\right)^{D} = L(p)^{D}{}_{D}\,mc = p^{D}\).
Continuity at \(\vect{p} = 0\). The only term in Equation (A.650) whose form is singular there is \(p^{i}p^{j}\left(p^{D}/(mc) - 1\right)/\abs{\vect{p}}^{2}\), and
so the quotient by \(\abs{\vect{p}}^{2}\) is bounded and tends to \(0\); the whole term vanishes with \(\vect{p}\), leaving \(L(k) = \identity\). The remaining entries are manifestly continuous.
Transitivity. For \(p,q \in \mathcal{M}_{m}\) the element \(L(q)L(p)^{-1}\) is in \(\SO(D-1,1)^{\uparrow}\) and carries \(p\) to \(q\). Conversely a Lorentz transformation preserves \(\eta_{AB}p^{A}p^{B}\), and an orthochronous one preserves the sign of \(p^{D}\) on timelike vectors, so it maps \(\mathcal{M}_{m}\) into itself.
∎The measure
on \(\mathcal{M}_{m}\) is invariant under \(\SO(D-1,1)^{\uparrow}\). Rests on Definition A.384 and Lemma A.387.
Derives Lemma A.388. Consider on \(\R^{D}\) the measure
with \(\theta\) the unit step. Every factor is invariant: \(\dd^{D}p\) because \(\abs{\det\Lambda} = 1\), the delta because its argument is a Lorentz scalar, and the step because \(\Lambda\) is orthochronous. Now carry out the \(p^{D}\) integration. Writing \(f(p^{D}) = \abs{\vect{p}}^{2} - (p^{D})^{2} + m^{2}c^{2}\), the only zero with \(p^{D} > 0\) is \(p^{D} = \sqrt{\abs{\vect{p}}^{2}+m^{2}c^{2}}\), at which \(\abs{f'} = 2p^{D}\); so \(\delta(f)\,\theta(p^{D})\,\dd p^{D}\) integrates to \(1/\left(2p^{D}\right)\) and \(\dd\nu = \dd^{D-1}\vect{p}/(2p^{D}) = \tfrac{1}{2}\dd\mu\). A constant multiple of an invariant measure is invariant.
∎Each \(p^{i}\) carries the unit \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), so \(\dd\mu\) of Equation (A.652) carries \(\left(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\right)^{D-2}\); in the observed \(D = 4\) that is \(\mathrm{kg}^{2}\,\mathrm{m}^{2}/\mathrm{s}^{2}\). A wave function on the mass shell therefore carries the reciprocal of the square root of that unit, so that \(\int\dd\mu(p)\,\abs{\psi(p)}^{2}\) is a pure number, as a probability must be. One may of course divide \(\dd\mu\) by \((mc)^{D-2}\) to obtain a dimensionless measure; nothing below depends on the choice, and the unnormalized form is kept because it is the one in which Lemma A.388 is transparent.
The induced representation
Let \(\sigma\) be a unitary representation of \(G_{k} = \SO(D-1)\) — or of \(\Spin(D-1)\), for the covering group — on a finite-dimensional Hilbert space \(V_{\sigma}\). On
define, for \(a \in \R^{D}\) and \(\Lambda \in \SO(D-1,1)^{\uparrow}\),
with the Wigner rotation
and \(U^{(m,\sigma)}(\Lambda,a) = U^{(m,\sigma)}(\identity,a)\,U^{(m,\sigma)}(\Lambda,0)\). Then \(W(\Lambda,q) \in G_{k}\), and \(U^{(m,\sigma)}\) is a strongly continuous unitary representation of \(G\) with
and positive energy, \(p^{D} > 0\) throughout the spectrum. Rests on Lemmas A.386, A.387 and A.388.
Derives Proposition A.390. The Wigner rotation lies in the little group. Using \(L(q)k = q\) from Lemma A.387,
so \(W(\Lambda,q) \in G_{k}\) by Definition A.384, and it depends continuously on \((\Lambda,q)\) because \(L\) does.
The composition law. Put \(q = \Lambda^{-1}\Lambda'^{-1}p\), so that \(\Lambda q = \Lambda'^{-1}p\) and \(\Lambda'\Lambda q = p\). Then
the two middle factors cancelling. Applying Equation (A.654) twice and using Equation (A.657) together with the homomorphism property of \(\sigma\),
which is \(\left(U(\Lambda'\Lambda,0)\psi\right)(p)\). The translations compose by inspection, \(\ee^{-\ii p\cdot a/\hbar} \ee^{-\ii p\cdot a'/\hbar} = \ee^{-\ii p\cdot(a+a')/\hbar}\), and the mixed relation required by Equation (A.648) is
because \(\eta\) is \(\Lambda\)-invariant: \((\Lambda^{-1}p)\cdot a = p\cdot(\Lambda a)\). That is \(U(\Lambda,0)U(\identity,a) = U(\identity,\Lambda a)U(\Lambda,0)\), the semidirect product law.
Unitarity. For the translations the factor is a phase. For the Lorentz part,
since \(\sigma\) is unitary on \(V_{\sigma}\), and the last integral equals \(\norm{\psi}^{2}\) by the invariance of \(\dd\mu\) (Lemma A.388). Surjectivity is \(U(\Lambda,0)^{-1} = U(\Lambda^{-1},0)\). Strong continuity follows from the continuity of \(L\), \(W\) and \(\sigma\) together with dominated convergence on the dense set of compactly supported continuous \(\psi\).
The momentum spectrum. The one-parameter groups of translations are \(U(\identity,a) = \exp\left(-\ii\,\eta_{AB}\hat{P}^{A}a^{B} /\hbar\right)\), which is the normalisation of Remark 14.66; comparing with the first line of Equation (A.654), whose phase is \(-\ii\,\eta_{AB}p^{A}a^{B}/\hbar\) for every \(a\), the generator \(\hat{P}^{A}\) acts as multiplication by the number \(p^{A}\), which is Equation (A.656); then \(\hat{P}^{2} = \eta^{AB}p_{A}p_{B} = -m^{2}c^{2}\) everywhere on \(\mathcal{M}_{m}\) by Equation (A.649), and \(p^{D}>0\) there by construction.
∎If \(V_{\sigma}\) has a proper nonzero \(\sigma\)-invariant subspace \(V'\), then \(L^{2}(\mathcal{M}_{m},\dd\mu;V')\) is a proper nonzero \(U^{(m,\sigma)}\)-invariant subspace of \(\mathcal{H}_{\sigma}\). Hence irreducibility of \(U^{(m,\sigma)}\) requires irreducibility of \(\sigma\). Rests on Proposition A.390.
Derives Proposition A.391. Both operations in Equation (A.654) act on the value \(\psi(q) \in V_{\sigma}\) either by a scalar phase or by \(\sigma(W)\) with \(W \in G_{k}\), and both preserve \(V'\); the argument \(q\) is merely relabelled. So the subspace of functions with values in \(V'\) is invariant, and it is proper and nonzero because \(V'\) is.
∎The classification, and the count
The converse of Proposition A.391, and the statement that the list Equation (A.654) is exhaustive, is the imported ingredient.
Let \(D \ge 4\), \(m > 0\), and let \(U\) be a strongly continuous unitary representation of \(\tilde{G}\) that is irreducible, has positive energy and satisfies \(\hat{P}^{2} = -m^{2}c^{2}\identity\). Then \(U\) is unitarily equivalent to \(U^{(m,\sigma)}\) of Equation (A.654) for some irreducible unitary representation \(\sigma\) of \(\Spin(D-1)\); conversely \(U^{(m,\sigma)}\) is irreducible whenever \(\sigma\) is; and \(U^{(m,\sigma)}\) and \(U^{(m,\sigma')}\) are equivalent if and only if \(\sigma\) and \(\sigma'\) are [Wigner:1939] [Weinberg:1995]. Rests on Proposition A.390 and Definition A.384.
Theorem A.392 is the one statement of this appendix that is not proved in this treatise, and it is worth being exact about what it costs.
Its content is Mackey's imprimitivity theorem [Mackey:1952], specialized by Wigner [Wigner:1939] to the inhomogeneous Lorentz group. Three analytic inputs go into it, none of which this treatise develops: the spectral theorem for a strongly continuous unitary representation of the translation group \(\R^{D}\) (the Stone–Naimark–Ambrose–Godement theorem, which supplies the projection-valued measure whose support is the momentum spectrum — Stone's one-parameter case is [Stone:1932] and the general statement is standard functional analysis [Reed:1972]); the identification of the commutant of that measure's multiplication algebra with the decomposable operators, which is what turns an operator commuting with all translations into a measurable field \(p \mapsto A(p)\) of matrices; and the disintegration of the Hilbert space over the orbit, which converts the field into a representation of the stabilizer. Given those three, the classification is the argument whose geometric half is proved above: the spectrum is a \(\Lambda\)-invariant subset of the hyperboloid Equation (A.649), which is one orbit by Lemma A.387, so the spectrum is the whole mass shell; the explicit continuous section \(L\) of Equation (A.650) then trivializes the bundle of fibres, and the residual freedom is exactly a representation of \(G_{k}\).
Two things follow from that accounting. First, the imported theorem is functional analysis, not group theory: the Lie-theoretic content — which subalgebra fixes a timelike momentum, what its rank is, and how many labels its representations carry — is proved, in Lemma 14.68 and in Theorem 14.21. Second, the existence half of the classification, which is what a physical application actually uses, is proved here without any import: Proposition A.390 constructs \(U^{(m,\sigma)}\) and computes its invariants, and Proposition A.391 supplies one direction of the irreducibility criterion. What is quoted is the exhaustiveness of the list, together with the converse irreducibility statement.
One further standard fact is used silently and is named here: a unitary representation of a compact group decomposes into finite-dimensional irreducible ones (the Peter–Weyl theorem [Weinberg:1995]), which is why \(V_{\sigma}\) may be taken finite-dimensional in Proposition A.390. Lemma 94.28 records the same point for \(\SU(2)\) in four dimensions.
Proof of Theorem A.385. Derives Theorem A.385. The bijection between equivalence classes is Theorem A.392. It remains to count.
The little algebra and its rank. By Lemma A.386 the little group is \(\SO(D-1)\), whose Lie algebra is the \(\mathfrak{so}(D-1)\) of Lemma 14.68; and since \(D \ge 4\), that algebra has \(D-1 \ge 3\) and is semisimple, of rank \(\lfloor(D-1)/2\rfloor\) by Proposition 14.22. Passing to \(\Spin(D-1)\) changes neither the algebra nor its rank.
How many labels \(\sigma\) carries. Theorem 14.21 applied to \(\mathfrak{so}(D-1)\) says that the centre of its enveloping algebra is a polynomial algebra in \(\lfloor(D-1)/2\rfloor\) algebraically independent Casimir elements, and that a complete set of invariant labels for its finite-dimensional irreducible representations consists of that many numbers and of no fewer. The representation \(\sigma\) is finite-dimensional by the Peter–Weyl theorem (Remark A.393), so exactly \(\lfloor(D-1)/2\rfloor\) numbers fix its class.
Adding the mass. The number \(m\) is an invariant of \(U\) independent of those: it is read off \(\hat{P}^{2}\), which by Equation (A.656) is \(-m^{2}c^{2}\) and involves the translation generators alone, whereas the little-group Casimirs are built from \(\hat{J}\). Distinct masses therefore give inequivalent representations for every \(\sigma\), and distinct \(\sigma\) give inequivalent representations for every mass, by Theorem A.392. The complete label set is thus \(m\) together with the \(\lfloor(D-1)/2\rfloor\) labels of \(\sigma\), and by Equation (14.108) of Lemma 14.68 their number is
which is Equation (A.647) and Equation (14.107).
No further invariant survives. Let \(C\) be any central element of the enveloping algebra \(U\!\left(\mathfrak{iso}(D-1,1)\right)\). On the irreducible \(U\) it acts as a scalar by Corollary 14.18, and that scalar depends only on the unitary equivalence class of \(U\) — an intertwiner conjugates \(C\) to itself. By the bijection the class is the pair \((m,[\sigma])\), so the scalar is a function of \(m\) and of the \(\lfloor(D-1)/2\rfloor\) labels of \(\sigma\): no invariant of the enveloping algebra can separate two representations that those numbers already identify.
∎The observed case
Take \(D = 4\), the observed spacetime. Then \(\lfloor(D-1)/2\rfloor = \lfloor 3/2 \rfloor = 1\) and \(\lceil D/2\rceil = 2\): a massive particle carries exactly two labels.
The little group is \(\SO(3)\) by Lemma A.386, its cover is \(\SU(2)\), and its algebra has rank one (Proposition 14.22 with \(D-1 = 3\)), so Theorem 14.21 allows a single Casimir — the \(\vect{J}^{2}\) of Definition 14.40. Its irreducible representations are labelled by \(s\) with \(2s \in \N\cup\set{0}\) (Theorem 14.43), of dimension \(2s+1\), and the Casimir takes the value
\(s\) itself being dimensionless. The two labels of Theorem A.385 are therefore \(m\), of unit \(\mathrm{kg}\), and \(s\).
In four dimensions — and, of the cases treated in Lie Groups, Lie Algebras, and Fibre Bundles, only there — both labels are realized by genuine Casimir elements of the inhomogeneous algebra itself: \(\hat{P}^{2}\) and the Pauli–Lubanski square \(\hat{W}^{2}\) of Theorem 14.65, whose values are Equation (14.105). That \(\hat{W}^{2}\) measures the rest-frame angular momentum, so that \(\hat{W}^{2} = -m^{2}c^{2}\hat{\vect{\mathcal{J}}}^{2}\) on the fibre over \(k\), is the computation carried out in Remark 14.66 and again in Theorem 94.29; combined with Equation (A.658) it gives the second line of Equation (14.105), and the ratio Equation (14.106) isolates the \(\hbar^{2}s(s+1)\) that survives. Theorem 94.27 is Proposition A.390 written out for this case, with \(\dd\mu\) of Equation (A.652) replaced by the equivalent normalization used there. Rests on Theorems 14.43, 14.65 and A.385.
Remark 14.67 states that in general \(D\) the invariants beyond \(\hat{P}^{2}\) are built from antisymmetrized products of \(\hat{J}\) with \(\hat{P}\), the Pauli–Lubanski construction Equation (14.103) being available only at \(D = 4\) because it contracts a \(D\)-index Levi-Civita symbol with one \(J\) and one \(P\). That construction is not carried out in this appendix, and it is not needed for the theorem: the count Equation (A.647) follows from the classification and from Theorem 14.21 applied to the little algebra, neither of which requires the invariants to be exhibited as elements of the enveloping algebra. What has been proved is that a massive irreducible representation is specified by \(\lceil D/2\rceil\) numbers and by no fewer, and that any central element of the enveloping algebra is a function of them. The explicit general-\(D\) construction of the higher invariants would say which polynomial in \(\hat{J}\) and \(\hat{P}\) realizes each label; only the \(D = 4\) case is used anywhere in this book, and it is Theorem 14.65.
Theorem A.385 discharges the obligation recorded after Lemma 14.68: the lemma supplies the little algebra and its rank, and the theorem converts that rank into the label count Equation (14.107) quoted in Remark 14.67 — two labels at \(D = 4\) and three at both \(D = 5\) and \(D = 6\), the ceiling and not the floor. The \(3+1\) instantiation, Example A.394, is the mass and the spin, and it is the only case the physical parts of this treatise use; its detailed development, including the massless representations that this appendix does not treat, is Particles as Poincaré Representations.
Whitehead's Lemmas and the Rigidity of Semisimple Algebras
This appendix proves Proposition 14.77 of Lie Groups, Lie Algebras, and Fibre Bundles: the first and second Chevalley–Eilenberg cohomology groups of a finite-dimensional semisimple Lie algebra, with coefficients in any finite-dimensional module, vanish. The second of these — Whitehead's second lemma — is what Section 14.4 needs, since by Proposition 14.76 the vanishing of \(H^{2}(\mathfrak{g},\R)\) says exactly that every central extension of \(\mathfrak{g}\) is trivial. The corollary for the Lorentz algebra follows from Corollary A.344.
Nothing is quoted. The engine is the Casimir operator of the representation: it is invertible on any module with no trivial summand, and averaging a cocycle against it produces the coboundary that was sought. Semisimplicity enters through Cartan's Criterion for Semisimplicity at three points — the Killing form and the trace form are nondegenerate, ideals split off orthogonally, and \(\comm{\mathfrak{g}}{\mathfrak{g}} = \mathfrak{g}\).
Section 14.4.1 defines \(H^{2}(\mathfrak{g},\R)\) directly, by writing down the cocycle condition Equation (14.121) and the coboundaries Equation (14.122) for the one case it needs; the complex those live in is named there but not built. It is built here (The Chevalley–Eilenberg complex), because the proof needs coefficients in an arbitrary module and needs degree three to say what a \(2\)-cocycle is. The Casimir operator of a representation constructs the Casimir of a representation, Averaging a cocycle against the Casimir carries out the averaging, Complete reducibility proves the complete reducibility needed to reduce a general module to irreducible ones, Trivial coefficients treats the trivial coefficients that averaging cannot reach, and The Whitehead lemmas assembles the lemmas and draws the corollary.
Throughout, \(\mathbb{K}\) is a field of characteristic zero with algebraic closure \(\mathbb{F}\), \(\mathfrak{g}\) is a finite-dimensional semisimple Lie algebra over \(\mathbb{K}\) with basis \(\set{T_{a}}\), and a module is a finite-dimensional vector space \(\mathbb{V}\) with a homomorphism \(\rho : \mathfrak{g} \to \mathfrak{gl}(\mathbb{V})\), that is, a linear map with \(\rho\left(\comm{X}{Y}\right) = \comm{\rho(X)}{\rho(Y)}\). We write \(\rho(X)v\) and \(X\cdot v\) interchangeably, and \(\rho_{a} = \rho(T_{a})\).
The Chevalley–Eilenberg complex
Let \(\mathbb{V}\) be a \(\mathfrak{g}\)-module. For \(n \ge 1\) let \(C^{n}(\mathfrak{g},\mathbb{V})\) be the space of alternating \(n\)-linear maps \(\mathfrak{g}^{n} \to \mathbb{V}\), and let \(C^{0}(\mathfrak{g},\mathbb{V}) = \mathbb{V}\). The Chevalley–Eilenberg differential \(\dd : C^{n} \to C^{n+1}\) is
the hat marking an omitted argument. Explicitly, in the three degrees used below,
A cochain with \(\dd f = 0\) is a cocycle, one of the form \(\dd h\) a coboundary, and \(H^{n}(\mathfrak{g},\mathbb{V})\) is the quotient of the cocycles by the coboundaries in degree \(n\). Rests on Definition 14.75.
\(\dd\left(\dd v\right) = 0\) for \(v \in C^{0}\) and \(\dd\left(\dd f\right) = 0\) for \(f \in C^{1}\). Hence \(H^{1}(\mathfrak{g},\mathbb{V})\) and \(H^{2}(\mathfrak{g},\mathbb{V})\) are defined. Rests on Definition A.397.
Derives Lemma A.398. In degree zero, Equations (A.660) and (A.661) give
which is precisely the statement that \(\rho\) is a homomorphism.
In degree one, write \(g = \dd f\) and expand \(\left(\dd g\right)(X,Y,Z)\) by Equation (A.662), substituting Equation (A.661) in each of the six terms. Suppressing \(\rho\) and writing \(Xu\) for \(\rho(X)u\), the six terms are
Add them and collect. The terms carrying \(f(Z)\) are \(XY - YX - \comm{X}{Y}\), which vanish because \(\rho\) is a homomorphism; likewise those carrying \(f(Y)\), namely \(-XZ + ZX + \comm{X}{Z}\), and those carrying \(f(X)\), namely \(YZ - ZY - \comm{Y}{Z}\). The terms \(\pm Xf\left(\comm{Y}{Z}\right)\), \(\pm Yf\left(\comm{X}{Z}\right)\) and \(\pm Zf\left(\comm{X}{Y}\right)\) cancel in pairs. What remains is
because \(-\comm{\comm{X}{Z}}{Y} = \comm{\comm{Z}{X}}{Y}\) and the argument is then the Jacobi identity.
∎Take \(\mathbb{V} = \R\) with the trivial action, \(\rho = 0\). Then Equation (A.662) reduces to
using the antisymmetry of \(c\) to write \(c\left(\comm{X}{Z},Y\right) = -c\left(\comm{Z}{X},Y\right)\); so \(\dd c = 0\) is exactly the cocycle condition Equation (14.121) of Proposition 14.73. And Equation (A.661) reduces to \(\left(\dd b\right)(X,Y) = -b\left(\comm{X}{Y}\right)\), whose image is the space of coboundaries Equation (14.122) of Definition 14.74 — the overall sign changes no subspace. The group \(H^{2}(\mathfrak{g},\R)\) of Definition 14.75 is therefore the group \(H^{2}\) of Definition A.397, and Proposition 14.76 may be applied to it without further comment.
The Casimir operator of a representation
A symmetric bilinear form \(\beta\) on \(\mathfrak{g}\) is invariant if
which for the Killing form is Equation (14.16) rewritten with the antisymmetry of the bracket. If \(\beta\) is in addition nondegenerate, the dual basis \(\set{T^{a}}\) is defined by \(\beta\left(T_{a},T^{b}\right) = \delta_{a}{}^{b}\), and the Casimir operator of the representation \(\rho\) relative to \(\beta\) is
summed over \(a\). Taking \(\beta = \kappa\) and \(\rho\) the identity map of \(\mathfrak{g}\) into \(U(\mathfrak{g})\) recovers \(C_{2} = \kappa^{ab}T_{a}T_{b}\) of Equation (14.20). Rests on Definition 14.10 and Lemma 14.11.
Let \(\beta\) be invariant and nondegenerate with dual basis \(\set{T^{a}}\). Then for every \(Y \in \mathfrak{g}\)
Consequently \(\Gamma\) commutes with \(\rho(Y)\) for every \(Y \in \mathfrak{g}\). Rests on Definition A.400.
Derives Lemma A.401. Expand \(\comm{T_{a}}{Y} = \mu_{a}{}^{b}T_{b}\) and \(\comm{T^{a}}{Y} = \nu^{a}{}_{b}T^{b}\), so that \(\mu_{a}{}^{b} = \beta\left(\comm{T_{a}}{Y},T^{b}\right)\) and \(\nu^{a}{}_{b} = \beta\left(\comm{T^{a}}{Y},T_{b}\right)\). By Equation (A.667) and the symmetry of \(\beta\),
The left-hand side of Equation (A.669) is therefore \(\sum_{a,b}\left(\nu^{a}{}_{b} + \mu_{b}{}^{a}\right) T^{b}\otimes T_{a}\), after renaming the summation indices in the second term, and each coefficient vanishes by Equation (A.670).
For the last statement apply the bilinear map \((u,w) \mapsto \rho(u)\rho(w)\) to Equation (A.669):
which is the image of Equation (A.669) and hence zero.
∎Let \(\mathfrak{g}\) be semisimple and \(\rho\) faithful. Then
is a symmetric invariant nondegenerate bilinear form on \(\mathfrak{g}\), and the corresponding Casimir operator satisfies \(\tr_{\mathbb{V}}\Gamma = \dim\mathfrak{g}\). Rests on Theorems A.330 and A.340.
Derives Lemma A.402. Symmetry is cyclicity of the trace, and invariance is the computation Equation (A.605): \(\tr\left(\rho\left(\comm{X}{Y}\right) \rho(Z)\right) = \tr\left(\rho(X)\rho\left(\comm{Y}{Z}\right)\right)\). Let \(\mathfrak{s}\) be the radical of \(\beta_{\mathbb{V}}\); it is an ideal by the argument of Equation (A.606) with \(\beta_{\mathbb{V}}\) in place of \(\kappa\). For \(W \in \comm{\mathfrak{s}}{\mathfrak{s}}\) and \(Y \in \mathfrak{s}\), \(\tr\left(\rho(W)\rho(Y)\right) = \beta_{\mathbb{V}}(W,Y) = 0\), so Theorem A.340 makes \(\rho(\mathfrak{s})\) solvable; as \(\rho\) is faithful, \(\mathfrak{s}\) is a solvable ideal of a semisimple algebra, hence zero (Remark A.333). Finally
Averaging a cocycle against the Casimir
Let \(\beta\) be an invariant nondegenerate symmetric form on \(\mathfrak{g}\) and let the Casimir \(\Gamma\) of Equation (A.668) be invertible on \(\mathbb{V}\). Then
Rests on Lemma A.401 and Definition A.397.
Derives Theorem A.403. Degree one. Let \(f\) be a \(1\)-cocycle, so that by Equation (A.661)
Put \(X = T_{a}\), apply \(\rho\left(T^{a}\right)\) and sum over \(a\). The first term gives \(\Gamma f(Y)\). In the second, commute \(\rho\left(T^{a}\right)\) past \(\rho(Y)\),
and set
The result is
The last two terms are the image of the invariance identity Equation (A.669) under the bilinear map \((u,w) \mapsto \rho(u)f(w)\), hence vanish. So \(\Gamma f(Y) = \rho(Y)v\) for every \(Y\). Since \(\Gamma\) commutes with every \(\rho(Y)\) (Lemma A.401) so does \(\Gamma^{-1}\), and therefore
by Equation (A.660): every \(1\)-cocycle is a coboundary.
Degree two. Let \(c\) be a \(2\)-cocycle. Put \(X = T_{a}\) in \(\left(\dd c\right)(X,Y,Z) = 0\), apply \(\rho\left(T^{a}\right)\), sum over \(a\), and set
The six terms of Equation (A.662) contribute as follows. The first gives \(\Gamma c(Y,Z)\) outright. In the second and the third, Equation (A.676) moves \(\rho\left(T^{a}\right)\) past \(\rho(Y)\) and \(\rho(Z)\), producing \(-\rho(Y)b(Z)\) and \(+\rho(Z)b(Y)\) together with one commutator term each. The fourth and fifth are unchanged. The sixth is \(b\left(\comm{Y}{Z}\right)\), since \(-c\left(\comm{Y}{Z},T_{a}\right) = +c\left(T_{a},\comm{Y}{Z}\right)\). Collecting,
where the leftover is
Both lines of Equation (A.682) vanish: applying the bilinear map \((u,w) \mapsto \rho(u)c(w,Z)\) to Equation (A.669) gives
and the same computation with \(Y\) and \(Z\) exchanged kills the second line. So Equation (A.681) leaves
that is, by Equation (A.661), \(\Gamma c = \dd b\). As in degree one, \(\Gamma^{-1}\) commutes with every \(\rho(Y)\) and therefore with \(\dd\), so \(c = \dd\left(\Gamma^{-1}b\right)\): every \(2\)-cocycle is a coboundary.
∎The name is worth justifying. For a compact group the standard proof of the same statement integrates a cochain over the group with the Haar measure, and the invariant so produced is a coboundary. There is no group and no measure here, but \(\Gamma\) plays the same role: it is the one operator built from the algebra alone that commutes with everything (Lemma A.401), and Equations (A.678) and (A.684) say that contracting a cocycle with the Casimir tensor \(T^{a}\otimes T_{a}\) produces the cochain whose coboundary it is. What has to be paid for is the invertibility of \(\Gamma\), and the price is exactly the trivial summands of \(\mathbb{V}\), on which \(\Gamma\) vanishes identically. Those are dealt with separately in Trivial coefficients, and they are the reason the proof needs complete reducibility as well.
Complete reducibility
Let \(\mathbb{V}\) be a finite-dimensional irreducible module over \(\mathbb{F}\) algebraically closed, and let \(A \in \operatorname{End}(\mathbb{V})\) commute with \(\rho(X)\) for every \(X \in \mathfrak{g}\). Then \(A = \lambda\identity\) for some \(\lambda \in \mathbb{F}\). Rests on Theorem 5.153.
Derives Lemma A.405. The argument is that of Theorem 5.153, with the family of operators \(\set{U(a)}\) replaced by \(\set{\rho(X)}\): the characteristic polynomial of \(A\) has a root \(\lambda\) because \(\mathbb{F}\) is algebraically closed and \(\dim\mathbb{V} \ge 1\), so \(\ker\left(A - \lambda\identity\right) \neq 0\); that kernel is a submodule, because \(A - \lambda\identity\) commutes with the action; and irreducibility forces it to be all of \(\mathbb{V}\).
∎Let \(\mathfrak{g}\) be semisimple over \(\mathbb{F}\) and let \(\mathbb{V}\) be a module with a submodule \(\mathbb{W}\) of codimension one. Then \(\mathbb{V} = \mathbb{W} \oplus \mathbb{X}\) for some one-dimensional submodule \(\mathbb{X}\). Rests on Lemma A.402, Lemma A.405 and Corollary A.342.
Derives Lemma A.406. First, \(\rho(\mathfrak{g})\mathbb{V} \subseteq \mathbb{W}\): the quotient \(\mathbb{V}/\mathbb{W}\) is a one-dimensional module, so the action on it is a homomorphism of \(\mathfrak{g}\) into an abelian algebra, which annihilates \(\comm{\mathfrak{g}}{\mathfrak{g}} = \mathfrak{g}\) by Corollary A.342.
Next, we may assume \(\rho\) faithful. Otherwise put \(\mathfrak{k} = \ker\rho\), an ideal; by Corollary A.342, \(\mathfrak{g}/\mathfrak{k} \cong \mathfrak{k}^{\perp}\) is again semisimple, and \(\mathbb{V}\) is a module over it with the same submodules. If \(\mathfrak{g}/\mathfrak{k} = 0\) the action is trivial, every subspace is a submodule and any complementary line serves.
Induct on \(\dim\mathbb{W}\).
Case 1: \(\mathbb{W}\) has a submodule \(\mathbb{W}'\) with \(0 \neq \mathbb{W}' \neq \mathbb{W}\). In \(\mathbb{V}/\mathbb{W}'\) the submodule \(\mathbb{W}/\mathbb{W}'\) has codimension one and smaller dimension, so by induction there is a one-dimensional submodule \(\widetilde{\mathbb{X}}/\mathbb{W}'\) with \(\mathbb{V}/\mathbb{W}' = \mathbb{W}/\mathbb{W}' \oplus \widetilde{\mathbb{X}}/\mathbb{W}'\). Inside \(\widetilde{\mathbb{X}}\) the submodule \(\mathbb{W}'\) has codimension one and \(\dim\mathbb{W}' < \dim\mathbb{W}\), so by induction again \(\widetilde{\mathbb{X}} = \mathbb{W}' \oplus \mathbb{X}\) with \(\mathbb{X}\) a one-dimensional submodule. Then \(\mathbb{X} \cap \mathbb{W} \subseteq \widetilde{\mathbb{X}} \cap \mathbb{W} = \mathbb{W}'\) and \(\mathbb{X} \cap \mathbb{W}' = 0\), so \(\mathbb{X} \cap \mathbb{W} = 0\) and dimensions give \(\mathbb{V} = \mathbb{W}\oplus\mathbb{X}\).
Case 2: \(\mathbb{W}\) is irreducible. If \(\rho(\mathfrak{g})\mathbb{W} = 0\) then \(\mathbb{W}\) is one-dimensional, \(\dim\mathbb{V} = 2\), and every product of two elements of \(\rho(\mathfrak{g})\) vanishes because \(\rho(\mathfrak{g})\mathbb{V}\subseteq\mathbb{W}\); so \(\rho(\mathfrak{g})\) is abelian and \(\rho(\mathfrak{g}) = \rho\left(\comm{\mathfrak{g}}{\mathfrak{g}}\right) = 0\), and any complementary line serves. So assume \(\rho(\mathfrak{g})\mathbb{W} \neq 0\), and let \(\Gamma\) be the Casimir built from the trace form \(\beta_{\mathbb{V}}\) of Lemma A.402, which is available because \(\rho\) is faithful. By Lemma A.401 \(\Gamma\) commutes with the action; \(\Gamma\mathbb{V} \subseteq \rho(\mathfrak{g})\mathbb{V} \subseteq \mathbb{W}\), since \(\Gamma\) is a sum of products of two operators \(\rho(\cdot)\); hence \(\Gamma\) acts as zero on \(\mathbb{V}/\mathbb{W}\) and
by Lemma A.402. On the irreducible \(\mathbb{W}\), Lemma A.405 gives \(\Gamma|_{\mathbb{W}} = \lambda\identity\) with \(\lambda\dim\mathbb{W} = \dim\mathfrak{g}\), so \(\lambda \neq 0\) in characteristic zero and \(\Gamma|_{\mathbb{W}}\) is invertible. Therefore \(\Gamma\mathbb{V} = \mathbb{W}\), \(\dim\ker\Gamma = \dim\mathbb{V} - \dim\mathbb{W} = 1\), and \(\ker\Gamma \cap \mathbb{W} = 0\). As \(\Gamma\) commutes with the action, \(\ker\Gamma\) is a submodule, and \(\mathbb{V} = \mathbb{W}\oplus\ker\Gamma\).
∎Let \(\mathfrak{g}\) be semisimple over \(\mathbb{F}\) algebraically closed of characteristic zero. Every finite-dimensional \(\mathfrak{g}\)-module is a direct sum of irreducible submodules. Rests on Lemma A.406 and Corollary A.342.
Derives Theorem A.407. It suffices to show that every submodule \(\mathbb{W} \subseteq \mathbb{V}\) has a complementary submodule; the decomposition then follows by induction on \(\dim\mathbb{V}\). Assume \(0 \neq \mathbb{W} \neq \mathbb{V}\) and let \(\mathcal{H} = \operatorname{Hom}(\mathbb{V},\mathbb{W})\), a \(\mathfrak{g}\)-module under
Let \(\mathcal{V} \subseteq \mathcal{H}\) be the set of \(\phi\) whose restriction to \(\mathbb{W}\) is a scalar multiple of the identity, and \(\mathcal{W} \subseteq \mathcal{V}\) the set of \(\phi\) vanishing on \(\mathbb{W}\). Both are subspaces, and the map \(\mathcal{V} \to \mathbb{F}\) sending \(\phi\) to its scalar is linear, surjective — any projection of \(\mathbb{V}\) onto \(\mathbb{W}\) lies in \(\mathcal{V}\) with scalar \(1\) — and has kernel \(\mathcal{W}\), so \(\mathcal{W}\) has codimension one in \(\mathcal{V}\).
They are submodules: if \(\phi|_{\mathbb{W}} = \lambda\identity\) then for \(w \in \mathbb{W}\),
using that \(\mathbb{W}\) is a submodule; so \(X\cdot\mathcal{V} \subseteq \mathcal{W} \subseteq \mathcal{V}\).
Apply Lemma A.406 to \(\mathcal{W} \subseteq \mathcal{V}\): there is a one-dimensional submodule \(\mathbb{F}\phi_{0}\) complementary to \(\mathcal{W}\), and \(\phi_{0}\) may be normalized so that \(\phi_{0}|_{\mathbb{W}} = \identity\). Being a one-dimensional submodule with \(X\cdot\mathcal{V}\subseteq\mathcal{W}\), it satisfies \(X\cdot\phi_{0} \in \mathbb{F}\phi_{0}\cap\mathcal{W} = 0\): that is, \(\phi_{0}\) is a homomorphism of modules. Its kernel is therefore a submodule; it meets \(\mathbb{W}\) in \(0\) and has dimension \(\dim\mathbb{V} - \dim\mathbb{W}\) because \(\phi_{0}\) is onto \(\mathbb{W}\). Hence \(\mathbb{V} = \mathbb{W}\oplus\ker\phi_{0}\).
∎Trivial coefficients
Let \(\mathfrak{g}\) be semisimple over \(\mathbb{F}\) and let \(\mathbb{F}\) carry the trivial action. Then \(H^{1}\left(\mathfrak{g},\mathbb{F}\right) = H^{2}\left(\mathfrak{g},\mathbb{F}\right) = 0\). Rests on Theorem A.407, Corollary A.342 and Proposition 14.76.
Derives Proposition A.408. Degree one. With \(\rho = 0\), Equation (A.661) makes a \(1\)-cocycle a linear map \(f\) with \(f\left(\comm{X}{Y}\right) = 0\), that is, a linear map vanishing on \(\comm{\mathfrak{g}}{\mathfrak{g}} = \mathfrak{g}\) (Corollary A.342); so \(f = 0\). The coboundaries are zero as well, and \(H^{1} = 0\).
Degree two. Let \(c\) be a \(2\)-cocycle. By the construction in the proof of Proposition 14.76 — “every class is realized” — the vector space \(\tilde{\mathfrak{g}} = \mathfrak{g}\oplus\mathbb{F}Z\) with \(Z\) central and the bracket Equation (14.120) is a Lie algebra, the Jacobi identity being Equation (14.121), which by Remark A.399 is \(\dd c = 0\).
Make \(\tilde{\mathfrak{g}}\) a \(\mathfrak{g}\)-module. For \(\xi \in \mathfrak{g}\) and \(u \in \tilde{\mathfrak{g}}\) set \(\xi \cdot u = \comm{\tilde{\xi}}{u}\), where \(\tilde{\xi}\) is any preimage of \(\xi\) under the projection \(\pi : \tilde{\mathfrak{g}} \to \mathfrak{g}\). This is well defined, because two preimages differ by a multiple of the central \(Z\); and it is a module structure, because \(\comm{\widetilde{\comm{\xi}{\eta}}}{u} = \comm{\comm{\tilde{\xi}}{\tilde{\eta}}}{u}\) — the two lifts differ by a central element — and the right-hand side expands by the Jacobi identity of \(\tilde{\mathfrak{g}}\) into \(\comm{\tilde{\xi}}{\comm{\tilde{\eta}}{u}} - \comm{\tilde{\eta}}{\comm{\tilde{\xi}}{u}}\).
The line \(\mathbb{F}Z\) is a submodule, \(Z\) being central. By Theorem A.407 it has a complementary submodule \(\mathfrak{m}\), so \(\tilde{\mathfrak{g}} = \mathbb{F}Z\oplus\mathfrak{m}\) with \(\comm{\tilde{\mathfrak{g}}}{\mathfrak{m}} \subseteq \mathfrak{m}\); in particular \(\mathfrak{m}\) is a subalgebra, and \(\pi|_{\mathfrak{m}} : \mathfrak{m} \to \mathfrak{g}\) is a bijective homomorphism of Lie algebras. Its inverse \(s = \left(\pi|_{\mathfrak{m}}\right)^{-1}\) is a linear splitting of \(\pi\) which is also a homomorphism, so the cocycle it defines through Equation (14.120) is \(\comm{s(\xi)}{s(\eta)} - s\left(\comm{\xi}{\eta}\right) = 0\).
Finally, two linear splittings of \(\pi\) differ by a linear map \(\mathfrak{g} \to \mathbb{F}\), and by the first part of the proof of Proposition 14.76 the cocycles they define differ by a coboundary. The splitting \(\xi \mapsto (\xi,0)\) defines \(c\) and the splitting \(s\) defines \(0\); hence \(c\) is a coboundary and \(H^{2}\left(\mathfrak{g},\mathbb{F}\right) = 0\).
∎The Whitehead lemmas
Let \(\mathfrak{g}\) be semisimple over \(\mathbb{F}\) and let \(\mathbb{V}\) be an irreducible module with \(\rho \neq 0\). Then there is an invariant nondegenerate symmetric form \(\beta\) on \(\mathfrak{g}\) whose Casimir operator Equation (A.668) is invertible on \(\mathbb{V}\). Rests on Corollary A.342, Corollary A.343 and Lemma A.405.
Derives Lemma A.409. Write \(\mathfrak{g} = \mathfrak{g}_{1}\oplus\cdots\oplus\mathfrak{g}_{s}\) as a direct sum of simple ideals (Corollary A.342) and let \(\kappa_{j}\) be the Killing form of \(\mathfrak{g}_{j}\), extended to \(\mathfrak{g}\) by declaring the other summands orthogonal to \(\mathfrak{g}_{j}\) and to each other. Each \(\kappa_{j}\) is invariant and, by Corollary A.343, \(\kappa = \sum_{j}\kappa_{j}\). For nonzero scalars \(t_{1},\ldots,t_{s}\) put \(\beta_{t} = \sum_{j}t_{j}\kappa_{j}\), an invariant symmetric form, nondegenerate because each \(\kappa_{j}\) is nondegenerate on \(\mathfrak{g}_{j}\). Choosing a basis adapted to the decomposition, the \(\beta_{t}\)-dual basis is \(t_{j}^{-1}\) times the \(\kappa\)-dual basis in block \(j\), so
the inner sum running over a basis of \(\mathfrak{g}_{j}\) with its \(\kappa_{j}\)-dual. Each \(\Gamma_{j}\) commutes with \(\rho\left(\mathfrak{g}_{j}\right)\) by Lemma A.401 applied inside \(\mathfrak{g}_{j}\), and with \(\rho\left(\mathfrak{g}_{i}\right)\) for \(i \neq j\) because \(\comm{\mathfrak{g}_{i}}{\mathfrak{g}_{j}} = 0\); so it commutes with the whole action, and Lemma A.405 gives \(\Gamma_{j} = \lambda_{j}\identity\) on the irreducible \(\mathbb{V}\).
The \(\lambda_{j}\) are not all zero. Let \(\beta_{\mathbb{V}}\) be the trace form Equation (A.672), which is invariant but need not be nondegenerate. Its restriction to the simple ideal \(\mathfrak{g}_{j}\) is a multiple of \(\kappa_{j}\): writing \(\beta_{\mathbb{V}}(X,Y) = \kappa_{j}(AX,Y)\) for \(X,Y \in \mathfrak{g}_{j}\) with \(A \in \operatorname{End} \left(\mathfrak{g}_{j}\right)\), invariance of both forms gives
for all \(Z \in \mathfrak{g}_{j}\), so \(A\) commutes with every \(\ad_{Z}\). The adjoint representation of a simple algebra is irreducible — its submodules are its ideals — so Lemma A.405 gives \(A = c_{j}\identity\) and \(\beta_{\mathbb{V}}|_{\mathfrak{g}_{j}} = c_{j}\kappa_{j}\). Now
If \(c_{j} = 0\) then \(\beta_{\mathbb{V}}\) vanishes on \(\mathfrak{g}_{j}\), so \(\tr\left(\rho(X)\rho(Y)\right) = 0\) for all \(X \in \comm{\mathfrak{g}_{j}}{\mathfrak{g}_{j}} = \mathfrak{g}_{j}\) and \(Y \in \mathfrak{g}_{j}\), and Theorem A.340 makes \(\rho\left(\mathfrak{g}_{j}\right)\) solvable; being a homomorphic image of the simple \(\mathfrak{g}_{j}\) it is either isomorphic to \(\mathfrak{g}_{j}\), which is not solvable, or zero. So \(c_{j} = 0\) forces \(\rho\left(\mathfrak{g}_{j}\right) = 0\). Since \(\rho \neq 0\), some \(\mathfrak{g}_{j_{0}}\) acts nontrivially and \(\lambda_{j_{0}} \neq 0\).
Choice of \(t\). The eigenvalue of \(\Gamma_{t}\) on \(\mathbb{V}\) is \(\sum_{j}t_{j}^{-1}\lambda_{j}\). If \(\sum_{j}\lambda_{j} \neq 0\) take every \(t_{j} = 1\). Otherwise take \(t_{j_{0}} = 2\) and \(t_{j} = 1\) for \(j \neq j_{0}\); the eigenvalue is then \(\sum_{j}\lambda_{j} - \tfrac{1}{2}\lambda_{j_{0}} = -\tfrac{1}{2}\lambda_{j_{0}} \neq 0\). In either case \(\Gamma_{t}\) is a nonzero scalar on \(\mathbb{V}\), hence invertible.
∎Let \(\mathfrak{g}\) be a finite-dimensional semisimple Lie algebra over a field of characteristic zero and let \(\mathbb{V}\) be any finite-dimensional \(\mathfrak{g}\)-module. Then
Rests on Theorem A.403, Theorem A.407 and Lemma A.409.
Derives Theorem A.410. Reduction to an algebraically closed field. Each \(C^{n}(\mathfrak{g},\mathbb{V})\) is a finite-dimensional \(\mathbb{K}\)-vector space and \(\dd\) is \(\mathbb{K}\)-linear; extending scalars to \(\mathbb{F}\) gives \(C^{n}(\mathfrak{g},\mathbb{V})\otimes_{\mathbb{K}}\mathbb{F} = C^{n}\left(\mathfrak{g}_{\mathbb{F}}, \mathbb{V}_{\mathbb{F}}\right)\), with the same differential, because a multilinear map is fixed by its values on a \(\mathbb{K}\)-basis. The ranks of \(\dd\) in each degree are unchanged by field extension, so \(H^{n}(\mathfrak{g},\mathbb{V})\otimes_{\mathbb{K}}\mathbb{F} = H^{n}\left(\mathfrak{g}_{\mathbb{F}}, \mathbb{V}_{\mathbb{F}}\right)\), and a space vanishes if and only if its extension does. Moreover \(\mathfrak{g}_{\mathbb{F}}\) is again semisimple by Remark A.341. It therefore suffices to prove the theorem over \(\mathbb{F}\).
Reduction to irreducible coefficients. By Theorem A.407, \(\mathbb{V} = \mathbb{V}_{1}\oplus\cdots\oplus\mathbb{V}_{r}\) with each \(\mathbb{V}_{i}\) irreducible. A cochain with values in a direct sum is the same thing as a tuple of cochains with values in the summands, and \(\dd\) acts componentwise, so
The two cases. If \(\mathfrak{g}\) acts as zero on the irreducible \(\mathbb{V}_{i}\), then every subspace is a submodule, so \(\dim\mathbb{V}_{i} = 1\) and \(\mathbb{V}_{i} \cong \mathbb{F}\) with the trivial action; Proposition A.408 gives \(H^{1} = H^{2} = 0\). Otherwise Lemma A.409 supplies an invariant nondegenerate form whose Casimir is invertible on \(\mathbb{V}_{i}\), and Theorem A.403 gives \(H^{1} = H^{2} = 0\). Summing over \(i\) in Equation (A.692) completes the proof.
∎A finite-dimensional semisimple Lie algebra \(\mathfrak{g}\) over a field of characteristic zero satisfies \(H^{1}(\mathfrak{g},\R) = H^{2}(\mathfrak{g},\R) = 0\) and admits no nontrivial central extension. In particular, for \(D = p+q \ge 3\) the algebra \(\mathfrak{so}(p,q)\) — and hence the Lorentz algebra \(\mathfrak{so}(D-1,1)\) of Equation (14.90), and the rotation algebra \(\mathfrak{so}(3)\) — admits no nontrivial central extension. Rests on Theorem A.410, Corollary A.344 and Proposition 14.76.
Derives Corollary A.411. Take \(\mathbb{V} = \R\) with the trivial action in Theorem A.410; by Remark A.399 the group \(H^{2}\) so computed is the group \(H^{2}(\mathfrak{g},\R)\) of Definition 14.75, and by Proposition 14.76 its vanishing says that every central extension of \(\mathfrak{g}\) by \(\R\) is trivial, that is, isomorphic to \(\mathfrak{g}\oplus\R\) as a Lie algebra. That \(\mathfrak{so}(p,q)\) is semisimple for \(D \ge 3\) is Corollary A.344, proved there by computing its Killing form and applying Theorem A.330.
∎Each hypothesis of Theorem A.410 earns its place, and the counterexamples are the physical cases of Section 14.4.2. Semisimplicity: the abelian algebra \(\R^{2f}\) of Example 14.79 has \(H^{2} = \Lambda^{2}\left(\R^{2f}\right)^{*}\), as far from zero as possible, and the extension is the Heisenberg algebra with \(Z = \ii\hbar\); the proof fails at the first step, since \(\comm{\mathfrak{g}}{\mathfrak{g}} = 0\) and the Killing form vanishes identically. Finite dimension: the Witt algebra of Example 14.81 has a one-dimensional \(H^{2}\) whose extension is the Virasoro algebra, and the Casimir Equation (A.668) is an infinite sum there, with no meaning as an operator on the algebra. Characteristic zero: the division by \(\lambda\) in Lemma A.406 and by \(\dim\mathbb{V}\) in Equation (A.690) both fail in characteristic \(p\), and so does the theorem.
The Galilei algebra of Example 14.80 is the instructive intermediate case: it is finite-dimensional over \(\R\) but not semisimple — it has the translations and the boosts as abelian subalgebras and is a contraction (Section 14.5) of an algebra that is — and it carries the mass as a central charge. So the mass is not an accident of the nonrelativistic limit that a better choice of generators could remove; it is an obstruction that exists because the algebra fails the hypothesis of this theorem.
Whitehead's Lemmas and the Rigidity of Semisimple Algebras discharges the proof obligation of Proposition 14.77. Its statement is used at Remark 14.78 to localize where nontrivial central charges can live — not in semisimple algebras, hence only in abelian algebras, in inhomogeneous semidirect sums such as Equation (14.99), in contractions (Section 14.5) and in infinite-dimensional algebras — and at Remark 14.85, where the cohomological reading of a contraction is set out. Two by-products proved along the way have independent value and are used tacitly elsewhere in the chapter: Weyl's complete reducibility theorem (Theorem A.407), which is why a finite-dimensional representation of a semisimple algebra may always be decomposed into irreducibles as Theorem 14.47 does for \(\mathfrak{su}(2)\), and the vanishing of \(H^{1}(\mathfrak{g},\mathbb{V})\), which is the statement that a semisimple algebra admits no nontrivial deformation of a module structure.
The Second Cohomology of the Galilei Algebra
This appendix proves the assertion left owed in Example 14.80 of Lie Groups, Lie Algebras, and Fibre Bundles: the second Chevalley–Eilenberg cohomology \(H^{2}(\mathfrak{g},\R)\) of the ten-generator Galilei algebra of three space dimensions is one-dimensional, so that up to a scale the mass extension is the only central extension the algebra admits.
Exactly one half of that statement is proved elsewhere in this treatise. Proposition 25.14, in the chapter bridging Poisson brackets and quantization, shows that the mass cocycle \(c(K_{i},P_{j}) = m\,\delta_{ij}\) of Equation (25.15) is not a coboundary — that its class in \(H^{2}\) is nonzero. It says nothing about how large \(H^{2}\) is, and the chapter records the difference honestly. What is added here is the other half: that the class of the mass cocycle exhausts the cohomology. Nothing in Lie Groups, Lie Algebras, and Fibre Bundles or in Part III needs the stronger statement, so this appendix closes a gap in the mathematics rather than supplying a missing physical input.
The computation is finite and is carried out in full. The space of alternating bilinear forms on a ten-dimensional algebra has dimension \(\binom{10}{2} = 45\); the cocycle condition Equation (14.121) cuts it down to \(10\), the coboundaries account for \(9\) of those, and the survivor is the mass.
The algebra, and the two spaces
Let \(\mathfrak{g}\) be the real ten-dimensional Lie algebra spanned by rotations \(J_{i}\), boosts \(K_{i}\), space translations \(P_{i}\) (\(i = 1,2,3\)) and the time translation \(H\), with the isotropy relations Equation (A.161),
and the single further nonvanishing bracket
This is the case \((\gamma,\mu,\nu) = (0,\ast,0)\) of the enumeration in Step four: the enumeration, with \(\alpha = -\gamma\mu = 0\) and \(\beta = \gamma\nu = 0\); the normalisation \(\mu = 1\) is the one fixed there. All indices are Cartesian, repeated indices are summed, and \(\epsilon_{ijk}\) is the Levi-Civita symbol. Rests on Equation (A.161), Equation (A.165) and Theorem 14.90.
\(\dim_{\R}H^{2}(\mathfrak{g},\R) = 1\), and the class is represented by the mass cocycle
Every central extension of \(\mathfrak{g}\) by \(\R\) is therefore equivalent to the one with bracket \(\comm{\tilde{K}_{i}}{\tilde{P}_{j}} = m\,\delta_{ij}Z\) for a single real number \(m\), and it is trivial exactly when \(m = 0\). Rests on Definition A.413, Definition 14.75 and Proposition 14.76.
A cocycle is fixed by its values on pairs of basis generators. Grouping those values by type, an alternating bilinear \(c : \mathfrak{g}\times\mathfrak{g} \rightarrow \R\) is exactly the datum
with \(A\), \(D\), \(F\) antisymmetric and \(B\), \(C\), \(E\) unrestricted: \(3+9+9+3+3+9+3+3+3 = 45\) real parameters, as it must be.
Three tensor lemmas
The cocycle conditions that involve one or two rotation generators are statements about \(\SO(3)\)-covariance of the arrays in Equation (A.696), and all of them reduce to two identities. They are isolated here so that the enumeration below reads cleanly. Throughout, the only inputs are the standard contractions
A real array \(T_{jk}\) satisfies
if and only if \(T_{jk} = t\,\delta_{jk}\) for a real number \(t\). Rests on Equation (A.697).
Derives Lemma A.415. Necessity. Contract Equation (A.698) with \(\epsilon_{ijm}\) and sum over \(i\) and \(j\). By the first identity of Equation (A.697) the first term gives \(2T_{mk}\); by the second, the remaining term gives
where \(\tr T = T_{jj}\). Hence
Exchanging \(m\) and \(k\) in Equation (A.699) and subtracting gives \(T_{mk} - T_{km} = 0\), so \(T\) is symmetric; adding instead gives \(3\left(T_{mk}+T_{km}\right) = 2\delta_{mk}\tr T\), that is \(T_{mk} = \tfrac{1}{3}\left(\tr T\right)\delta_{mk}\).
Sufficiency. For \(T = \delta\) the left side of Equation (A.698) is \(\epsilon_{ijk} + \epsilon_{ikj} = 0\).
∎A real array \(T_{jk}\) satisfies
if and only if \(T\) is antisymmetric, that is \(T_{jk} = \epsilon_{jkm}t_{m}\) for a real vector \(t\). Rests on Equation (A.697).
Derives Lemma A.416. Necessity. Contract with \(\epsilon_{ijm}\) as before. The first two terms give \(2T_{mk}\) and \(T_{km} - \delta_{mk}\tr T\) exactly as in Lemma A.415. For the third, use \(\epsilon_{ijm}\epsilon_{jkl} = -\epsilon_{jim}\epsilon_{jkl} = -\left(\delta_{ik}\delta_{ml}-\delta_{il}\delta_{mk}\right)\), so that
Summing the three,
Setting \(m = k\) and summing gives \(2\tr T = 3\tr T\), so \(\tr T = 0\), and then Equation (A.701) reads \(T_{mk} = -T_{km}\).
Sufficiency. Put \(T_{jk} = \epsilon_{jkm}t_{m}\) and evaluate the three terms of Equation (A.700) with the third identity of Equation (A.697), applied after cycling the symbols into the shape \(\epsilon_{ab l}\epsilon_{lcd}\):
The first two add to \(\delta_{ik}t_{j} - \delta_{ij}t_{k}\), which is the third; the alternating sum vanishes.
∎Lemmas A.415 and A.416 are the only steps of the computation that are not pure bookkeeping, and both are statements about the invariant two-tensors of the rotation group of three space dimensions: they use the Levi-Civita symbol with three indices and the contraction identities Equation (A.697), which are three-dimensional. The theorem proved here is accordingly a theorem about the Galilei algebra of the observed three space dimensions, which is the only case this treatise instantiates; no claim is made, and none is needed, for any other number of space dimensions.
Solving the cocycle condition
Write \(\delta c(x,y,z)\) for the left-hand side of Equation (14.121),
which is alternating in \((x,y,z)\), so it suffices to impose \(\delta c = 0\) on unordered triples of basis generators; a triple containing \(H\) twice gives nothing, since there is only one \(H\). Sixteen types remain, and every one is evaluated below.
Triples with no bracket. For \((K_{i},K_{j},K_{k})\), \((P_{i},P_{j},P_{k})\), \((K_{i},K_{j},P_{k})\), \((K_{i},P_{j},P_{k})\) and \((P_{i},P_{j},H)\) every bracket appearing in Equation (A.702) vanishes by Equation (A.694), so the condition is empty.
The triple \((J_{i},J_{j},J_{k})\). Writing \(A_{ij} = \epsilon_{ijm}a_{m}\) and using the third identity of Equation (A.697), \(\epsilon_{ijl}A_{lk} = \delta_{ik}a_{j} - \delta_{jk}a_{i}\), so the cyclic sum is
identically. The array \(A\) is unconstrained, and it appears in no other triple: a term \(c(\comm{\cdot}{\cdot},J)\) can involve \(A\) only when the bracket produces a rotation, and only \(\comm{J}{J}\) does.
The triple \((J_{i},J_{j},K_{k})\). Using Equation (A.693) three times,
which is Equation (A.700). By Lemma A.416,
for an arbitrary vector \(\beta\), and nothing more.
The triple \((J_{i},J_{j},P_{k})\). Identical in form, with \(C\) in place of \(B\): it forces \(C\) to be antisymmetric and imposes nothing else. It will be subsumed by the stronger condition below.
The triple \((J_{i},J_{j},H)\). Since \(\comm{J_{i}}{H} = \comm{H}{J_{j}} = 0\), only the first term survives: \(\delta c = \epsilon_{ijl}u_{l} = 0\) for all \(i,j\), hence
The triple \((J_{i},K_{j},K_{k})\). Here \(\comm{K_{j}}{K_{k}} = 0\), so
the second step using the antisymmetry of \(D\). This is Equation (A.698), so \(D = t\delta\) by Lemma A.415; but \(D\) is antisymmetric and \(\delta\) is not, so \(t = 0\) and
The triple \((J_{i},P_{j},P_{k})\). Word for word the same with \(F\) in place of \(D\), giving \(F_{ij} = c(P_{i},P_{j}) = 0\).
The triple \((J_{i},K_{j},H)\). Now the bracket \(\comm{K_{j}}{H} = P_{j}\) contributes:
so that
The array \(C\) is thus determined by the vector \(v\), and is automatically antisymmetric — consistent with, and stronger than, the \((J,J,P)\) condition.
The triple \((J_{i},P_{j},H)\). Both \(\comm{P_{j}}{H}\) and \(\comm{H}{J_{i}}\) vanish, leaving \(\delta c = \epsilon_{ijl}w_{l} = 0\), hence
The triple \((J_{i},K_{j},P_{k})\). With \(\comm{K_{j}}{P_{k}} = 0\),
which is Equation (A.698) again. By Lemma A.415,
The triple \((K_{i},K_{j},H)\). Only the brackets with \(H\) survive:
which says \(E\) is symmetric — already true of Equation (A.708), so no new information.
The triple \((K_{i},P_{j},H)\). Here \(\comm{K_{i}}{P_{j}} = 0\) and \(\comm{P_{j}}{H} = 0\), and \(\comm{H}{K_{i}} = -P_{i}\), so \(\delta c = c(-P_{i},P_{j}) = -F_{ij} = 0\): again already known.
That exhausts the sixteen types. Collecting Equations (A.703), (A.704), (A.705), (A.706), (A.707) and (A.708) and the two vanishing statements for \(F\):
A bilinear form on \(\mathfrak{g}\) is a \(2\)-cocycle in the sense of Proposition 14.73 if and only if
with \(u = w = 0\) and \(D = F = 0\), for arbitrary vectors \(a,\beta,v \in \R^{3}\) and an arbitrary real number \(e\). Hence \(\dim_{\R}Z^{2}(\mathfrak{g},\R) = 3+3+3+1 = 10\). Rests on Definition A.413, Lemma A.415 and Lemma A.416.
The coboundaries, and the quotient
\(\dim_{\R}B^{2}(\mathfrak{g},\R) = 9\), and a coboundary is exactly a cocycle Equation (A.709) with \(e = 0\). Rests on Definition 14.74, Definition A.413 and Proposition A.418.
Derives Proposition A.419. Let \(b : \mathfrak{g} \rightarrow \R\) be linear, with components \(b(J_{i}) = \beta^{J}_{i}\), \(b(K_{i}) = \beta^{K}_{i}\), \(b(P_{i}) = \beta^{P}_{i}\) and \(b(H) = \eta\). Evaluating \((\partial b)(x,y) = b\left(\comm{x}{y}\right)\) of Equation (14.122) on each pair of basis generators with Equations (A.693) and (A.694),
and \((\partial b)(K_{i},K_{j}) = (\partial b)(P_{i},P_{j}) = (\partial b)(P_{i},H) = 0\). In the coordinates of Equation (A.709) this is
where the third entry is consistent in both places it occurs: \(C_{ij} = \epsilon_{ijk}\beta^{P}_{k}\) and \(c(K_{i},H) = \beta^{P}_{i}\) are exactly the pair required by Equation (A.709). Two consequences follow at once. First, every coboundary is a cocycle with \(e = 0\) — which also re-proves Proposition 14.73 for this algebra. Second, the map \(b \mapsto \partial b\) carries \(\left(\beta^{J},\beta^{K},\beta^{P}\right)\) onto the whole nine-dimensional subspace \(\set{e = 0}\) of \(Z^{2}\), bijectively, while \(\eta\) is annihilated. Hence \(\dim B^{2} = 9\).
The kernel is visible structurally as well: \(\partial b = 0\) if and only if \(b\) vanishes on the derived algebra \(\comm{\mathfrak{g}}{\mathfrak{g}}\), which by Equations (A.693) and (A.694) is the span of the \(J_{i}\), \(K_{i}\) and \(P_{i}\) — the \(P_{i}\) arising both from \(\comm{J}{P}\) and from \(\comm{K}{H}\) — of dimension \(9\). No bracket produces \(H\), so \(H\) is not in the derived algebra and the kernel is the line \(\set{\beta^{J}=\beta^{K}=\beta^{P}=0}\), of dimension \(10 - 9 = 1\).
∎Proof of Theorem A.414. Derives Theorem A.414. By Propositions A.418 and A.419 the coboundaries are the hyperplane \(\set{e = 0}\) inside the ten-dimensional space of cocycles, so
and the coordinate \(e\) descends to an isomorphism \(H^{2}(\mathfrak{g},\R) \cong \R\). The cocycle \(c_{M}\) of Equation (A.695) has \(a = \beta = v = 0\) and \(e = 1\): it satisfies Equation (A.709), hence is a cocycle, and its class is a generator. The final statement is Proposition 14.76: central extensions of \(\mathfrak{g}\) by \(\R\) correspond to classes in \(H^{2}\), so they form a one-parameter family \(m\,[c_{M}]\), realized by \(\comm{\tilde{K}_{i}}{\tilde{P}_{j}} = m\,\delta_{ij}Z\) through Equation (14.120), and by Equation (14.123) the extension is trivial precisely for the zero class, \(m = 0\).
∎Proposition 25.14 proves that the class of \(c_{M}\) is nonzero, by exhibiting the Poisson-bracket realisation of a free particle and showing that no shift of the generators by constants removes the defect Equation (25.15). The present appendix reproves that as the single line \(e \neq 0\) in Equation (A.712) — coboundaries have \(e = 0\), \(c_{M}\) has \(e = 1\) — and adds the converse, that there is nothing else: the whole of \(H^{2}\) is the mass.
Two structural readings of the computation are worth recording, because they explain the answer rather than merely producing it. The first is that the surviving parameter is the one component of \(c\) that no coboundary can reach, namely \(c(K_{i},P_{j})\), and it is unreachable because \(\comm{K_{i}}{P_{j}} = 0\) in \(\mathfrak{g}\): a coboundary is a linear functional composed with the bracket, and there is no bracket there to compose with. That is exactly the argument of Proposition 25.14, and the enumeration above shows it is the only such place. The second is that the answer is forced to be at most one-dimensional by rotational covariance: Lemma A.415 says the surviving array \(c(K_{i},P_{j})\) has nowhere to live but the invariant \(\delta_{ij}\), so the mass is a single number and not a tensor.
Example 14.80 may therefore be read without its final caveat, and Remark 14.88 — which describes how the mass emerges from the Inönü–Wigner limit of a trivially extended Poincaré algebra — is now accompanied by an independent statement of where the limit must land: \(H^{2}\) of the contracted algebra is one-dimensional, so any nontrivial extension produced by any route is a multiple of \(c_{M}\). This is a concrete instance of the phenomenon named in Remark 14.78: a contraction can create a cohomology class that the parent semisimple algebra, rigid by Proposition 14.77, did not have.
Central Extensions of Three Infinite-Dimensional Algebras
Example 14.81 of Lie Groups, Lie Algebras, and Fibre Bundles lists three infinite-dimensional algebras whose central extensions are nontrivial: the Witt algebra of vector fields on the circle, whose extension is the Virasoro algebra; the loop algebra of a simple Lie algebra, whose extension is affine Kac–Moody; and the equal-time current algebra of a field theory, whose extension is the Schwinger term. This appendix supplies the three computations.
They are of two different kinds, and the difference is the point of the section. The first two are finite algebraic computations in the sense of Definition 14.75: a cocycle is written down, the cocycle condition Equation (14.121) is solved, and the coboundaries Equation (14.122) are divided out. The third is not. There is no finite computation of a Schwinger term, because the object whose commutator is asked for — a product of operator-valued distributions at a point — does not exist until it is regularized. What can be proved, and is proved below, is a no-go: in any theory with a Hamiltonian bounded below, a normalizable ground state and a local conserved current, the equal-time commutator of the charge density with the current density cannot vanish, and the leading term it must carry is a central one of definite sign. Remark A.434 states exactly which hypotheses that argument consumes and which question it leaves open.
Throughout, the ground field is written \(F\) and may be taken to be \(\R\) or \(\C\); Definition 14.72, Proposition 14.73, Definition 14.74, Definition 14.75 and Proposition 14.76 are stated in Section 14.4.1 for \(\R\), and their statements and proofs use nothing but the field axioms in characteristic zero, so they hold verbatim over \(\C\). Rule 1 of this treatise applies here as everywhere: what follows is mathematics about Lie algebras, and no claim is made about any physical theory in which these algebras have been proposed to act.
The Virasoro cocycle
The Witt algebra \(\mathfrak{w}\) is the Lie algebra over \(F\) with basis \(\set{L_{n}}_{n\in\Z}\) and brackets
It is the algebra of polynomial vector fields on the circle: with \(t = \ee^{\ii\theta}\) and \(L_{n} = -t^{n+1}\dv{}{t}\) acting on \(F[t,t^{-1}]\), a direct computation reproduces Equation (A.713). The Jacobi identity is the identity \((m-n)(m+n-k) + (n-k)(n+k-m) + (k-m)(k+m-n) = 0\), which holds because every quadratic monomial cancels in pairs. Rests on Definition 14.72.
Note that \(\mathfrak{w}\) is infinite-dimensional, so a bilinear form on it is an arbitrary array of values \(c(L_{m},L_{n})\); bilinearity constrains only finite linear combinations, and no continuity or boundedness is imposed anywhere below.
\(H^{2}(\mathfrak{w},F)\) is one-dimensional, generated by the class of
Every central extension of \(\mathfrak{w}\) by \(F\) is therefore equivalent, for exactly one \(C \in F\), to the Virasoro algebra
Rests on Definition A.421, Definition 14.75 and Proposition 14.76.
Derives Theorem A.422. Let \(c\) be a \(2\)-cocycle on \(\mathfrak{w}\).
Step 1: normalize against \(L_{0}\). Define a linear functional \(b\) on \(\mathfrak{w}\) by
Its coboundary is \((\partial b)(L_{m},L_{n}) = b\left(\comm{L_{m}}{L_{n}}\right) = (m-n)\,b(L_{m+n})\), so for \(n \neq 0\)
while \((\partial b)(L_{0},L_{0}) = 0 = c(L_{0},L_{0})\) by antisymmetry. Replacing \(c\) by \(c - \partial b\) — which changes nothing in \(H^{2}\) — we may and do assume
Step 2: only the diagonal \(m+n = 0\) survives. Impose the cocycle condition Equation (14.121) on the triple \(\left(L_{0},L_{m},L_{n}\right)\). With \(\comm{L_{0}}{L_{m}} = -mL_{m}\), \(\comm{L_{n}}{L_{0}} = nL_{n}\) and \(\comm{L_{m}}{L_{n}} = (m-n)L_{m+n}\),
the middle term vanishing by Equation (A.717) and the last by antisymmetry. Hence
and antisymmetry gives \(h(-m) = -h(m)\), with \(h(0) = 0\).
Step 3: the recursion. A cocycle of the shape Equation (A.718) makes every term of Equation (14.121) vanish on a triple \(\left(L_{m},L_{n},L_{k}\right)\) with \(m+n+k \neq 0\), since each term pairs \(L_{m+n}\) with \(L_{k}\) and the two indices sum to \(m+n+k\). For \(m+n+k = 0\) the three terms are
so the cocycle condition is \((m-n)h(k) + (n-k)h(m) + (k-m)h(n) = 0\) with \(k = -m-n\), that is
Setting \(n = 1\),
which for \(m \ge 2\) determines \(h(m+1)\) from \(h(m)\) and \(h(1)\). Since \(h(2)\) is not reached by Equation (A.720) — at \(m = 1\) the left-hand side vanishes identically, and indeed the relation then reads \(3h(1) = 3h(1)\) — the solution space has dimension at most \(2\), the free data being \(h(1)\) and \(h(2)\); negative arguments follow from \(h(-m) = -h(m)\).
Two solutions are exhibited directly. For \(h(m) = m\),
and for \(h(m) = m^{3}\), expanding both sides,
where the middle expression is obtained by multiplying out \((m-n)\left(m^{3}+3m^{2}n+3mn^{2}+n^{3}\right)\) and cancelling the two \(3m^{2}n^{2}\) terms. The two are linearly independent, so the solution space of Equation (A.719) is exactly \(\set{h(m) = \alpha m + \beta m^{3}}\).
Step 4: the linear solution is a coboundary. Take \(b_{\gamma}\left(L_{n}\right) = \gamma\,\delta_{n,0}\). Then \(\left(\partial b_{\gamma}\right)(L_{m},L_{n}) = (m-n)\gamma\,\delta_{m+n,0} = 2m\gamma\,\delta_{m+n,0}\), which is Equation (A.718) with \(h(m) = 2\gamma m\); and it respects Equation (A.717), since \(\left(\partial b_{\gamma}\right)(L_{n},L_{0}) = n\,b_{\gamma}(L_{n}) = 0\) for \(n \neq 0\). Choosing \(\gamma = \alpha/2\) removes the \(\alpha\) term, so every class is represented by a multiple of \(h(m) = m^{3}\), and equally by a multiple of Equation (A.714), since \((m^{3}-m)/12\) is such a combination. Thus \(\dim H^{2}(\mathfrak{w},F) \le 1\).
Step 5: \(\omega\) is not a coboundary. Every coboundary satisfies \(\left(\partial b\right)\left(L_{m},L_{-m}\right) = b\left(\comm{L_{m}}{L_{-m}}\right) = 2m\,b\left(L_{0}\right)\), a linear function of \(m\). If \(\omega = \partial b\) then \((m^{3}-m)/12 = 2m\,b(L_{0})\) for every \(m\); at \(m = 2\) this gives \(b(L_{0}) = \tfrac{1}{8}\) and at \(m = 3\) it gives \(b(L_{0}) = \tfrac{1}{3}\), a contradiction. Hence \(\dim H^{2}(\mathfrak{w},F) = 1\) and \(\omega\) generates it.
The last assertion is Proposition 14.76: extensions up to equivalence are in bijection with classes, the class of \(C\,\omega\) corresponds to the bracket Equation (A.715) through Equation (14.120), and by Equation (14.123) the extension is trivial precisely for \(C = 0\).
∎Step 4 shows that the representative of a class is fixed only up to a multiple of \(h(m) = m\), so the choice Equation (A.714) is a convention and not a result. It is the convention that makes \(\omega\) vanish on the three-dimensional subalgebra spanned by \(L_{-1}, L_{0}, L_{1}\): those are the values \(m \in \set{-1,0,1}\), at which \(m^{3}-m = 0\). That subalgebra is closed — \(\comm{L_{1}}{L_{-1}} = 2L_{0}\), \(\comm{L_{0}}{L_{\pm1}} = \mp L_{\pm1}\) — and is isomorphic to \(\mathfrak{sl}(2,F)\), which by Proposition 14.77 admits no nontrivial central extension at all; the normalization therefore makes the restriction of the extension to it not merely trivial in cohomology but zero on the nose. The factor \(12\) is chosen so that the coefficient \(C\) in Equation (A.715) takes the value \(1\) for the smallest extension usually singled out; nothing in the mathematics depends on it.
The Kac–Moody cocycle
Let \(\mathfrak{g}\) be a finite-dimensional simple Lie algebra over \(\C\) with Killing form \(\kappa\) (Definition 14.10). The loop algebra is
which is a Lie algebra because the bracket of \(\mathfrak{g}\) is and \(\C[t,t^{-1}]\) is commutative and associative. For a Laurent polynomial \(f = \sum_{n}a_{n}t^{n}\) the residue is \(\operatorname{Res}f = a_{-1}\). Rests on Definitions 14.10 and 14.72.
\(\operatorname{Res}f' = 0\) for every \(f \in \C[t,t^{-1}]\), and consequently \(\operatorname{Res}\left(f'g\right) = -\operatorname{Res}\left(fg'\right)\) for all \(f,g\). Rests on Definition A.424.
Derives Lemma A.425. \(\left(t^{n}\right)' = n\,t^{n-1}\), whose exponent is \(-1\) only for \(n = 0\), where the coefficient \(n\) vanishes; so no \(t^{-1}\) term is ever produced and \(\operatorname{Res}f' = 0\) by linearity. Applying this to \(fg\) and using the Leibniz rule, \(0 = \operatorname{Res}\left(fg\right)' = \operatorname{Res}\left(f'g\right) + \operatorname{Res}\left(fg'\right)\).
∎The bilinear map
is an antisymmetric \(2\)-cocycle on \(L\mathfrak{g}\) and is not a coboundary. The corresponding central extension is the affine Kac–Moody algebra
with \(\ell \in \C\) the level and \(Z\) central. Rests on Definition A.424, Lemma A.425 and Lemma 14.11.
Derives Theorem A.426. The two forms agree. For \(f = t^{m}\), \(g = t^{n}\) one has \(f'g = m\,t^{m+n-1}\), whose residue is \(m\,\delta_{m+n,0}\).
Antisymmetry. \(\kappa\) is symmetric (Definition 14.10), so
by Lemma A.425.
The cocycle condition. Let \(u = x\otimes f\), \(v = y\otimes g\), \(w = z\otimes h\). By Equation (A.721),
and likewise for the two cyclic permutations. The three Killing-form factors are equal. Indeed Lemma 14.11 reads \(\kappa\left(\comm{X}{Y},Z\right) = -\kappa\left(Y,\comm{X}{Z}\right)\); taking \(X = y\), \(Y = x\), \(Z = z\) and using antisymmetry of the bracket gives \(\kappa\left(\comm{x}{y},z\right) = \kappa\left(x,\comm{y}{z}\right)\), and the symmetry of \(\kappa\) turns the right side into \(\kappa\left(\comm{y}{z},x\right)\); applying the same step again gives \(\kappa\left(\comm{z}{x},y\right)\). Write \(\mathcal{K}\) for the common value. The cyclic sum is therefore
Expanding each derivative by the Leibniz rule,
and \(\operatorname{Res}\) of a derivative vanishes by Lemma A.425. So Equation (A.724) is zero, which is Equation (14.121).
It is not a coboundary. Since \(\mathfrak{g}\) is simple, \(\comm{\mathfrak{g}}{\mathfrak{g}}\) is a nonzero ideal, hence all of \(\mathfrak{g}\); so \(\kappa \neq 0\), and there is \(x \in \mathfrak{g}\) with \(\kappa(x,x) \neq 0\) — were \(\kappa(x,x) = 0\) for every \(x\), the polarization identity \(2\kappa(x,y) = \kappa(x+y,x+y) - \kappa(x,x) - \kappa(y,y)\) would make \(\kappa\) vanish identically, contradicting its nondegeneracy (Theorem 14.13). Put \(u = x\otimes t\) and \(v = x\otimes t^{-1}\). Then \(\comm{u}{v} = \comm{x}{x}\otimes 1 = 0\), so \(\left(\partial b\right)(u,v) = b\left(\comm{u}{v}\right) = 0\) for every linear \(b\), whereas \(\omega_{\kappa}(u,v) = 1\cdot\kappa(x,x) \neq 0\) by Equation (A.722). The extension Equation (A.723) is then Equation (14.120) for the cocycle \(\ell\,\omega_{\kappa}\), and is nontrivial for \(\ell \neq 0\) by Equation (14.123).
∎The structure of Equation (A.722) is worth naming: it is the product of the unique invariant pairing on \(\mathfrak{g}\) with the unique invariant pairing on \(\C[t,t^{-1}]\) that is antisymmetric, namely \((f,g)\mapsto\operatorname{Res}(f'g)\). The next proposition makes the first “unique” precise and shows that the product is forced.
Let \(\mathfrak{g}\) be finite-dimensional and simple over \(\C\). Every bilinear \(B : \mathfrak{g}\times\mathfrak{g}\rightarrow\C\) with
is a multiple of the Killing form; in particular it is symmetric. Rests on Lemma 14.11, Theorem 14.13 and Theorem 5.153.
Derives Lemma A.427. Give \(\mathfrak{g}^{*}\) the coadjoint action \(\left(z\cdot\phi\right)(y) = -\phi\left(\comm{z}{y}\right)\). For a bilinear \(B\) satisfying Equation (A.725) the map \(\Phi_{B} : \mathfrak{g}\rightarrow\mathfrak{g}^{*}\), \(\Phi_{B}(x) = B(x,\cdot)\), obeys
so \(\Phi_{B}\) intertwines the adjoint and coadjoint actions. The Killing form satisfies Equation (A.725) by Lemma 14.11 and is nondegenerate by Theorem 14.13, so \(\Phi_{\kappa}\) is an isomorphism of \(\mathfrak{g}\)-modules. Hence \(T = \Phi_{\kappa}^{-1}\circ\Phi_{B}\) is an endomorphism of \(\mathfrak{g}\) commuting with every \(\ad_{z}\). The adjoint representation of a simple algebra on the finite-dimensional complex space \(\mathfrak{g}\) is irreducible — an invariant subspace is an ideal — so Schur's lemma, in the form of Theorem 5.153 (whose proof uses only that \(T\) commutes with an irreducible family of operators on a finite-dimensional complex space), gives \(T = \lambda\,\identity\) and therefore \(B = \lambda\,\kappa\).
∎Every \(2\)-cocycle \(c\) on \(L\mathfrak{g}\) that is homogeneous of degree zero — meaning \(c\left(x\otimes t^{m},y\otimes t^{n}\right) = 0\) whenever \(m+n \neq 0\) — is cohomologous to a multiple of \(\omega_{\kappa}\). Rests on Theorem A.426, Lemma A.427 and Proposition 14.77.
Derives Proposition A.428. Write \(B_{m}(x,y) = c\left(x\otimes t^{m},y\otimes t^{-m}\right)\). Antisymmetry of \(c\) gives
so \(B_{0}\) is an antisymmetric bilinear form on \(\mathfrak{g}\otimes 1 \cong \mathfrak{g}\).
Step 1: kill \(B_{0}\). The restriction of a cocycle to a subalgebra is a cocycle, so \(B_{0}\) is a \(2\)-cocycle on \(\mathfrak{g}\); by Proposition 14.77, \(H^{2}(\mathfrak{g},\C) = 0\), so \(B_{0} = \partial\beta\) for a linear \(\beta\) on \(\mathfrak{g}\). Extend \(\beta\) to \(b\) on \(L\mathfrak{g}\) by \(b(x\otimes t^{0}) = \beta(x)\) and \(b(x\otimes t^{n}) = 0\) for \(n\neq0\). Then \(\left(\partial b\right)\left(x\otimes t^{m},y\otimes t^{n}\right) = \delta_{m+n,0}\,\beta\left(\comm{x}{y}\right)\), which is homogeneous of degree zero, so \(c - \partial b\) is still homogeneous of degree zero and now has \(B_{0} = 0\). Replace \(c\) by it.
Step 2: each \(B_{m}\) is invariant. Apply Equation (14.121) to \(\left(z\otimes t^{0},\ x\otimes t^{m},\ y\otimes t^{-m}\right)\):
and \(B_{0} = 0\) while \(B_{-m}\left(\comm{y}{z},x\right) = -B_{m}\left(x,\comm{y}{z}\right) = B_{m}\left(x,\comm{z}{y}\right)\) by Equation (A.726). What is left is exactly Equation (A.725), so \(B_{m} = \lambda_{m}\,\kappa\) by Lemma A.427. Feeding that back into Equation (A.726) and using the symmetry of \(\kappa\),
Step 3: \(\lambda\) is additive. Apply Equation (14.121) to \(\left(x\otimes t^{a},\ y\otimes t^{b},\ z\otimes t^{e}\right)\) with \(a+b+e = 0\). The three terms pair \(t^{a+b} = t^{-e}\) with \(t^{e}\), and so on, giving
the three Killing factors being the common value \(\mathcal{K}\) computed in the proof of Theorem A.426. Some choice of \(x,y,z\) makes \(\mathcal{K} \neq 0\): \(\mathfrak{g} = \comm{\mathfrak{g}}{\mathfrak{g}}\) for a simple algebra, so a nonzero \(w\) with \(\kappa(w,z)\neq0\) — available by nondegeneracy — is a sum of brackets, one of which must pair nontrivially with \(z\). Hence \(\lambda_{p}+\lambda_{q}+\lambda_{r} = 0\) whenever \(p+q+r = 0\); putting \(r = -p-q\) and using Equation (A.727), \(\lambda_{p+q} = \lambda_{p}+\lambda_{q}\), so \(\lambda_{m} = m\,\lambda_{1}\).
Therefore \(c\left(x\otimes t^{m},y\otimes t^{n}\right) = \lambda_{1}\,m\,\delta_{m+n,0}\,\kappa(x,y)\), which is \(\lambda_{1}\omega_{\kappa}\) by Equation (A.722).
∎The hypothesis of Proposition A.428 is less restrictive than it looks, and it is worth saying how much less. Define, for a cocycle \(c\) and \(k \in \Z\), the component \(c_{k}\) by \(c_{k}\left(x\otimes t^{m},y\otimes t^{n}\right) = c\left(x\otimes t^{m},y\otimes t^{n}\right)\) when \(m+n = k\) and \(0\) otherwise. Each \(c_{k}\) is bilinear and antisymmetric, and each is separately a cocycle: every term of Equation (14.121) evaluated on \(\left(x\otimes t^{a},y\otimes t^{b},z\otimes t^{e}\right)\) pairs two factors whose exponents sum to \(a+b+e\), so the cocycle condition never mixes degrees. Since only one component is nonzero on any given pair of basis elements, \(c = \sum_{k}c_{k}\) makes sense with no convergence question. So the classification of cocycles reduces exactly to the classification of homogeneous ones, and Proposition A.428 settles the degree-zero part.
What is not proved here is that a homogeneous cocycle of nonzero degree is a coboundary — equivalently, that \(\dim H^{2}(L\mathfrak{g},\C) = 1\). That is a theorem of Garland (The arithmetic theory of loop groups, Publications mathématiques de l'IHÉS 52, 1980, 5–136), reproved in several places since. Its proof at degree \(k \neq 0\) turns on the nonvanishing of the third cohomology of a simple Lie algebra, which this treatise does not develop; the present appendix therefore establishes that the level is a central charge and the only one that survives the natural grading, not that it is the only one whatever.
The Schwinger term
The third case is different in kind, and the difference is not a defect of the exposition. A current algebra is a bracket relation between operator-valued distributions at equal times. The naive canonical evaluation of that bracket multiplies two distributions at one point, which is not an operation, so the “formal” current algebra is not a Lie algebra whose cohomology one can compute; the question is instead whether the central term that any legitimate definition produces can be made to vanish. Schwinger's answer is that it cannot. The argument below is his; it is stated in J. Schwinger, Field theory commutators, Physical Review Letters 3 (1959), 296–297, a work for which this treatise carries no bibliography entry, so the attribution is made here in prose and the reader is warned that it is uncited.
Throughout this subsection the setting is \(3+1\) dimensions, in accordance with rule 7: the argument is dimension-independent, but it is instantiated where the evidence is.
Assume a quantum theory on a Hilbert space \(\mathcal{H}\) with
-
a self-adjoint Hamiltonian \(\hat{H}\) bounded below, with a normalizable ground state \(\ket{0}\), \(\hat{H}\ket{0} = E_{0}\ket{0}\), and a complete orthonormal family of eigenstates \(\set{\ket{n}}\) with \(\hat{H}\ket{n} = E_{n}\ket{n}\) and \(E_{n} \ge E_{0}\);
-
a conserved current: Hermitian operator-valued distributions \(\hat{\jmath}^{0}(\vect{x},t)\) (charge density, of SI unit \(\mathrm{C}/\mathrm{m}^{3}\)) and \(\hat{\jmath}^{k}(\vect{x},t)\) (current density, \(\mathrm{C}/\mathrm{m}^{2}/\mathrm{s}\)), satisfying the operator continuity equation
\begin{equation}\tag{A.728} \pdv{\hat{\jmath}^{0}}{t} + \pp_{k}\hat{\jmath}^{k} = 0\ep \end{equation}
For a real smooth compactly supported \(f\) on \(\R^{3}\), dimensionless, write
a Hermitian operator of SI unit \(\mathrm{C}\). Rests on Definition 14.72.
Under Definition A.430,
with equality if and only if \(\hat{A}_{f}\ket{0}\) lies in the ground eigenspace of \(\hat{H}\). Rests on Definition A.430.
Derives Lemma A.431. Abbreviate \(\hat{A} = \hat{A}_{f}\). Expanding the nested commutator,
Taking the ground-state expectation and using \(\hat{H}\ket{0} = E_{0}\ket{0}\) on the last two terms,
Inserting the resolution of the identity \(\sum_{n}\ketbra{n}{n}\) between the operators and using \(\hat{A}^{\dagger} = \hat{A}\), which gives \(\bra{0}\hat{A}\ket{n} = \overline{\bra{n}\hat{A}\ket{0}}\),
which is Equation (A.730). Every summand is nonnegative because \(E_{n} \ge E_{0}\), so the sum vanishes only if \(\bra{n}\hat{A}\ket{0} = 0\) for every \(n\) with \(E_{n} > E_{0}\), that is only if \(\hat{A}\ket{0}\) has no component outside the ground eigenspace.
∎Under Definition A.430, suppose the equal-time commutator of the charge density with the current density has the vacuum expectation
with \(S\) a constant (the Schwinger term), of SI unit \(\mathrm{C}^{2}/\mathrm{m}/\mathrm{s}\). Then
for every test function \(f\). Hence \(S \ge 0\), and \(S = 0\) forces \(\hat{\jmath}^{0}(\vect{x},0)\ket{0}\) to lie in the ground eigenspace for every \(\vect{x}\) — that is, the charge density must have no vacuum fluctuations at all. In any theory in which it does, the equal-time commutator cannot vanish. Rests on Lemma A.431 and Definition A.430.
Derives Theorem A.432. The Heisenberg equation \(\ii\hbar\,\pp_{t}\hat{O} = \comm{\hat{O}}{\hat{H}}\) and the continuity equation Equation (A.728) give
so that, integrating by parts against the compactly supported \(f\),
Therefore
whose ground-state expectation, by Equation (A.731) and \(\int\dd^{3}y\,f(\vect{y})\,\pp_{k}^{(\vect{x})} \delta^{3}(\vect{x}-\vect{y}) = \left(\pp_{k}f\right)(\vect{x})\), is
the two imaginary units combining to \(-\ii\cdot\ii = 1\). Equating this with Equation (A.730) gives Equation (A.732). Since \(\int\abs{\nabla f}^{2} > 0\) for any nonconstant \(f\), the sign of \(S\) follows; and \(S = 0\) makes the right-hand side vanish for every \(f\), so by the equality clause of Lemma A.431 the vector \(\hat{A}_{f}\ket{0}\) lies in the ground eigenspace for every \(f\), which is the stated conclusion.
The dimensions check: \(\hat{A}_{f}\) carries \(\mathrm{C}\), so the right-hand side of Equation (A.732) carries \(\mathrm{J}\times\mathrm{C}^{2}\); on the left, \(\hbar S \int\abs{\nabla f}^{2}\dd^{3}x\) carries \(\mathrm{J}\,\mathrm{s}\times \mathrm{C}^{2}/\mathrm{m}/\mathrm{s}\times\mathrm{m}\), the same.
∎Two features of Equation (A.731) put it in the frame of Section 14.4. First, the right-hand side is a \(c\)-number: it commutes with every operator of the theory, so as an element of the algebra generated by the smeared densities it is central, exactly as the \(Z\) of Equation (14.120). Second, it cannot be removed by redefining the generators. The would-be redefinition is a shift of each smeared density by a constant, \(\hat{A}_{f}\mapsto\hat{A}_{f}+b(f)\), which changes the commutator by \(b\) of a bracket — the coboundary of Equation (14.122) — and the naive algebra has \(\comm{\hat{\jmath}^{0}}{\hat{\jmath}^{k}} = 0\), so that bracket is zero and no shift produces anything. The situation is formally identical to the Galilei mass of Example 14.80: a central term sitting on a pair of generators whose classical bracket vanishes, and therefore beyond the reach of any coboundary (The Second Cohomology of the Galilei Algebra).
What Theorem A.432 adds, and what has no analogue in the two finite computations above, is that the term is forced. There the question was which cocycles exist; here the cocycle that appears to be zero is proved to be nonzero by an argument that never computes it — only positivity of the energy above the ground state and conservation of the current are used.
Three things are imported in this section, and the three are of different weight.
(i) Whitehead's second lemma, used once, in Step 1 of Proposition A.428: \(H^{2}(\mathfrak{g},F) = 0\) for a finite-dimensional semisimple \(\mathfrak{g}\) in characteristic zero. It is Proposition 14.77 of the chapter, whose own derivation is recorded there and not repeated here. It is used only to remove the degree-zero-in-\(t\) piece \(B_{0}\), and nothing else in this appendix depends on it.
(ii) Garland's theorem, that \(\dim H^{2}(L\mathfrak{g},\C) = 1\) for simple \(\mathfrak{g}\), which would upgrade Proposition A.428 from the degree-zero statement to the full one. It is not proved here and it is not used: Theorem A.426 constructs the cocycle and proves it nontrivial without it, and Remark A.429 says precisely which case remains. The Virasoro uniqueness, by contrast, is proved in full — Theorem A.422 imports nothing.
(iii) The framework in which Theorem A.432 is stated. This is the substantial one, and it is an assumption rather than a theorem. It is assumed that the charge and current densities are operator-valued distributions on a common dense domain, smearable as in Equation (A.729), and that the equal-time commutator of two of them is supported on the diagonal \(\vect{x} = \vect{y}\), so that it is a finite sum of derivatives of \(\delta^{3}(\vect{x}-\vect{y}) \) with operator coefficients; the leading such term with a \(c\)-number coefficient is what Equation (A.731) isolates. That framework is the Wightman axiomatization [Streater:1964], which this treatise uses but does not construct. The spectral hypothesis of Definition A.430 — a self-adjoint Hamiltonian bounded below with a complete family of eigenstates — is likewise the standard spectral theorem for a self-adjoint operator [Reed:1972], applied in the form in which the sum over \(n\) is legitimate.
What is proved above, given that framework, is the inequality Equation (A.732) and its two consequences: \(S \ge 0\), and \(S = 0\) only in the degenerate case where the charge density does not fluctuate in the ground state. What is not proved, and is not attempted, is the evaluation of \(S\) in any particular theory. That evaluation requires a regularization — a point splitting of the product of two field operators, followed by the removal of the splitting — and it is a computation in quantum field theory, not in Lie algebra cohomology. This appendix therefore establishes the statement made in Example 14.81 in the exact form “a formally central extension survives regularization”: the extension is central by Remark A.433, and it survives because Theorem A.432 forbids the value zero that the unregularized computation returns.
Finally, the evidential status of all three algebras is settled in Remark 14.82 and is not reopened here. Writing down a consistent central extension is a mathematical act; that remark records what is and is not measured, in particular that the Schwinger term reaches experiment only through the anomalies [Adler:1969] [Bell:1969] and their measured consequences [Larin:2020], and that this treatise records no measurement of a Virasoro central charge.
Theorems A.422, A.426 and A.432 discharge the three obligations recorded after Example 14.81 of Lie Groups, Lie Algebras, and Fibre Bundles. Read together with Example 14.79 — where every antisymmetric form on an abelian algebra is a cocycle and none is a coboundary — and with The Second Cohomology of the Galilei Algebra, they fill in Remark 14.78's list of the four places a central charge can live: abelian algebras, inhomogeneous algebras and their contractions, and infinite-dimensional algebras. The last is the case treated here, and the three instances differ instructively. The Witt algebra has a one-dimensional \(H^{2}\) proved outright; the loop algebra has a distinguished class built from the Killing form and the residue, whose uniqueness is settled here only within the natural grading; and the current algebra has an extension that no cohomological computation produces, because the bracket it extends does not exist until the theory is regularized.
The Strengthened Jacobi Condition and Conjugate Points
This appendix proves the equivalence asserted in Remark 16.57 of Calculus of Variations: for a regular problem, the hypothesis of Theorem 16.56 — the existence of a solution of the Jacobi accessory equation Equation (16.49) with no zero on the closed interval \([a,b]\) — holds if and only if no point of \(\left(a,b\right]\) is conjugate to \(a\) in the sense of Definition 16.53. That equivalence is what joins Jacobi's necessary condition (Theorem 16.55) to his sufficient one, and without it the two theorems test different hypotheses.
Three ingredients are needed and all three are proved here: global existence and uniqueness for the accessory equation, which is Corollary 9.9 of Ordinary Differential Equations and Sturm–Liouville Theory once the equation is written as a first-order system; the continuous dependence of a solution on the point at which its initial data are posed, which is proved from the integral equation Equation (9.7) underlying the Picard–Lindelöf theorem Theorem 9.8 together with a Grönwall estimate; and the Sturm separation theorem, which is proved from the Lagrange identity Lemma 9.56. The notation is that of Section 16.4.2 throughout: \(P\) and \(Q\) are the coefficients Equation (16.46) evaluated along the extremal, the problem is regular, meaning \(P>0\) on \([a,b]\), and the accessory equation is
Only the continuity of \(P\) and \(Q\) and the positivity of \(P\) are used; in particular \(P\) is never differentiated, which is why the results below hold for every \(C^{3}\) integrand and \(C^{2}\) extremal without a further regularity hypothesis.
Statement
Let \(P,Q\in C^{0}[a,b]\) with \(P>0\) on \([a,b]\), and let \(u_{a}\) denote the solution of Equation (A.734) determined by
The following three statements are equivalent.
-
There exists a solution \(w\) of Equation (A.734) with \(w(x)\neq0\) for every \(x\in[a,b]\) (the strengthened Jacobi condition).
-
No point of \(\left(a,b\right]\) is conjugate to \(a\) (Definition 16.53).
-
\(u_{a}>0\) on \(\left(a,b\right]\).
Rests on Definition 16.53 and Equation (16.49).
Definition 16.53 normalises the Jacobi solution by \(u'(a)=1\), whereas Equation (A.735) normalises the quantity \(P\,\dd u/\dd x\), which is the natural momentum variable of Equation (A.734). The two differ by the positive factor \(P(a)\), and a nonzero constant multiple of a solution has exactly the same zeros, so the set of points conjugate to \(a\) is the same under either normalisation. The momentum normalisation is used below because \(P\) is not assumed differentiable and \(P\,\dd u/\dd x\), unlike \(\dd u/\dd x\) alone, is a component of the first-order system Equation (A.736).
The proof occupies the rest of the section: the system form and the elementary structure of the zeros (The accessory equation as a first-order system), continuous dependence on the initial point (Continuous dependence on the initial point), the Sturm separation theorem (The Sturm separation theorem), and the equivalence itself (Proof of the equivalence), after which The displaced initial point recovers the exact phrasing used in the chapter, in which the initial point is displaced to the left of \(a\).
The accessory equation as a first-order system
For a solution \(u\) of Equation (A.734) put \(v=P\,\dd u/\dd x\) and \(\vect{Y}=\left(u,v\right)\). Then \(\vect{Y}\) solves the linear system
whose coefficient matrix is continuous on \([a,b]\) because \(P\) and \(Q\) are continuous and \(P\) does not vanish. Conversely, if \(\vect{Y}=(u,v)\) solves Equation (A.736) then \(u\) is of class \(C^{1}\) with \(P\,\dd u/\dd x=v\) of class \(C^{1}\), and \(\dd\left(P\,\dd u/\dd x\right)/\dd x=Qu\), which is Equation (A.734). Rests on Equations (9.3) and (16.49).
For every \(s\in[a,b]\) and every \(\vect{Y}_{0}\in\R^{2}\) the system Equation (A.736) has exactly one solution on the whole of \([a,b]\) with \(\vect{Y}(s)=\vect{Y}_{0}\). Consequently, for a solution \(u\) of Equation (A.734) that does not vanish identically:
-
at every zero \(x_{0}\) of \(u\) one has \(\left(P\,\dd u/\dd x\right)(x_{0})\neq0\), so the zero is simple;
-
the zeros of \(u\) in \([a,b]\) are isolated, and there are finitely many of them.
Rests on Definition A.438 and Corollary 9.9.
Derives Lemma A.439. The system Equation (A.736) is linear with a continuous coefficient matrix on the compact interval \([a,b]\), so Corollary 9.9 applies verbatim with \(\vect{g}=\vect{0}\): there is exactly one solution on all of \([a,b]\) for each choice of \(s\) and \(\vect{Y}_{0}\).
(1) If \(u(x_{0})=0\) and \(\left(P\,\dd u/\dd x\right)(x_{0})=0\) then \(\vect{Y}(x_{0})=\vect{0}\); but \(\vect{Y}\equiv\vect{0}\) is a solution with those data, so uniqueness forces \(\vect{Y}\equiv\vect{0}\) and hence \(u\equiv0\), which was excluded.
(2) Let \(x_{0}\) be a zero. By (1) and \(P>0\) we have \(\left(\dd u/\dd x\right)(x_{0})\neq0\), so by the definition of the derivative there is a punctured neighbourhood of \(x_{0}\) in which \(u(x)/(x-x_{0})\) keeps the sign of \(\left(\dd u/\dd x\right)(x_{0})\) and in particular \(u(x)\neq0\): the zero is isolated. If there were infinitely many zeros in \([a,b]\), Bolzano–Weierstrass (Theorem 7.7) would give a sequence of distinct zeros converging to some \(x_{\ast}\in[a,b]\); continuity of \(u\) makes \(x_{\ast}\) a zero as well, and it is not isolated — a contradiction.
∎Continuous dependence on the initial point
The estimate below is the one place where a quantitative statement is needed rather than a qualitative one, and it is obtained from the integral form Equation (9.7) of the initial value problem, which is where the proof of Theorem 9.8 begins. Throughout, \(\norm{\cdot}\) is any fixed norm on \(\R^{2}\) and \(\norm{A(x)}\) the induced operator norm, so that \(\norm{A(x)\vect{Z}}\le\norm{A(x)}\,\norm{\vect{Z}}\); put
which is finite because \(A\) is continuous on a compact interval (Theorem 7.24).
Let \(\varphi\in C^{0}[a,b]\) be nonnegative, let \(x_{0}\in[a,b]\), and let \(C\ge0\) and \(L\ge0\) satisfy
Then \(\varphi(x)\le C\,\ee^{L\abs{x-x_{0}}}\) on \([a,b]\). Rests on Theorems 7.35 and 7.42.
Derives Lemma A.440. Take first \(x\ge x_{0}\) and set \(\Phi(x)=C+L\int_{x_{0}}^{x}\varphi(t)\,\dd t\), which is of class \(C^{1}\) with \(\Phi'=L\varphi\) by the fundamental theorem of calculus (Theorem 7.42). Hypothesis Equation (A.738) reads \(\varphi\le\Phi\) on \([x_{0},b]\), so \(\Phi'=L\varphi\le L\Phi\) there, and therefore
By the mean value theorem (Theorem 7.35), a function whose derivative is nowhere positive on an interval satisfies \(g(x)-g\left(x_{0}\right)=g'(\xi)\left(x-x_{0}\right)\le0\) for some interior \(\xi\) whenever \(x>x_{0}\); applied to \(g(x)=\Phi(x)\,\ee^{-L\left(x-x_{0}\right)}\) this gives \(\Phi(x)\,\ee^{-L\left(x-x_{0}\right)}\le\Phi(x_{0})=C\), that is \(\varphi(x)\le\Phi(x)\le C\,\ee^{L\left(x-x_{0}\right)}\).
For \(x\le x_{0}\) apply what has just been proved to \(\tilde\varphi(t)=\varphi\left(x_{0}-t\right)\) on \(\left[0,x_{0}-a\right]\), which satisfies the same hypothesis with the same constants because the substitution \(t\longmapsto x_{0}-t\) turns \(\left|\int_{x_{0}}^{x}\varphi\right|\) into \(\left|\int_{0}^{x_{0}-x}\tilde\varphi\right|\).
∎For \(s\in[a,b]\) and \(\vect{Y}_{0}\in\R^{2}\) write \(\vect{Y}\left(\cdot\,;s,\vect{Y}_{0}\right)\) for the solution of Equation (A.736) with \(\vect{Y}(s)=\vect{Y}_{0}\). Then, with \(L\) as in Equation (A.737), for all \(s,\tilde s\in[a,b]\) and all \(\vect{Y}_{0},\tilde{\vect{Y}}_{0}\),
In particular the map \(\left(x,s,\vect{Y}_{0}\right)\longmapsto \vect{Y}\left(x;s,\vect{Y}_{0}\right)\) is continuous on \([a,b]\times[a,b]\times\R^{2}\). Rests on Lemma A.440, Lemma A.439 and Equation (9.7).
Derives Theorem A.441. Write \(\vect{Y}=\vect{Y}\left(\cdot\,;s,\vect{Y}_{0}\right)\) and \(\vect{Z}=\vect{Y}\left(\cdot\,;\tilde s,\tilde{\vect{Y}}_{0}\right)\). Both are continuous and, by the equivalence of the initial value problem with its integral form Equation (9.7) — which is the first step of the proof of Theorem 9.8 and uses only the fundamental theorem of calculus —
A bound for \(\vect{Z}\) alone. Taking norms in the second identity of Equation (A.740),
so Lemma A.440 with \(\varphi=\norm{\vect{Z}}\), \(C=\norm{\tilde{\vect{Y}}_{0}}\) and \(x_{0}=\tilde s\) gives
The difference. Subtracting the two identities of Equation (A.740) and splitting the second integral at \(s\),
The last term is bounded in norm by \(L\,\abs{s-\tilde s}\,\max_{[a,b]}\norm{\vect{Z}}\), which Equation (A.741) bounds by \(L\,\ee^{L\left(b-a\right)}\norm{\tilde{\vect{Y}}_{0}}\, \abs{s-\tilde s}\). Hence, writing \(\varphi=\norm{\vect{Y}-\vect{Z}}\) and \(C=\norm{\vect{Y}_{0}-\tilde{\vect{Y}}_{0}} +L\,\ee^{L\left(b-a\right)}\norm{\tilde{\vect{Y}}_{0}}\, \abs{s-\tilde s}\),
and Lemma A.440 gives Equation (A.739).
For the final assertion, fix \(\left(x_{1},s_{1},\vect{Y}_{1}\right)\). The estimate Equation (A.739) shows that \(\vect{Y}\left(\cdot\,;s,\vect{Y}_{0}\right)\) converges uniformly in \(x\) to \(\vect{Y}\left(\cdot\,;s_{1},\vect{Y}_{1}\right)\) as \(\left(s,\vect{Y}_{0}\right)\longrightarrow \left(s_{1},\vect{Y}_{1}\right)\), and each solution is continuous in \(x\); a uniform limit of continuous functions is continuous, and the triangle inequality
makes both contributions small.
∎The Sturm separation theorem
Let \(u_{1},u_{2}\) solve Equation (A.734) on \([a,b]\) and put
Then \(W\) is constant on \([a,b]\), and \(W=0\) if and only if \(u_{1}\) and \(u_{2}\) are linearly dependent. Rests on Lemmas 9.56 and A.439.
Derives Lemma A.442. Equation (A.734) is the Sturm–Liouville equation \(L[u]=0\) of Lemma 9.56 with the data \(p=P\) and \(q=-Q\). The Lagrange identity Equation (9.56) therefore gives
so \(W\) is constant. In the notation of Definition A.438, \(W\) is the determinant of the matrix with columns \(\vect{Y}_{1}=\left(u_{1},v_{1}\right)\) and \(\vect{Y}_{2}=\left(u_{2},v_{2}\right)\), since \(u_{1}v_{2}-u_{2}v_{1}=P\left(u_{1}u_{2}'-u_{2}u_{1}'\right)\). If \(W=0\), the two vectors are parallel at every point; picking a point \(x_{0}\) at which \(\vect{Y}_{1}\left(x_{0}\right)\neq\vect{0}\) — one exists unless \(u_{1}\equiv0\), in which case the dependence is trivial — gives \(\vect{Y}_{2}\left(x_{0}\right) =\kappa\,\vect{Y}_{1}\left(x_{0}\right)\) for a scalar \(\kappa\), and uniqueness in Lemma A.439 then forces \(\vect{Y}_{2}=\kappa\vect{Y}_{1}\) identically, hence \(u_{2}=\kappa u_{1}\). Conversely \(u_{2}=\kappa u_{1}\) makes Equation (A.742) vanish outright.
∎Let \(u_{1},u_{2}\) be linearly independent solutions of Equation (A.734) on \([a,b]\). Then between any two consecutive zeros of \(u_{1}\) there lies exactly one zero of \(u_{2}\), and \(u_{1},u_{2}\) have no zero in common. In particular the zeros of the two solutions strictly interlace. Rests on Lemma A.442 and Theorem 7.23.
Derives Theorem A.443. By Lemma A.442 the constant \(W=P\left(u_{1}u_{2}'-u_{2}u_{1}'\right)\) is nonzero. Neither solution vanishes identically, so Lemma A.439 applies to both.
No common zero. If \(u_{1}\left(x_{0}\right) =u_{2}\left(x_{0}\right)=0\) then both terms of Equation (A.742) vanish at \(x_{0}\) and \(W=0\), a contradiction.
At least one zero in between. Let \(x_{1}<x_{2}\) be consecutive zeros of \(u_{1}\) in \([a,b]\), so that \(u_{1}\neq0\) on \(\left(x_{1},x_{2}\right)\). Evaluating Equation (A.742) at the two ends, where \(u_{1}\) vanishes,
Both derivatives are nonzero by Lemma A.439(1), and they have opposite signs: \(u_{1}\) keeps one sign, say positive, throughout \(\left(x_{1},x_{2}\right)\), so \(\left(\dd u_{1}/\dd x\right)\left(x_{1}\right) =\lim_{x\to x_{1}^{+}}u_{1}(x)/\left(x-x_{1}\right)\ge0\) and hence is \(>0\), while \(\left(\dd u_{1}/\dd x\right)\left(x_{2}\right) =\lim_{x\to x_{2}^{-}}u_{1}(x)/\left(x-x_{2}\right)\le0\) and hence is \(<0\); if \(u_{1}<0\) in between, both signs reverse. Since \(W\neq0\) is the same number at both ends of Equation (A.743) and \(P>0\), the values \(u_{2}\left(x_{1}\right)\) and \(u_{2}\left(x_{2}\right)\) are both nonzero and of opposite sign. The intermediate value theorem (Theorem 7.23) gives a zero of \(u_{2}\) in \(\left(x_{1},x_{2}\right)\).
At most one. Suppose \(u_{2}\) had two zeros \(y_{1}<y_{2}\) in \(\left(x_{1},x_{2}\right)\); choosing them consecutive among the zeros of \(u_{2}\) — possible because by Lemma A.439(2) there are finitely many — and applying the previous paragraph with the roles of \(u_{1}\) and \(u_{2}\) exchanged produces a zero of \(u_{1}\) in \(\left(y_{1},y_{2}\right)\subset\left(x_{1},x_{2}\right)\), contradicting the assumption that \(x_{1}\) and \(x_{2}\) are consecutive zeros of \(u_{1}\).
∎Proof of the equivalence
Proof of Theorem A.436. Derives Theorem A.436. The three implications are proved in the cycle \((1)\Rightarrow(2)\Rightarrow(3)\Rightarrow(1)\).
\((1)\Rightarrow(2)\). Let \(w\) be a solution with no zero on \([a,b]\) and suppose, for contradiction, that some \(c\in\left(a,b\right]\) is conjugate to \(a\), i.e.\ \(u_{a}(c)=0\). The solution \(u_{a}\) does not vanish identically, since \(\left(P\,\dd u_{a}/\dd x\right)(a)=1\neq0\). Moreover \(u_{a}\) and \(w\) are linearly independent: a multiple of \(w\) cannot vanish at \(a\) as \(u_{a}\) does, unless the multiple is zero. By Lemma A.439(2) the zeros of \(u_{a}\) in \([a,b]\) are finite in number, so among those lying in \(\left(a,c\right]\) there is a smallest one, \(c_{1}\); then \(a\) and \(c_{1}\) are consecutive zeros of \(u_{a}\). The Sturm separation theorem Theorem A.443 places a zero of \(w\) in \(\left(a,c_{1}\right)\subset[a,b]\), contradicting the hypothesis on \(w\).
\((2)\Rightarrow(3)\). Since \(\left(P\,\dd u_{a}/\dd x\right)(a)=1\) and \(P(a)>0\), we have \(\left(\dd u_{a}/\dd x\right)(a)=1/P(a)>0\), and with \(u_{a}(a)=0\) this makes \(u_{a}(x)>0\) for \(x>a\) close enough to \(a\). By (2), \(u_{a}\) has no zero in \(\left(a,b\right]\); were \(u_{a}\) negative at some point of \(\left(a,b\right]\), the intermediate value theorem (Theorem 7.23) would produce a zero between that point and one where \(u_{a}>0\). Hence \(u_{a}>0\) throughout \(\left(a,b\right]\).
\((3)\Rightarrow(1)\). Let \(z\) be the solution of Equation (A.734) with
which exists by Lemma A.439, and consider the one-parameter family of solutions
each of which solves Equation (A.734) by linearity. We show that \(w_{\delta}>0\) on \([a,b]\) for all small \(\delta>0\).
Write \(v_{a}=P\,\dd u_{a}/\dd x\), a continuous function with \(v_{a}(a)=1\), and recall \(z(a)=1\) with \(z\) continuous. Choose \(\eta>0\) so small that
On that subinterval \(\dd u_{a}/\dd x=v_{a}/P>0\), so \(u_{a}(x)=\int_{a}^{x}\left(v_{a}/P\right)\dd t\ge0\) by Theorem 7.43, and therefore
If \(a+\eta\ge b\) this already proves the claim for every \(\delta>0\). Otherwise, on the compact interval \(\left[a+\eta,b\right]\) the continuous function \(u_{a}\) is strictly positive by (3), so by the extreme value theorem (Theorem 7.24) it has a minimum \(m>0\) there; let \(M=\max_{[a,b]}\abs{z}\), also finite by Theorem 7.24. Choosing any \(\delta\) with \(0<\delta<m/\left(M+1\right)\) gives, for \(x\in\left[a+\eta,b\right]\),
So \(w=w_{\delta}\) has no zero on \([a,b]\), which is (1).
∎With Theorem A.436 the two halves of the second-order theory close on each other. For a regular problem, Theorem 16.55 says that a weak minimum admits no conjugate point in the open interval \((a,b)\), and Theorem 16.56 says that a solution without zeros on the closed interval makes \(\delta^{2}J\) positive definite; by statement (2) of the theorem just proved, the hypothesis of the second is exactly the absence of conjugate points in \(\left(a,b\right]\). The one remaining gap between them is the single point \(x=b\): an extremal whose first conjugate point is the right endpoint itself satisfies the necessary condition and fails the sufficient one, and the second variation is then positive semidefinite with the Jacobi solution \(u_{a}\) in its null space — which is precisely the computation carried out in the proof of Theorem 16.55, where \(\delta^{2}J\left[y;\eta\right]=0\) for the broken variation built from \(u_{a}\). The threshold case is not a defect of the theory: it is the kinetic focus, and Remark 16.74 exhibits it as the exact instant \(T=\pi\sqrt{m/k}\) at which the harmonic oscillator's action stops being a minimum.
The displaced initial point
Remark 16.57 states the equivalence in the form in which it is usually met: the Jacobi solution vanishing slightly to the left of \(a\) has no zero in \([a,b]\) exactly when the solution vanishing at \(a\) has none in \(\left(a,b\right]\). That form presupposes that the coefficients are defined on a slightly larger interval, which happens whenever the extremal itself extends; under that hypothesis it is a corollary of Theorems A.436 and A.441, and it is the form in which continuous dependence on the initial point does real work.
Suppose \(P\) and \(Q\) extend continuously to \(\left[a-\varepsilon_{0},b\right]\) for some \(\varepsilon_{0}>0\), with \(P>0\) there, and for \(s\in\left[a-\varepsilon_{0},b\right)\) let \(u_{s}\) be the solution of Equation (A.734) with \(u_{s}(s)=0\) and \(\left(P\,\dd u_{s}/\dd x\right)(s)=1\). Then the strengthened Jacobi condition holds on \([a,b]\) if and only if there is an \(\varepsilon\in\left(0,\varepsilon_{0}\right)\) such that \(u_{a-\varepsilon}\) has no zero in \([a,b]\) — and in that case no zero in \(\left[a,b\right]\) for every smaller \(\varepsilon\) as well. Rests on Theorems A.436 and A.441.
Derives Corollary A.445. Sufficiency is immediate: \(u_{a-\varepsilon}\) restricted to \([a,b]\) is a solution of Equation (A.734) without zeros there, which is statement (1) of Theorem A.436.
Necessity. Assume the strengthened Jacobi condition, hence statement (3): \(u_{a}>0\) on \(\left(a,b\right]\). Work on the extended interval \(I=\left[a-\varepsilon_{0},b\right]\), on which Theorem A.441 holds with the constant \(L\) of Equation (A.737) computed over \(I\). Write \(v_{s}=P\,\dd u_{s}/\dd x\), so that the initial vector is \(\vect{Y}_{0}=\left(0,1\right)\) for every \(s\) and only the initial point varies; Equation (A.739) then reads
where the norm on \(\R^{2}\) has been taken to be the sum of the moduli of the components and \(\norm{\vect{Y}_{0}}=1\).
Since \(v_{a}\) is continuous on \(I\) with \(v_{a}(a)=1\), fix \(\eta>0\) with \(\eta<\varepsilon_{0}\) and
If \(a+\eta<b\), the function \(u_{a}\) is continuous and strictly positive on the compact interval \(\left[a+\eta,b\right]\) and so has a minimum \(m>0\) there (Theorem 7.24); if \(a+\eta\ge b\) put \(m=+\infty\) and ignore the second condition below. Now choose \(\varepsilon>0\) subject to
By Equation (A.746) the solution \(u=u_{a-\varepsilon}\) and its momentum \(v=v_{a-\varepsilon}\) then satisfy
On \(\left[a-\varepsilon,a+\eta\right]\) the first inequality gives \(\dd u/\dd x=v/P>0\), so \(u\) is strictly increasing there and, since \(u\left(a-\varepsilon\right)=0\), it is strictly positive on \(\left(a-\varepsilon,a+\eta\right]\) — in particular on \(\left[a,a+\eta\right]\), because \(a>a-\varepsilon\). Together with the second inequality, \(u>0\) on all of \([a,b]\), so \(u_{a-\varepsilon}\) has no zero there. The argument used only the three smallness conditions on \(\varepsilon\), each of which persists when \(\varepsilon\) is decreased, which proves the last clause.
∎Corollary A.445 is the analytic content of the picture in Remark 16.54. A solution of Equation (A.734) is the derivative \(\pp y/\pp\alpha\) of a one-parameter family of extremals through a common initial point, so \(u_{s}\) describes, to first order, the spreading of the pencil of extremals issuing from the point \(s\). The corollary says that the pencil issuing from \(a\) stays spread out over \([a,b]\) if and only if a pencil issuing from a slightly earlier point does, and Theorem A.441 is what makes “slightly earlier” a legitimate perturbation: the focusing distance depends continuously on the point of emission. On the unit sphere, where the extremals are great circles and every pencil refocuses at the antipode, the corollary reproduces the familiar statement that an arc is a minimising geodesic exactly while it is shorter than half a great circle.
Theorem A.436 discharges the equivalence asserted in Remark 16.57 of Section 16.4.2, and Corollary A.445 supplies the displaced-point phrasing used there. The two conditions the equivalence links are Jacobi's necessary condition (Theorem 16.55) and the sufficiency of the second variation (Theorem 16.56); with the equivalence in hand, a regular extremal with no conjugate point in \(\left(a,b\right]\) has a positive definite second variation, which is the form in which the result is used in the field-theoretic sufficiency argument of Section 16.4.4 and in the discussion of the kinetic focus of Hamilton's principle in Section 16.7.1.
Tonelli's Existence Theorem for an Integrand Depending on the Function
This appendix proves Theorem 16.73 of Calculus of Variations: an integrand \(f\left(x,y,q\right)\) that is convex in the slope \(q\) and coercive in it makes the functional Equation (16.11) attain its minimum in the Sobolev class with prescribed endpoint values. The abstract half of the argument is already in the chapter — the direct method Theorem 16.70 and the convexity lemma Lemma 16.72 — and what that lemma leaves owed is the passage from an integrand \(f\left(x,q\right)\) to one that also depends on \(y\). That passage is what is carried out here, and it turns on one fact about the one-dimensional Sobolev space: a bounded set of \(W^{1,r}(a,b)\) is precompact in the continuous functions with the uniform norm. That compact embedding is proved below in full (Theorem A.455) from the fundamental theorem of calculus, Hölder's inequality and the Arzelà–Ascoli theorem, the last of which is not available anywhere else in this treatise and is therefore also proved here (Lemma A.454).
Uniform convergence of the competitors is exactly the leverage the \(y\)-dependence needs: it makes the values \(f\left(x,u_{n}(x),q\right)\) converge pointwise to \(f\left(x,u(x),q\right)\) and confines them to a compact range of \(y\), on which the growth hypothesis on \(\pp f/\pp q\) can be applied uniformly. With that in hand the classical Scorza–Dragoni argument — which is the route taken when no such growth hypothesis is assumed — is not needed; Remark A.458 states precisely what it would supply and why it is avoided here, and Remark A.459 lists the analytic inputs that are quoted rather than proved, all of them from the Lebesgue theory and the duality of \(L^{r}\) which this treatise declares as imports in Remark 12.1.
Statement
Throughout, \(1<r<\infty\) and \(r'=r/\left(r-1\right)\) is the conjugate exponent, so that \(1/r+1/r'=1\); \(a<b\) are real and \(y_{a},y_{b}\) are the prescribed endpoint values of Equation (16.10).
Let \(f:[a,b]\times\R\times\R\longrightarrow\R\) satisfy
-
[(H1)] \(f\) is continuous, and continuously differentiable with respect to its third argument, with derivative written \(f_{q}=\pp f/\pp q\);
-
[(H2)] \(q\longmapsto f\left(x,y,q\right)\) is convex for every \(\left(x,y\right)\);
-
[(H3)] \(f\left(x,y,q\right)\ge\alpha\abs{q}^{r}-\beta\) for constants \(\alpha>0\) and \(\beta\in\R\) and all \(\left(x,y,q\right)\);
-
[(H4)] for every \(R>0\) there is a constant \(C_{R}\) with \(\abs{f_{q}\left(x,y,q\right)} \le C_{R}\left(1+\abs{q}^{r-1}\right)\) whenever \(x\in[a,b]\), \(\abs{y}\le R\) and \(q\in\R\).
Then the functional Equation (16.11) attains its minimum on
that is, there is \(u_{\ast}\in\mathcal{A}_{r}\) with \(J\left[u_{\ast}\right]=\inf_{\mathcal{A}_{r}}J\), and the infimum is finite. Rests on Theorem 16.70, Lemma 16.72 and Equation (16.11).
(H1)–(H4) are the hypotheses of Theorem 16.73 as the chapter states it. (H2) and (H3) — convexity and coercivity in the slope — are Tonelli's own. (H1) and (H4) are the hypotheses under which the chapter's convexity lemma Lemma 16.72 was already stated in the \(y\)-free case: there \(f\) is of class \(C^{1}\) in \(q\) and \(\abs{\pp f/\pp q\left(x,q\right)}\le C\left(1+\abs{q}^{r-1}\right)\). (H4) is that same bound made locally uniform in \(y\), which is the least that can be asked once \(y\) enters, and it is automatic for every integrand met in this book — for a mechanical Lagrangian \(\Lag=\tfrac12 m_{ij}(q)\dot q^{i}\dot q^{j}-V(q)\) the derivative in the velocities is linear in them, so (H4) holds with \(r=2\) and \(C_{R}\) the maximum of the mass matrix over \(\abs{y}\le R\). The theorem as Tonelli stated it [Tonelli:1915] assumes only continuity and convexity; Remark A.458 says what closing that last gap costs.
The Sobolev space in one dimension
For \(1<r<\infty\), \(W^{1,r}(a,b)\) is the set of functions \(u:[a,b]\longrightarrow\R\) for which there exists \(v\in L^{r}(a,b)\) with
normed by \(\norm{u}_{W^{1,r}}=\norm{u}_{L^{r}}+\norm{v}_{L^{r}}\). The function \(v\) is unique as an element of \(L^{r}\) and is written \(u'\); every \(u\in W^{1,r}(a,b)\) is continuous on \([a,b]\), by Equation (A.748) and the absolute continuity of the Lebesgue integral. Rests on Definition 10.85 and Theorem 7.42.
Definition 10.85 introduces \(H^{1}=W^{1,2}\) through distributional derivatives, and in one dimension the two definitions agree: if \(u\in L^{r}(a,b)\) has a distributional derivative \(v\in L^{r}\), then \(w(x)=\int_{a}^{x}v\) has the same distributional derivative, so \(u-w\) has distributional derivative zero and is therefore almost everywhere equal to a constant — which is the du Bois-Reymond lemma Lemma 16.19 in its integrable form, the one step of this identification that needs Lebesgue theory rather than the continuous argument given in the chapter. Nothing below uses the distributional definition, so Definition A.450 is taken as the definition and the identification is recorded only to place the space where the reader expects it. In particular each element of \(W^{1,r}(a,b)\) is here a genuine continuous function, not an equivalence class, so the boundary conditions in Equation (A.747) have their literal meaning.
Let \(1<r<\infty\) and \(r'=r/\left(r-1\right)\). For all \(s,t\ge0\),
and consequently, for measurable \(g,h\) on \((a,b)\),
Rests on Theorem 7.38 and Definition A.450.
Derives Lemma A.452. Young. For \(s=0\) or \(t=0\) the inequality is trivial, so let \(s,t>0\) and put \(\lambda=1/r\in(0,1)\), so \(1-\lambda=1/r'\). The function \(\Phi=-\log\) is of class \(C^{2}\) on \((0,\infty)\) with \(\Phi''(\tau)=\tau^{-2}>0\), hence convex there: writing \(c=\lambda\,\sigma+\left(1-\lambda\right)\tau\) for \(\sigma,\tau>0\), Taylor's theorem with Lagrange remainder (Theorem 7.38) about \(c\) gives \(\Phi(\sigma)=\Phi(c)+\Phi'(c)\left(\sigma-c\right) +\tfrac12\Phi''(\xi)\left(\sigma-c\right)^{2} \ge\Phi(c)+\Phi'(c)\left(\sigma-c\right)\) and the same with \(\tau\) in place of \(\sigma\); multiplying the first by \(\lambda\), the second by \(1-\lambda\) and adding annihilates the linear terms, because \(\lambda\left(\sigma-c\right) +\left(1-\lambda\right)\left(\tau-c\right)=0\), and leaves \(\lambda\,\Phi(\sigma)+\left(1-\lambda\right)\Phi(\tau)\ge\Phi(c)\). Taking \(\sigma=s^{r}\) and \(\tau=t^{r'}\) and undoing the minus sign,
and \(\log\) is increasing, which is Equation (A.749).
Hölder. If \(\norm{g}_{L^{r'}}=0\) or \(\norm{h}_{L^{r}}=0\) then \(gh=0\) almost everywhere and there is nothing to prove; if either norm is infinite the right side is \(+\infty\). Otherwise replace \(g\) by \(g/\norm{g}_{L^{r'}}\) and \(h\) by \(h/\norm{h}_{L^{r}}\), which reduces the claim to the case \(\norm{g}_{L^{r'}}=\norm{h}_{L^{r}}=1\). Apply Equation (A.749) pointwise with \(s=\abs{h(x)}\), \(t=\abs{g(x)}\) and integrate:
Let \(u\in W^{1,r}(a,b)\). Then for all \(x,x'\in[a,b]\),
and consequently
Rests on Definition A.450 and Lemma A.452.
Derives Lemma A.453. By Equation (A.748), for \(x'<x\), \(u(x)-u\left(x'\right)=\int_{x'}^{x}u'(t)\,\dd t\). Apply Equation (A.750) on the interval \(\left(x',x\right)\) with \(h=u'\) and \(g\) the constant function \(1\), whose \(L^{r'}\) norm over that interval is \(\left(x-x'\right)^{1/r'}\):
Taking \(x'=a\) and adding \(\abs{u(a)}\) gives Equation (A.752).
∎The Arzelà–Ascoli theorem and the compact embedding
Let \(\left(u_{n}\right)\) be a sequence of continuous functions on \([a,b]\) that is
-
uniformly bounded: \(\abs{u_{n}(x)}\le K\) for all \(n\) and all \(x\); and
-
equicontinuous: for every \(\varepsilon>0\) there is \(\delta>0\), independent of \(n\), such that \(\abs{x-x'}<\delta\) implies \(\abs{u_{n}(x)-u_{n}\left(x'\right)}<\varepsilon\) for every \(n\).
Then some subsequence converges uniformly on \([a,b]\), and its limit is continuous. Rests on Theorems 7.7 and 7.8.
Derives Lemma A.454. A subsequence converging on a countable dense set. Enumerate the rational points of \([a,b]\) together with \(a\) and \(b\) as \(q_{1},q_{2},\ldots\); the set is countable and dense in \([a,b]\). The numerical sequence \(\left(u_{n}\left(q_{1}\right)\right)_{n}\) is bounded by \(K\), so by Bolzano–Weierstrass (Theorem 7.7) there is a subsequence \(\left(u^{(1)}_{n}\right)\) of \(\left(u_{n}\right)\) for which \(\left(u^{(1)}_{n}\left(q_{1}\right)\right)\) converges. Applying the same argument to \(\left(u^{(1)}_{n}\left(q_{2}\right)\right)\) produces a subsequence \(\left(u^{(2)}_{n}\right)\) of \(\left(u^{(1)}_{n}\right)\) converging at \(q_{1}\) and \(q_{2}\), and so on, giving nested subsequences \(\left(u^{(j)}_{n}\right)_{n}\), the \(j\)-th converging at \(q_{1},\ldots,q_{j}\). The diagonal sequence \(w_{k}=u^{(k)}_{k}\) is, from its \(j\)-th term onwards, a subsequence of \(\left(u^{(j)}_{n}\right)_{n}\), so \(\left(w_{k}\left(q_{j}\right) \right)_{k}\) converges for every \(j\).
Uniform convergence. Let \(\varepsilon>0\) and take \(\delta>0\) from equicontinuity for \(\varepsilon/3\). Partition \([a,b]\) into finitely many subintervals of length less than \(\delta\) and pick in each one a point of the dense set, obtaining \(q_{j_{1}},\ldots,q_{j_{m}}\) such that every \(x\in[a,b]\) lies within \(\delta\) of some \(q_{j_{i}}\). Because each of the \(m\) sequences \(\left(w_{k}\left(q_{j_{i}}\right)\right)_{k}\) converges, it is Cauchy (Theorem 7.8), and there is an \(N\) — the largest of the \(m\) thresholds — with \(\abs{w_{k}\left(q_{j_{i}}\right)-w_{l}\left(q_{j_{i}}\right)} <\varepsilon/3\) for all \(k,l\ge N\) and all \(i\). For an arbitrary \(x\in[a,b]\) choose \(i\) with \(\abs{x-q_{j_{i}}}<\delta\); then, for \(k,l\ge N\),
the outer terms by equicontinuity. The bound does not involve \(x\), so \(\left(w_{k}\right)\) is uniformly Cauchy; at each fixed \(x\) it therefore converges (Theorem 7.8), to a limit \(u(x)\), and letting \(l\to\infty\) in the display gives \(\abs{w_{k}(x)-u(x)}\le\varepsilon\) for every \(k\ge N\) and every \(x\) — uniform convergence.
Continuity of the limit. Given \(\varepsilon>0\), pick \(k\) with \(\max_{[a,b]}\abs{w_{k}-u}<\varepsilon/3\) and \(\delta\) from equicontinuity for \(\varepsilon/3\); then \(\abs{x-x'}<\delta\) gives
Let \(\left(u_{n}\right)\subset W^{1,r}(a,b)\) with
Then a subsequence of \(\left(u_{n}\right)\) converges uniformly on \([a,b]\) to a continuous limit \(u\). If in addition \(u_{n}'\rightharpoonup v\) weakly in \(L^{r}(a,b)\) along that subsequence — that is, \(\int_{a}^{b}g\,u_{n}'\,\dd x\longrightarrow\int_{a}^{b}g\,v\,\dd x\) for every \(g\in L^{r'}(a,b)\) — then \(u\in W^{1,r}(a,b)\) with \(u'=v\). Rests on Lemmas A.453 and A.454.
Derives Theorem A.455. Precompactness. By Equation (A.752) the sequence is uniformly bounded by \(K=M_{0}+M\left(b-a\right)^{1/r'}\), and by Equation (A.751)
for every \(n\), so the sequence is equicontinuous: given \(\varepsilon>0\), the choice \(\delta=\left(\varepsilon/M\right)^{r'}\) serves every \(n\) at once, and it is here that \(r>1\) is used, since \(1/r'>0\). Lemma A.454 now supplies a uniformly convergent subsequence, which we do not relabel, with continuous limit \(u\).
Identification of the derivative. Fix \(x\in[a,b]\) and let \(\chi_{x}\) be the function equal to \(1\) on \(\left(a,x\right)\) and to \(0\) on \(\left(x,b\right)\). It belongs to \(L^{r'}(a,b)\), being bounded on a bounded interval, so weak convergence gives
On the other hand Equation (A.748) for \(u_{n}\) says the left side equals \(u_{n}(x)-u_{n}(a)\), which converges to \(u(x)-u(a)\) by uniform — indeed pointwise — convergence. Hence \(u(x)=u(a)+\int_{a}^{x}v\) for every \(x\), which is Equation (A.748) for \(u\) with \(v\) in the role of the derivative; since \(v\in L^{r}\) by hypothesis, \(u\in W^{1,r}(a,b)\) and \(u'=v\).
∎Weak lower semicontinuity with $y$ present
Let \(f\) satisfy (H1)–(H4) of Theorem A.448. Let \(\left(u_{n}\right)\subset W^{1,r}(a,b)\) and \(u\in W^{1,r}(a,b)\) be such that
-
\(\sup_{n}\norm{u_{n}'}_{L^{r}}<\infty\);
-
\(u_{n}\longrightarrow u\) uniformly on \([a,b]\);
-
\(u_{n}'\rightharpoonup u'\) weakly in \(L^{r}(a,b)\).
Then
Rests on Lemma 16.72, Equation (16.62) and Theorem A.455.
Derives Theorem A.456. Write \(M=\sup_{n}\norm{u_{n}'}_{L^{r}}\) and \(R=\sup_{n}\max_{[a,b]}\abs{u_{n}}\), the latter finite because a uniformly convergent sequence of continuous functions is uniformly bounded; enlarge \(R\) if necessary so that also \(\max_{[a,b]}\abs{u}\le R\).
Step 1: the tangent inequality, taken at the limit slope. Convexity of a \(C^{1}\) function of \(q\) puts its graph above every tangent, which is Equation (16.62); applied at the fixed value \(y=u_{n}(x)\) of the second argument, with \(q_{1}=u'(x)\) and \(q_{2}=u_{n}'(x)\), it gives, for almost every \(x\),
Note that the tangent is taken at the limit slope \(u'\) but at the approximating value \(u_{n}\) of the function; this is the one departure from the \(y\)-free argument of Lemma 16.72, and it is what makes both terms on the right tractable. Integrating Equation (A.755),
Both integrals on the right are well defined: the second because \(g_{n}\in L^{r'}\) by Step 3 below and \(u_{n}'-u'\in L^{r}\), so Equation (A.750) applies; the first because its integrand is bounded below by \(-\beta\) through (H3), so the integral exists in \(\left(-\infty,+\infty\right]\).
Step 2: the first term. By (H1) the map \(y\longmapsto f\left(x,y,u'(x)\right)\) is continuous, and \(u_{n}(x)\longrightarrow u(x)\) for every \(x\), so
By (H3) the functions \(f\left(x,u_{n},u'\right)+\beta\) are nonnegative and measurable, so Fatou's lemma applies to them and gives
the constants \(\beta\left(b-a\right)\) cancelling from both sides. The inequality holds in \(\left(-\infty,+\infty\right]\), so nothing is assumed about the finiteness of either side.
Step 3: the second term vanishes. Put \(g=f_{q}\left(\cdot\,,u,u'\right)\). By (H4) with the constant \(C_{R}\),
and the right-hand side lies in \(L^{r'}(a,b)\), because \(\left(\abs{u'}^{r-1}\right)^{r'}=\abs{u'}^{r}\) is integrable and constants are integrable on a bounded interval. In particular \(g_{n},g\in L^{r'}\), as promised in Step 1. Moreover \(g_{n}\longrightarrow g\) pointwise almost everywhere, by the continuity of \(f_{q}\) in its second argument (H1) and the pointwise convergence \(u_{n}\to u\). Since \(\abs{g_{n}-g}^{r'}\le\left(2C_{R}\right)^{r'} \left(1+\abs{u'}^{r-1}\right)^{r'}\in L^{1}\) and tends to \(0\) almost everywhere, the dominated convergence theorem gives
Now split
The first integral tends to \(0\): \(g\in L^{r'}\) and \(u_{n}'\rightharpoonup u'\) in \(L^{r}\), which is exactly the statement that \(\int g\,u_{n}'\to\int g\,u'\). The second is bounded, by Hölder's inequality Equation (A.750), by
which tends to \(0\) by Equation (A.759) and hypothesis (1). Hence
Conclusion. Take the lower limit in Equation (A.756). The lower limit of a sum is at least the lower limit of the first summand plus the limit of the second when the latter exists, so by Equations (A.757) and (A.760)
which is Equation (A.754).
∎Proof of the theorem
Proof of Theorem A.448. Derives Theorem A.448. The infimum is finite. The affine function \(\ell(x)=y_{a}+\left(y_{b}-y_{a}\right)\left(x-a\right)/ \left(b-a\right)\) lies in \(\mathcal{A}_{r}\), with constant derivative, and \(f\) is continuous, so \(J[\ell]\) is finite and \(m=\inf_{\mathcal{A}_{r}}J\le J[\ell]<\infty\). By (H3), \(J[u]\ge-\beta\left(b-a\right)\) for every admissible \(u\), so \(m>-\infty\).
Coercivity bounds a minimising sequence. Let \(\left(u_{n}\right)\subset\mathcal{A}_{r}\) satisfy \(J\left[u_{n}\right]\longrightarrow m\); discarding finitely many terms we may assume \(J\left[u_{n}\right]\le m+1\) for every \(n\). By (H3),
so \(\norm{u_{n}'}_{L^{r}}^{r} \le\left[m+1+\beta\left(b-a\right)\right]/\alpha=:M^{r}\), a bound independent of \(n\). Since \(u_{n}(a)=y_{a}\) for every \(n\), the hypotheses Equation (A.753) of Theorem A.455 hold with \(M_{0}=\abs{y_{a}}\), and by Equation (A.752) the sequence is also bounded in \(L^{r}\), hence bounded in \(W^{1,r}(a,b)\).
Extraction of a limit. The sequence \(\left(u_{n}'\right)\) is bounded in \(L^{r}(a,b)\), a reflexive Banach space for \(1<r<\infty\), so some subsequence converges weakly, \(u_{n_{k}}'\rightharpoonup v\) in \(L^{r}(a,b)\) — this is the compactness input of Theorem 16.70, quoted there in the same form (see Remark A.459). Passing to a further subsequence, not relabelled, Theorem A.455 makes \(u_{n_{k}}\) converge uniformly on \([a,b]\) to a continuous \(u_{\ast}\), and identifies \(u_{\ast}\in W^{1,r}(a,b)\) with \(u_{\ast}'=v\).
The limit is admissible. Uniform convergence is in particular pointwise at the endpoints, so \(u_{\ast}(a)=\lim_{k}u_{n_{k}}(a)=y_{a}\) and \(u_{\ast}(b)=\lim_{k}u_{n_{k}}(b)=y_{b}\): the class \(\mathcal{A}_{r}\) is closed under this mode of convergence, which is the second place where the compact embedding earns its keep — weak convergence in \(W^{1,r}\) alone says nothing about the value of a function at a point.
The limit minimises. The three hypotheses of Theorem A.456 hold for \(\left(u_{n_{k}}\right)\) and \(u_{\ast}\), so
while \(J\left[u_{\ast}\right]\ge m\) because \(u_{\ast}\in\mathcal{A}_{r}\). Hence \(J\left[u_{\ast}\right]=m\) and the infimum is attained.
∎The proof just given is the abstract direct method with each of its three ingredients made concrete, and it is worth seeing which concrete statement plays which abstract role. Coercivity of \(J\) in the sense of Theorem 16.70 is Equation (A.761), and it controls only the derivative; that this suffices to control the function as well is Equation (A.752), i.e. the boundary condition plus Hölder's inequality. Sequential weak lower semicontinuity is Theorem A.456. The compactness that the abstract theorem imports from reflexivity is used here in two different topologies at once — weakly in \(L^{r}\) for the derivatives, uniformly in \(C^{0}\) for the functions — and the second is not a convenience but a necessity: without it neither the endpoint values nor the \(y\)-dependence of the integrand would survive the limit. That is the precise sense in which Lemma 16.72 left something owed.
If (H4) is dropped and \(f\) is asked only to be continuous and convex in \(q\) — Tonelli's own hypotheses — Step 3 above breaks down, because \(\abs{f_{q}\left(x,u_{n},u'\right)}\) can then grow with \(n\) faster than any fixed \(L^{r'}\) function, and the classical repair is a different one. It rests on the theorem of Scorza-Dragoni: a Carathéodory integrand — measurable in \(x\), continuous in \(\left(y,q\right)\) — is, after the deletion from \([a,b]\) of a set of arbitrarily small measure, continuous jointly in all its arguments on what remains; on that set the modulus of continuity in \(y\) is then uniform in \(q\) over compacta, which is what Step 3 needs, and the deleted set is handled by the coercivity bound (H3). This treatise's bibliography carries no entry for Scorza-Dragoni's paper, and the statement above is therefore quoted here without a citation — the honest record of a result that is named but neither used nor verified in this book. Nothing in Theorem A.448 depends on it: the theorem proved above is the theorem under (H1)–(H4), and every integrand this treatise minimises satisfies (H4) (Remark A.449).
Three analytic inputs are used and not proved in this treatise. Each belongs to the measure theory and functional analysis that Remark 12.1 of Hilbert Spaces already declares as imports, and all three are in Reed and Simon [Reed:1972].
-
The Lebesgue integral and its convergence theorems. Fatou's lemma is used once, in Equation (A.757), and the dominated convergence theorem once, in Equation (A.759). The absolute continuity of the integral, used in Definition A.450 to see that an element of \(W^{1,r}\) is continuous, is of the same family.
-
The duality \(\left(L^{r}\right)^{\ast}=L^{r'}\) and the reflexivity of \(L^{r}\) for \(1<r<\infty\), from which a bounded sequence in \(L^{r}\) has a weakly convergent subsequence. This is the only compactness assumption in the proof, and it is the same one Theorem 16.70 makes in the chapter, where it is named as the Eberlein–Smulian theorem; the special case \(r=2\) is the Hilbert-space statement of Hilbert Spaces.
-
The du Bois-Reymond lemma for integrable functions, used only in Remark A.451 to identify Definition A.450 with the distributional definition Definition 10.85. No step of the proof uses it; the continuous version proved in the chapter, Lemma 16.19, is not strong enough for the identification, which is why the space is defined here by Equation (A.748).
Everything else — the Arzelà–Ascoli theorem, the compact embedding, Young's and Hölder's inequalities, and the semicontinuity theorem itself — is proved above from Real Analysis and the chapter.
Theorem A.448 discharges Theorem 16.73 of Section 16.5.2, completing the direct method of Theorem 16.70 for the concrete functional Equation (16.11). Its hypotheses are what Remark 16.74 tests against the Lagrangians of mechanics, and the answer given there is unchanged by anything proved here: convexity in the velocities holds for a kinetic energy with a positive definite mass matrix, while coercivity (H3) demands a potential bounded above, so the theorem certifies existence for a free or repelled particle and not for a confining one. The failure of convexity is equally untouched — the functional Equation (16.64) satisfies (H1), (H3) and (H4) but not (H2), and has no minimiser.
Completeness of the Sturm–Liouville Eigenfunctions in the Energy Norm
This appendix proves the statement quoted in Remark 16.77 of Calculus of Variations: the eigenfunctions of a regular Sturm–Liouville problem are complete not merely in the weighted \(L^{2}\) space, but in the energy norm built from the quadratic form of the problem, and with that completeness the inequality Equation (16.70) sharpens into the identity Equation (16.71) for the energy form. The identity is what the Rayleigh–Ritz error estimate Corollary 16.78 needs; the inequality, which is what the variational characterisation Theorem 16.76 needs, is proved in the chapter from Bessel's inequality alone and is independent of everything below — a point Remark A.473 returns to, since a circular appendix would silently destroy the chapter's argument.
The route is the one the subject was born from [Hilbert:1912]: invert the differential operator. The inverse of a regular Sturm–Liouville operator with Dirichlet data is an integral operator whose kernel is the Green function of the problem; that kernel is continuous, so the operator is compact and self-adjoint, and the Hilbert–Schmidt theorem Theorem 12.44, proved in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators, hands back an orthonormal basis of eigenfunctions. Completeness in the weighted \(L^{2}\) space — which Theorem 16.76 assumes — then comes out as a theorem, and completeness in the energy norm follows from it by one identity relating the two inner products.
Throughout, the data are those of Definition 16.75: \(p\), \(q\), \(r\) are continuous and real on \([a,b]\) with \(p>0\) and \(r>0\), and the boundary conditions are \(y(a)=y(b)=0\). As in the proof of Theorem 16.76, fix once and for all a constant \(K\) with
which is possible because \(q\) is continuous and \(r>0\) on a compact interval, and write
The operator of the problem, shifted by \(K\), is
so that \(y\) solves Equation (16.66) with eigenvalue \(\lambda\) if and only if \(L_{K}[y]=\left(\lambda+K\right)r\,y\).
Statement
Let \(p,q,r\in C^{0}[a,b]\) with \(p>0\) and \(r>0\). Then:
-
the Dirichlet problem Equation (16.66) has an infinite sequence of real eigenvalues \(\lambda_{1}\le\lambda_{2}\le\cdots\) with \(\lambda_{n}\longrightarrow+\infty\), each of finite multiplicity, and eigenfunctions \(y_{n}\) that are orthonormal in \(\avg{\cdot,\cdot}_{r}\) and form an orthonormal basis of the weighted space \(L^{2}_{r}(a,b)\);
-
the rescaled eigenfunctions \(y_{n}/\sqrt{\lambda_{n}+K}\) form an orthonormal basis of the energy space \(H_{E}\) of Definition A.462 with the inner product \(B_{K}\);
-
consequently, for every \(y\in H_{E}\) — in particular for every \(y\in C^{1}_{0}[a,b]\) — the partial sums \(S_{N}=\sum_{n\le N}c_{n}y_{n}\), \(c_{n}=\avg{y,y_{n}}_{r}\), satisfy \(B_{K}\left[y-S_{N},y-S_{N}\right]\longrightarrow0\), and
\begin{equation}\tag{A.765} \int_{a}^{b}\left(p\,y'^{2}+q\,y^{2}\right)\dd x =\sum_{n\ge1}\lambda_{n}\,c_{n}^{2}\ec \end{equation}the series on the right converging.
Rests on Definition 16.75, Theorem 12.44 and Equation (16.66).
Equation (A.765) is Equation (16.71). The proof occupies the rest of the section: the energy space (The energy space), the Green function and the inverse operator (The Green function and the inverse operator), its compactness and the spectral decomposition (Compactness and the spectral decomposition), and the transfer to the energy norm (From the weighted norm to the energy norm).
The energy space
\(L^{2}_{r}(a,b)\) is the space of square-integrable functions with the inner product \(\avg{\cdot,\cdot}_{r}\) of Equation (A.763), and
is the energy space, where \(W^{1,2}(a,b)\) is the space of Definition A.450 with \(r=2\): the continuous functions \(u\) for which there is \(u'\in L^{2}(a,b)\) with \(u(x)=u(a)+\int_{a}^{x}u'\). On \(H_{E}\) the energy form \(B_{K}\) of Equation (A.763) is used as an inner product, with norm \(\norm{u}_{E}=B_{K}[u,u]^{1/2}\). Rests on Definitions 16.75 and A.450.
Write \(p_{-}=\min_{[a,b]}p>0\), \(p_{+}=\max_{[a,b]}p\), \(r_{-}=\min_{[a,b]}r>0\), \(r_{+}=\max_{[a,b]}r\) and \(\kappa_{+}=\max_{[a,b]}\left(q+Kr\right)\). For every \(u\in H_{E}\),
and consequently
In particular \(B_{K}\) is an inner product on \(H_{E}\): it is bilinear and symmetric by inspection, and \(B_{K}[u,u]=0\) forces \(u=0\). Rests on Definition A.462 and Lemma A.452.
Derives Lemma A.463. For \(u\in H_{E}\) and \(x\in[a,b]\), \(u(a)=0\) gives \(u(x)=\int_{a}^{x}u'\), and the Cauchy–Schwarz case \(r=2\) of Hölder's inequality Equation (A.750) bounds this by \(\left(x-a\right)^{1/2}\norm{u'}_{L^{2}} \le\left(b-a\right)^{1/2}\norm{u'}_{L^{2}}\), which is the first inequality of Equation (A.767). Integrating its square against \(r\) over \([a,b]\) gives the second.
For Equation (A.768), the lower bound drops the term \(\left(q+Kr\right)u^{2}\ge0\), which is nonnegative by Equation (A.762), and bounds \(p\ge p_{-}\); the upper bound uses \(p\le p_{+}\), \(q+Kr\le\kappa_{+}\) and the second inequality of Equation (A.767) with \(r\) replaced by the constant \(1\), i.e.\ \(\int u^{2}\le\left(b-a\right)^{2}\norm{u'}^{2}_{L^{2}}\). If \(B_{K}[u,u]=0\) then \(\norm{u'}_{L^{2}}=0\) by the lower bound, so \(u'=0\) almost everywhere and \(u(x)=u(a)+\int_{a}^{x}u'=0\).
∎\(\left(H_{E},B_{K}\right)\) is complete, hence a Hilbert space in the sense of Definition 12.2, and the inclusion \(H_{E}\subset L^{2}_{r}(a,b)\) is continuous. Rests on Lemma A.463 and Theorem 12.12.
Derives Lemma A.464. Continuity of the inclusion is the second inequality of Equation (A.767) combined with the lower bound in Equation (A.768): \(\avg{u,u}_{r}\le r_{+}\left(b-a\right)^{2}p_{-}^{-1}B_{K}[u,u]\).
Let \(\left(u_{k}\right)\) be Cauchy in \(\norm{\cdot}_{E}\). By the lower bound in Equation (A.768) the sequence \(\left(u_{k}'\right)\) is Cauchy in \(L^{2}(a,b)\), which is complete (Theorem 12.12), so \(u_{k}'\longrightarrow v\) in \(L^{2}\). By the first inequality of Equation (A.767) applied to \(u_{k}-u_{l}\), the sequence \(\left(u_{k}\right)\) is uniformly Cauchy on \([a,b]\), hence converges uniformly to a continuous \(u\) with \(u(a)=u(b)=0\). Passing to the limit in \(u_{k}(x)=\int_{a}^{x}u_{k}'\) — legitimate because \(\abs{\int_{a}^{x}\left(u_{k}'-v\right)} \le\left(b-a\right)^{1/2}\norm{u_{k}'-v}_{L^{2}} \longrightarrow0\) by Equation (A.750) — gives \(u(x)=\int_{a}^{x}v\), so \(u\in H_{E}\) with \(u'=v\). Finally the upper bound in Equation (A.768) applied to \(u_{k}-u\) gives \(\norm{u_{k}-u}_{E}\longrightarrow0\).
∎\(C^{1}_{0}[a,b]\subset H_{E}\), since a continuously differentiable function vanishing at both ends satisfies Equation (A.748) with \(u'\) its classical derivative. So every statement below applies to the trial functions of Theorem 16.76 and Proposition 16.79 without further comment. The space \(H_{E}\) is the \(H^{1}_{0}(a,b)\) of Definition 10.85: the identification of the two descriptions is Remark A.451 together with the fact that a \(W^{1,2}\) function vanishing at both endpoints is a uniform limit of smooth functions of compact support, which is not needed here and is not proved.
The Green function and the inverse operator
Let \(\varphi\) solve \(L_{K}[\varphi]=0\) on \([a,b]\) with \(\varphi(a)=\varphi(b)=0\). Then \(\varphi\equiv0\). Rests on Lemma A.463 and Equation (A.764).
Derives Lemma A.466. Multiply \(L_{K}[\varphi]=0\) by \(\varphi\) and integrate over \([a,b]\); one integration by parts (Equation (7.28)), whose boundary term \(\left[p\,\varphi\,\varphi'\right]_{a}^{b}\) vanishes because \(\varphi\) does at both ends, gives \(B_{K}[\varphi,\varphi]=0\). Since \(\varphi\in C^{1}[a,b]\) vanishes at \(a\), it lies in \(H_{E}\), and Lemma A.463 forces \(\varphi=0\).
∎Let \(\phi\) and \(\psi\) be the solutions of \(L_{K}[u]=0\) determined by
which exist, are unique and are defined on the whole of \([a,b]\) by Lemma A.439 applied to the system form of \(L_{K}[u]=0\) — the equation \(\left(p\,u'\right)'=\left(q+Kr\right)u\) is Equation (A.734) with \(P=p\) and \(Q=q+Kr\). Put
a nonzero constant, and define the Green function
The constant \(C\) of Equation (A.770) is independent of \(x\) and nonzero; \(G\) is well defined by Equation (A.771) (the two branches agree on \(x=s\)), symmetric, \(G\left(x,s\right)=G\left(s,x\right)\), continuous on \([a,b]\times[a,b]\), and there is a constant \(L_{G}\) with
Rests on Definition A.467 and Theorem 7.35.
Derives Lemma A.468. That \(C\) is constant is Lemma A.442 applied with \(P=p\) and \(Q=q+Kr\). If \(C=0\) the same lemma makes \(\phi\) and \(\psi\) linearly dependent, so \(\phi\) would vanish at \(b\) as \(\psi\) does, and \(\phi\) would be a nontrivial solution of the homogeneous Dirichlet problem — nontrivial because \(\left(p\phi'\right)(a)=1\) — contradicting Lemma A.466. Hence \(C\neq0\).
The two branches of Equation (A.771) coincide when \(x=s\), so \(G\) is well defined, and exchanging \(x\) and \(s\) exchanges the two branches, which is the symmetry. Both \(\phi\) and \(\psi\) are of class \(C^{1}\) on the compact interval, hence Lipschitz there by the mean value theorem (Theorem 7.35) with constants \(\max\abs{\phi'}\) and \(\max\abs{\psi'}\); each branch of Equation (A.771) is therefore Lipschitz in \(x\) uniformly in \(s\), with \(L_{G}=\abs{C}^{-1}\max\set{\max\abs{\phi'}\max\abs{\psi}, \max\abs{\phi}\max\abs{\psi'}}\), and since the branches agree where they meet, the bound Equation (A.772) holds for \(x,x'\) on opposite sides of \(s\) as well, by splitting the increment at \(s\). Continuity follows, jointly in \(\left(x,s\right)\), from Equation (A.772) and the symmetry.
∎Define
Then:
-
for continuous \(f\), \(u=Tf\) is of class \(C^{1}\) with \(p\,u'\) of class \(C^{1}\), and it is the unique solution of
\begin{equation}\tag{A.774} L_{K}[u]=r\,f\ec\qquad u(a)=u(b)=0\ep \end{equation} -
\(T\) maps \(L^{2}_{r}(a,b)\) into \(H_{E}\), with \(\norm{Tf}_{E}\le C_{P}\norm{f}_{r}\) where \(C_{P}=\left(r_{+}\left(b-a\right)^{2}/p_{-}\right)^{1/2}\), and
\begin{equation}\tag{A.775} B_{K}\left[Tf,\chi\right]=\avg{f,\chi}_{r} \qquad\text{for every }\chi\in H_{E}\ep \end{equation} -
\(T\) is self-adjoint and positive on \(L^{2}_{r}(a,b)\), and injective.
Rests on Definition A.467, Lemma A.463 and Lemma A.466.
Derives Theorem A.469. (1) Split Equation (A.773) at \(x\) using Equation (A.771):
the first branch of Equation (A.771) governing \(s\ge x\) and the second \(s\le x\). Both integrals have continuous integrands, so by the fundamental theorem of calculus (Theorem 7.42) the right side is differentiable and
the last bracket vanishing identically: the two boundary terms produced by differentiating the limits of integration cancel. Hence
which is again differentiable because \(p\phi'\) and \(p\psi'\) are of class \(C^{1}\), with \(\left(p\phi'\right)'=\left(q+Kr\right)\phi\) and likewise for \(\psi\). Differentiating,
using Equation (A.776) for the first two terms and Equation (A.770) for the bracket, which is \(\phi\left(p\psi'\right)-\psi\left(p\phi'\right) =p\left(\phi\psi'-\psi\phi'\right)=C\). Dividing by \(-C\) gives \(\left(p\,u'\right)'=\left(q+Kr\right)u-r\,f\), that is Equation (A.774). The boundary conditions hold because \(x=a\) kills the first integral in Equation (A.776) and leaves \(\phi(a)=0\) in the second, and symmetrically at \(x=b\) with \(\psi(b)=0\). Uniqueness is Lemma A.466 applied to the difference of two solutions.
(2) Let first \(f\) be continuous and \(u=Tf\). Then \(u\in C^{1}[a,b]\) with \(u(a)=u(b)=0\), so \(u\in H_{E}\), and for \(\chi\in C^{1}_{0}[a,b]\) one integration by parts (Equation (7.28)) applied to Equation (A.774) gives \(B_{K}[u,\chi]=\avg{f,\chi}_{r}\), the boundary term \(\left[p\,u'\chi\right]_{a}^{b}\) vanishing because \(\chi\) does. The same identity for arbitrary \(\chi\in H_{E}\) follows because both sides are continuous in \(\chi\) for \(\norm{\cdot}_{E}\) — the left side by the Cauchy–Schwarz inequality for the inner product \(B_{K}\), the right by Lemma A.464 — and, by the argument of Lemma A.464 run backwards, each \(\chi\in H_{E}\) is an \(\norm{\cdot}_{E}\)-limit of members of \(C^{1}_{0}[a,b]\): take \(\chi_{k}(x)=\int_{a}^{x}v_{k}\) with \(v_{k}\) continuous, \(\int_{a}^{b}v_{k}=0\) and \(v_{k}\longrightarrow\chi'\) in \(L^{2}\), which is possible because the continuous functions are dense in \(L^{2}\) (Remark A.474) and the mean may be subtracted off with a loss tending to zero.
Taking \(\chi=u\) in Equation (A.775) and using Equation (A.767),
so \(\norm{Tf}_{E}\le C_{P}\norm{f}_{r}\) for continuous \(f\). For general \(f\in L^{2}_{r}\) take continuous \(f_{k}\longrightarrow f\) in \(L^{2}_{r}\); by Equation (A.777) the sequence \(\left(Tf_{k}\right)\) is Cauchy in \(H_{E}\), hence convergent there (Lemma A.464), while by Equation (A.773) and the Cauchy–Schwarz inequality \(\abs{Tf_{k}(x)-Tf(x)}\le\max\abs{G} \left(r_{+}\left(b-a\right)\right)^{1/2}\norm{f_{k}-f}_{r}\) tends to zero uniformly. The two limits agree, so \(Tf\in H_{E}\), the bound persists, and Equation (A.775) passes to the limit.
(3) For \(f,g\in L^{2}_{r}\) put \(u=Tf\), \(w=Tg\). Then Equation (A.775) twice, with the symmetry of \(B_{K}\) in between, gives
so \(T\) is self-adjoint; and \(\avg{Tf,f}_{r}=B_{K}[u,u]\ge0\), so \(T\) is positive. If \(Tf=0\) then \(B_{K}[0,\chi]=\avg{f,\chi}_{r}=0\) for every \(\chi\in H_{E}\), and since \(H_{E}\) is dense in \(L^{2}_{r}\) — it contains the functions \(\chi_{k}\) built above from an arbitrary continuous \(v\), and the continuous functions are dense — this forces \(f=0\). Hence \(T\) is injective.
∎Compactness and the spectral decomposition
\(T\) maps every bounded sequence of \(L^{2}_{r}(a,b)\) to a sequence with a convergent subsequence in \(L^{2}_{r}(a,b)\); that is, \(T\) is compact in the sense of Definition 12.41. Rests on Lemmas A.454 and A.468.
Derives Theorem A.470. Let \(\norm{f_{n}}_{r}\le M\). By the Cauchy–Schwarz inequality in \(L^{2}_{r}\) applied to Equation (A.773),
so the sequence \(\left(Tf_{n}\right)\) is uniformly bounded, and by the Lipschitz estimate Equation (A.772),
so it is equicontinuous, with a modulus independent of \(n\). By the Arzelà–Ascoli theorem Lemma A.454 a subsequence converges uniformly on \([a,b]\), and uniform convergence implies convergence in \(L^{2}_{r}\), since \(\norm{w}_{r}^{2}\le r_{+}\left(b-a\right)\max\abs{w}^{2}\).
∎There is an orthonormal basis \(\set{y_{n}}_{n\ge1}\) of \(L^{2}_{r}(a,b)\) and numbers \(\mu_{n}>0\) with \(\mu_{n}\longrightarrow0\) such that \(Ty_{n}=\mu_{n}y_{n}\). Setting \(\lambda_{n}=\mu_{n}^{-1}-K\), each \(y_{n}\) is a classical solution of the Sturm–Liouville problem Equation (16.66) with eigenvalue \(\lambda_{n}\), and \(\lambda_{n}\longrightarrow+\infty\). Every eigenvalue has finite multiplicity, and the \(\lambda_{n}\) may be enumerated in nondecreasing order. Rests on Theorems 12.44, A.469 and A.470.
Derives Theorem A.471. \(L^{2}_{r}(a,b)\) is a Hilbert space: its inner product differs from that of \(L^{2}(a,b)\) by the factor \(r\), which is bounded between the positive constants \(r_{-}\) and \(r_{+}\), so the two norms are equivalent and completeness transfers from Theorem 12.12. On it, \(T\) is bounded (by Equation (A.777) and Equation (A.767)), self-adjoint and compact, by Theorems A.469 and A.470. The Hilbert–Schmidt theorem Theorem 12.44, proved in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators, therefore supplies real numbers \(\mu_{n}\longrightarrow0\) and an orthonormal system \(\set{y_{n}}\) with
The system is a basis. If \(f\) is orthogonal to every \(y_{n}\) then Equation (A.778) gives \(Tf=0\), and \(T\) is injective by Theorem A.469(3), so \(f=0\). By Definition 12.29 the system is maximal, i.e. an orthonormal basis, and Theorem 12.30 applies to it.
The eigenvalues are positive. Taking \(f=y_{n}\) in Equation (A.778) gives \(Ty_{n}=\mu_{n}y_{n}\), and \(\mu_{n}=\avg{Ty_{n},y_{n}}_{r}\ge0\) by positivity; \(\mu_{n}=0\) is excluded by injectivity. Hence \(\mu_{n}>0\) and \(\lambda_{n}=\mu_{n}^{-1}-K\) is well defined, with \(\lambda_{n}\longrightarrow+\infty\) because \(\mu_{n}\longrightarrow0^{+}\).
The eigenfunctions solve the differential equation. From \(y_{n}=\mu_{n}^{-1}Ty_{n}\) and Theorem A.469(2), \(y_{n}\in H_{E}\); in particular \(y_{n}\) is continuous, so Theorem A.469(1) applies to \(f=y_{n}\) and makes \(u=Ty_{n}=\mu_{n}y_{n}\) a classical solution of \(L_{K}[u]=r\,y_{n}\) with \(u(a)=u(b)=0\). Dividing by \(\mu_{n}\),
which, on subtracting \(K\,r\,y_{n}\) from both sides, is Equation (16.66) with eigenvalue \(\lambda_{n}\) and Dirichlet data.
Finite multiplicity and ordering. An eigenvalue \(\lambda\) of Equation (16.66) corresponds to the eigenvalue \(\mu=\left(\lambda+K\right)^{-1}\) of \(T\), whose eigenspace is finite dimensional by Theorem 12.44. Since \(\lambda_{n}\longrightarrow+\infty\), only finitely many \(\lambda_{n}\) lie below any given bound, so the sequence may be reordered nondecreasingly, and the \(y_{n}\) with it.
∎From the weighted norm to the energy norm
For every \(\chi\in H_{E}\) and every \(n\),
In particular \(B_{K}\left[y_{m},y_{n}\right] =\left(\lambda_{n}+K\right)\delta_{mn}\), and \(\lambda_{n}+K>0\) for every \(n\). Rests on Theorems A.469 and A.471.
Derives Lemma A.472. Apply Equation (A.775) with \(f=y_{n}\): \(B_{K}\left[Ty_{n},\chi\right]=\avg{y_{n},\chi}_{r}\). Since \(Ty_{n}=\mu_{n}y_{n}\) and \(B_{K}\) is bilinear, the left side is \(\mu_{n}B_{K}\left[y_{n},\chi\right]\); dividing by \(\mu_{n}>0\) and writing \(\mu_{n}^{-1}=\lambda_{n}+K\) gives Equation (A.779). Taking \(\chi=y_{m}\) and using orthonormality in \(\avg{\cdot,\cdot}_{r}\) gives the second statement, and \(\lambda_{n}+K=\mu_{n}^{-1}>0\).
∎Proof of Theorem A.461. Derives Theorem A.461. Part (1) is Theorem A.471. For (2), put
By Lemma A.472 the system \(\set{\tilde y_{n}}\) is orthonormal for the inner product \(B_{K}\):
It is maximal in \(H_{E}\): if \(\chi\in H_{E}\) satisfies \(B_{K}\left[\tilde y_{n},\chi\right]=0\) for every \(n\), then Equation (A.779) gives \(\avg{y_{n},\chi}_{r}=0\) for every \(n\), and \(\set{y_{n}}\) is an orthonormal basis of \(L^{2}_{r}\) by part (1), so \(\chi=0\) as an element of \(L^{2}_{r}\); being continuous, \(\chi\) vanishes identically. Since \(\left(H_{E},B_{K}\right)\) is a Hilbert space (Lemma A.464), maximality is completeness: Theorem 12.30 applies and \(\set{\tilde y_{n}}\) is an orthonormal basis of \(H_{E}\).
For (3), let \(y\in H_{E}\) and \(c_{n}=\avg{y,y_{n}}_{r}\). The energy-Fourier coefficients of \(y\) are, by Equation (A.779),
so the \(N\)-th partial sum of the expansion of \(y\) in the basis \(\set{\tilde y_{n}}\) is
exactly the partial sum formed in the chapter. By Equation (12.18) the expansion converges in the norm of \(H_{E}\), which says \(B_{K}\left[y-S_{N},y-S_{N}\right]\longrightarrow0\), and by Parseval's identity Equation (12.19) together with Equation (A.781),
the series converging because it is a Parseval sum. Finally, \(\set{y_{n}}\) is an orthonormal basis of \(L^{2}_{r}\), so Parseval's identity there gives \(\avg{y,y}_{r}=\sum_{n}c_{n}^{2}\); subtracting \(K\) times that from Equation (A.782) and using \(B_{K}[y,y]=\int\left(p\,y'^{2}+q\,y^{2}\right)\dd x +K\avg{y,y}_{r}\) leaves Equation (A.765). Both series converge separately — the first as a Parseval sum, the second because \(\sum_{n}c_{n}^{2}=\avg{y,y}_{r}<\infty\) — so the subtraction is legitimate.
∎Remark 16.77 insists that the inequality Equation (16.70), which is what the variational characterisation Theorem 16.76 rests on, is independent of the identity proved here, and that remains true: nothing above is used in the chapter's proof of Theorem 16.76, which argues from Equation (16.68) and the pairing Equation (16.69) to \(B_{K}[y,y]\ge\sum_{n\le N}\left(\lambda_{n}+K\right)c_{n}^{2}\) for every finite \(N\) — Bessel's inequality in the energy inner product, which needs no completeness at all. Conversely nothing in this appendix uses Theorem 16.76, so there is no circle. What this appendix does remove is a hypothesis: Theorem 16.76 assumes that the eigenfunctions are complete in \(L^{2}_{r}\), and Theorem A.471 proves it. The difference between the two statements is exactly the difference between Bessel's inequality Equation (12.15) and Parseval's identity Equation (12.19), which is the four-way criterion Theorem 12.30 in the one place where the distinction has a consequence a reader can feel: by Corollary 16.78 the Rayleigh–Ritz error is second order in the deviation measured in the energy norm, and the estimate is vacuous without Equation (A.765).
Two inputs are used and not proved in this section, and one further result is proved elsewhere in this appendix.
-
The completeness of \(L^{2}\) (Theorem 12.12) and, with it, the density of the continuous functions in \(L^{2}(a,b)\). The first is quoted in Hilbert Spaces itself (Remark 12.1) as resting on the convergence theorems of Lebesgue integration, which this treatise does not develop; the second belongs to the same package and is used twice above, both times to pass from continuous \(f\) to \(f\in L^{2}_{r}\) in Theorem A.469. Reed and Simon [Reed:1972] is the reference of record for both.
-
Nothing else. The Hilbert–Schmidt theorem Theorem 12.44 is not an import: it is proved, from the numerical-radius formula and sequential compactness, in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators of this appendix. The Arzelà–Ascoli theorem used for the compactness of \(T\) is proved in Lemma A.454, and the existence and uniqueness theory for the second-order equation defining \(\phi\) and \(\psi\) is Lemma A.439, both in this appendix.
Theorem A.461 discharges the derivation owed in Remark 16.77 of Section 16.5.3: the eigenfunctions of a regular Sturm–Liouville problem are complete in the energy norm, and Equation (16.71) holds for every admissible trial function. The identity is used in the chapter to make the error of a Ritz estimate second order (Corollary 16.78), and through that corollary it underwrites the convergence of the Ritz scheme Equation (16.72) and of its descendants — the finite element method, the linear variational method for molecular orbitals, and the variational principle for the ground-state energy of Approximation Methods. The eigenvalue problem itself is that of Ordinary Differential Equations and Sturm–Liouville Theory, whose orthogonality theorem Theorem 9.57 is here recovered as a by-product: eigenfunctions belonging to different eigenvalues are orthogonal in \(\avg{\cdot,\cdot}_{r}\) because they are eigenvectors of a self-adjoint operator for different eigenvalues.
The Plateau Problem: the Architecture of Douglas' Proof
This appendix accompanies Theorem 16.81 of Calculus of Variations: every rectifiable Jordan curve in \(\R^{3}\) bounds a disc-type minimal surface. It differs from every other section of this appendix in what it claims. Douglas' proof [Douglas:1931] and Radó's [Rado:1930] are each of monograph length and neither is reproduced; what is given here is the architecture of Douglas' argument, with every step that can be carried out inside this treatise carried out in full and every step that cannot stated precisely, as a displayed theorem, and named as an import. Remark A.495 lists the imports exactly, and the reader who wants the honest one-sentence summary should read that remark first: this section reduces Theorem 16.81 to two named results it does not prove, namely the solvability of the Dirichlet problem for the disc by the Poisson integral and Douglas' inner-variation argument for conformality, together with the classification of the conformal automorphisms of the disc.
What is proved here, from scratch, is the chain that makes those two imports the only ones: that the Dirichlet integral dominates the area with equality exactly for conformal maps (Proposition A.478); that the Dirichlet integral is conformally invariant, which is what makes the non-compactness of the conformal group a real obstruction rather than a technicality (Lemma A.479); that Douglas' boundary functional equals the Dirichlet integral of the harmonic extension, with the constant made explicit (Theorem A.482); that it is lower semicontinuous (Lemma A.484); that the three-point normalisation restores compactness, through the Courant–Lebesgue estimate and the Arzelà–Ascoli theorem (Compactness: the Courant–Lebesgue estimate); and that a harmonic conformal map has vanishing mean curvature, so that a minimiser really is a minimal surface (Proposition A.493).
Setting and statement
Let \(B=\set{z\in\C\mid\abs{z}<1}\) be the open unit disc, with \(z=u+\ii v=\rho\,\ee^{\ii\theta}\), and \(S^{1}=\pp B\) its boundary circle. A disc-type surface is a map \(\vect{X}:\overline{B}\longrightarrow\R^{3}\), continuous on \(\overline{B}\) and of class \(C^{1}\) on \(B\). Its first fundamental coefficients (Definition 13.22) are
and its area and Dirichlet functionals are
\(\vect{X}\) is conformal at a point if \(E=G\) and \(F=0\) there. Rests on Definition 13.22 and Equation (16.60).
Let \(\Gamma\subset\R^{3}\) be a rectifiable Jordan curve. Then there exists a disc-type surface \(\vect{X}\), harmonic and conformal on \(B\), whose restriction to \(S^{1}\) is a monotone parametrisation of \(\Gamma\) and which minimises the area among all disc-type surfaces spanning \(\Gamma\). Its image is a minimal surface. Rests on Equation (16.7) and Theorem 16.70.
The strategy is not to minimise \(A\), which is invariant under all reparametrisations of the disc and therefore has no compactness whatever, but to minimise \(D\), which is invariant only under the conformal ones — and then to arrange that the minimiser be conformal, so that the two functionals agree on it by Proposition A.478. Douglas' innovation is to push the minimisation onto the boundary circle, where the competitors are parametrisations of \(\Gamma\) and the functional is Equation (A.786).
The Dirichlet integral dominates the area
For every disc-type surface, \(D\left[\vect{X}\right]\ge A\left[\vect{X}\right]\), with equality if and only if \(\vect{X}\) is conformal almost everywhere on \(B\). Rests on Definition A.476.
Derives Proposition A.478. Pointwise on \(B\), the arithmetic–geometric mean inequality \(\left(\sqrt{E}-\sqrt{G}\right)^{2}\ge0\) gives \(E+G\ge2\sqrt{EG}\), with equality if and only if \(E=G\); and \(\sqrt{EG}\ge\sqrt{EG-F^{2}}\), with equality if and only if \(F=0\), since \(F^{2}\ge0\) and both radicands are nonnegative — indeed \(EG-F^{2}=\abs{\vect{X}_{u}}^{2}\abs{\vect{X}_{v}}^{2} -\left(\vect{X}_{u}\cdot\vect{X}_{v}\right)^{2} =\abs{\vect{X}_{u}\times\vect{X}_{v}}^{2}\ge0\) by Lagrange's identity. Chaining the two,
and integrating over \(B\) gives \(D\ge A\). Equality of the integrals of two continuous functions ordered pointwise forces equality of the functions almost everywhere, and by the two equality cases just identified that means \(E=G\) and \(F=0\) almost everywhere, i.e.\ conformality.
∎Let \(\tau:B\longrightarrow B\) be a bijective holomorphic map, written in real coordinates as \(\tau\left(u,v\right)=\left(\sigma\left(u,v\right), \varrho\left(u,v\right)\right)\). Then \(D\left[\vect{X}\circ\tau\right]=D\left[\vect{X}\right]\) for every disc-type surface \(\vect{X}\). Rests on Definition A.476 and Proposition 7.31.
Derives Lemma A.479. Write \(\vect{Y}=\vect{X}\circ\tau\). By the chain rule (Proposition 7.31), for each of the three components,
the derivatives of \(\vect{X}\) being evaluated at \(\tau\left(u,v\right)\). The Cauchy–Riemann equations for \(\tau\) read \(\sigma_{u}=\varrho_{v}\) and \(\sigma_{v}=-\varrho_{u}\), so with \(J=\sigma_{u}^{2}+\varrho_{u}^{2}\) — which is \(\abs{\tau'}^{2}\) and also the Jacobian determinant \(\sigma_{u}\varrho_{v}-\sigma_{v}\varrho_{u}\) — one finds
because \(\sigma_{u}^{2}+\sigma_{v}^{2} =\varrho_{u}^{2}+\varrho_{v}^{2}=J\) and \(\sigma_{u}\varrho_{u}+\sigma_{v}\varrho_{v} =\sigma_{u}\varrho_{u}-\varrho_{u}\sigma_{u}=0\) by Cauchy–Riemann. Hence the integrand of \(D\left[\vect{Y}\right]\) is the integrand of \(D\left[\vect{X}\right]\) composed with \(\tau\) and multiplied by the Jacobian \(J\), and the change-of-variables formula for a bijective \(C^{1}\) substitution gives \(D\left[\vect{Y}\right]=D\left[\vect{X}\right]\).
∎The area functional is invariant under every diffeomorphism of the disc, an infinite-dimensional group, and a minimising sequence for \(A\) may therefore be dragged about by reparametrisations without any change in the value being minimised: no compactness argument can survive that. Passing to \(D\) cuts the invariance group down to the conformal automorphisms of \(B\), which form only a three-parameter family — but it is still a non-compact one, and that residual non-compactness is exactly the difficulty Douglas' three-point normalisation removes. Lemma A.485 makes the statement precise.
The Douglas functional
For a bounded measurable \(\vect{g}:S^{1}\longrightarrow\R^{3}\), written as a function of the angle, \(\vect{g}=\vect{g}(\theta)\), put
Rests on Definition A.476.
Let \(\vect{g}\in L^{2}\left(S^{1};\R^{3}\right)\) have Fourier coefficients \(\vect{c}_{n}=\left(2\pi\right)^{-1}\int_{0}^{2\pi} \vect{g}(\theta)\,\ee^{-\ii n\theta}\,\dd\theta\in\C^{3}\), and let
be its harmonic extension to \(B\). Then
both sides being \(+\infty\) together. Rests on Definition A.481 and Equation (12.19).
Derives Theorem A.482. Throughout, \(\set{\left(2\pi\right)^{-1/2}\ee^{\ii n\theta}}_{n\in\Z}\) is the standard orthonormal basis of \(L^{2}\left(S^{1}\right)\) (Section 9.7.4), and Parseval's identity Equation (12.19) is applied componentwise to the three components of \(\vect{g}\), so that \(\abs{\vect{c}_{n}}^{2}\) denotes \(\sum_{k=1}^{3}\abs{c_{n}^{k}}^{2}\).
Step 1: the Dirichlet integral of the extension. In polar coordinates \(\dd u\,\dd v=\rho\,\dd\rho\,\dd\theta\) and
Differentiating Equation (A.787) term by term — legitimate on every disc \(\rho\le\rho_{0}<1\), where the series and each of its derived series are dominated termwise by \(\abs{n}^{2}\abs{\vect{c}_{n}}\rho_{0}^{\abs{n}-2}\), whose sum converges because \(\abs{\vect{c}_{n}}\) is bounded and \(\rho_{0}<1\), so that the Weierstrass \(M\)-test Lemma 9.6 gives uniform convergence — and integrating in \(\theta\) by orthogonality,
The two are equal — a first sign that the extension is conformally balanced — so
the term-by-term integration in \(\rho\) being legitimate because every term is nonnegative.
Step 2: the Douglas integral. Substitute \(\theta=\varphi+t\) and integrate first in \(\varphi\) at fixed \(t\). The Fourier coefficients of \(\varphi\longmapsto\vect{g}\left(\varphi+t\right)-\vect{g}(\varphi)\) are \(\vect{c}_{n}\left(\ee^{\ii nt}-1\right)\), so Parseval Equation (12.19) gives
whence
the interchange of sum and integral being legitimate because every term is nonnegative.
Step 3: the kernel integral. Using \(\sin^{2}\left(t/2\right)=\left(1-\cos t\right)/2\) and, for \(n\ge1\), the Dirichlet kernel \(\Delta_{n}(t)=\sum_{j=0}^{n-1}\ee^{\ii jt} =\left(\ee^{\ii nt}-1\right)/\left(\ee^{\ii t}-1\right)\),
where \(\abs{\ee^{\ii\alpha}-1}^{2}=2\left(1-\cos\alpha\right)\) has been used twice. By orthogonality of the exponentials, \(\int_{0}^{2\pi}\abs{\Delta_{n}}^{2}\dd t=2\pi n\), so the integral in Equation (A.790) equals \(4\pi\abs{n}\) — the case \(n=0\) giving \(0\), and negative \(n\) the same value as \(\abs{n}\) because \(\cos\) is even. Substituting,
which with Equation (A.789) is Equation (A.788). Both computations produce the same nonnegative series, so one side is infinite exactly when the other is.
∎Equation (A.788) converts a variational problem over maps of a two-dimensional domain into one over functions on a circle: by Proposition A.478 the area of a disc-type surface is at most its Dirichlet integral, and by Equation (A.788) the Dirichlet integral of the harmonic extension of a given boundary parametrisation is computable from that parametrisation alone. The harmonic extension is moreover the minimiser of \(D\) among all extensions of \(\vect{g}\), which is the Dirichlet principle of Definition 16.67 for the disc; so the boundary functional \(\mathcal{D}\) is exactly the value of the two-dimensional problem, and Douglas' competitors are parametrisations, not surfaces.
Let \(\vect{g}_{k}\longrightarrow\vect{g}\) pointwise almost everywhere on \(S^{1}\). Then \(\mathcal{D}\left[\vect{g}\right] \le\liminf_{k}\mathcal{D}\left[\vect{g}_{k}\right]\). Rests on Definition A.481.
Derives Lemma A.484. The integrand of Equation (A.786) is nonnegative, and for almost every pair \(\left(\theta,\varphi\right)\) it converges to the integrand built from \(\vect{g}\), the denominator being independent of \(k\) and the numerator continuous in the values of the map. Fatou's lemma — quoted, as throughout this appendix (Remark A.459) — applied on the square \(\left[0,2\pi\right]^{2}\) gives the claim.
∎The conformal group and the three-point normalisation
For \(\alpha\in B\) and \(\vartheta\in\R\) define
Each \(M_{\alpha,\vartheta}\) is a holomorphic bijection of \(B\) onto itself, extending to a homeomorphism of \(\overline{B}\) that maps \(S^{1}\) onto \(S^{1}\); these maps form a group under composition; the group is not compact; and it acts simply transitively on the ordered triples of distinct points of \(S^{1}\) that are in positive cyclic order. The first three assertions are proved below; the fourth is quoted (Remark A.495). Rests on Lemma A.479.
Derives Lemma A.485. The circle is preserved. For \(\abs{z}=1\), \(\overline{z}=z^{-1}\), so
whence \(\abs{M_{\alpha,\vartheta}(z)}=1\). The denominator never vanishes on \(\overline{B}\), since \(\abs{\overline{\alpha}z} \le\abs{\alpha}<1\), so \(M_{\alpha,\vartheta}\) is holomorphic on a neighbourhood of \(\overline{B}\) and continuous on \(\overline{B}\).
The disc is preserved, and the map is bijective. A direct computation gives, for \(\abs{z}<1\),
the middle equality by expanding both squared moduli, so \(M_{\alpha,\vartheta}\) maps \(B\) into \(B\). Its inverse is \(M_{\beta,\eta}\) with \(\beta=-\ee^{-\ii\vartheta}\alpha\) suitably normalised — explicitly, solving \(w=\ee^{\ii\vartheta}\left(z-\alpha\right)/ \left(1-\overline{\alpha}z\right)\) for \(z\) gives \(z=\left(\ee^{-\ii\vartheta}w+\alpha\right)/ \left(1+\overline{\alpha}\ee^{-\ii\vartheta}w\right)\), which is again of the form Equation (A.792); so each \(M_{\alpha,\vartheta}\) is a bijection of \(B\) and of \(S^{1}\), and the family is closed under inversion. It is closed under composition as well, since a composition of two quotients of linear polynomials is again one and it maps \(B\) bijectively onto \(B\) — that such a map must again be of the form Equation (A.792) is the classification quoted in the fourth assertion.
Non-compactness. Take \(\vartheta=0\) and \(\alpha=1-1/k\) along the real axis. For any fixed \(z\in S^{1}\) with \(z\ne1\),
whereas \(M_{\alpha,0}(1)=1\) for every \(k\). The sequence therefore has no subsequence converging uniformly on \(S^{1}\) to a homeomorphism: its pointwise limit collapses the whole circle except one point onto the single value \(-1\). This is precisely the degeneration a minimising sequence for \(\mathcal{D}\) may undergo at no cost, since by Lemma A.479 and Equation (A.788) the value of \(\mathcal{D}\) is unchanged by such a reparametrisation of the boundary.
∎Fix a homeomorphism \(\gamma:S^{1}\longrightarrow\Gamma\) and three points \(\vect{Q}_{1},\vect{Q}_{2},\vect{Q}_{3}\) on \(\Gamma\), in positive cyclic order. A monotone parametrisation of \(\Gamma\) is a map \(\vect{g}=\gamma\circ\psi\), where \(\psi:\left[0,2\pi\right]\longrightarrow\R\) is continuous and nondecreasing with \(\psi\left(2\pi\right)=\psi(0)+2\pi\), angles being read modulo \(2\pi\). The normalised class \(\mathcal{C}^{\ast}\left(\Gamma\right)\) consists of those monotone parametrisations with
Rests on Lemma A.485 and Definition A.481.
\(\inf\set{\mathcal{D}\left[\vect{g}\right]\mid \vect{g}\in\mathcal{C}^{\ast}\left(\Gamma\right)} =\inf\set{\mathcal{D}\left[\vect{g}\right]\mid \vect{g}\text{ a monotone parametrisation of }\Gamma}\). Rests on Lemmas A.479 and A.485.
Derives Lemma A.487. The class \(\mathcal{C}^{\ast}\) is contained in the larger class, so the left infimum is at least the right one. Conversely let \(\vect{g}\) be any monotone parametrisation. The three points \(\vect{g}^{-1}\left(\vect{Q}_{i}\right)\) may be chosen as three points \(\zeta_{1},\zeta_{2},\zeta_{3}\) of \(S^{1}\) in positive cyclic order, because \(\vect{g}\) is monotone of degree one; by the simple transitivity in Lemma A.485 there is a conformal automorphism \(M\) of \(B\) carrying \(\left(1,\ii,-1\right)\) to \(\left(\zeta_{1},\zeta_{2},\zeta_{3}\right)\). Then \(\vect{g}\circ M\) is again a monotone parametrisation, now satisfying Equation (A.793), and its harmonic extension is \(\vect{h}\circ M\) — the composition of a harmonic map with a holomorphic one is harmonic, since \(\nabla^{2}\left(\vect{h}\circ M\right) =\abs{M'}^{2}\left(\nabla^{2}\vect{h}\right)\circ M\) by the computation of Lemma A.479 applied to second derivatives, and its boundary values are \(\vect{g}\circ M\), so it is the harmonic extension of \(\vect{g}\circ M\) by the uniqueness of the solution of the Dirichlet problem (Corollary 10.76). Hence, by Equation (A.788) and Lemma A.479,
so every value attained on the larger class is attained on \(\mathcal{C}^{\ast}\).
∎Compactness: the Courant–Lebesgue estimate
Let \(\Gamma\) be a Jordan curve with homeomorphic parametrisation \(\gamma:S^{1}\longrightarrow\Gamma\). For every \(\varepsilon>0\) there is \(\delta>0\) such that any two points \(\vect{P},\vect{P}'\) of \(\Gamma\) with \(\abs{\vect{P}-\vect{P}'}<\delta\) divide \(\Gamma\) into two arcs, at least one of which has diameter less than \(\varepsilon\). Rests on Theorems 6.31 and 7.25.
Derives Lemma A.488. By definition a Jordan curve is the image of a homeomorphism \(\gamma\) of \(S^{1}\), so \(\gamma^{-1}\) exists and is continuous, and \(\Gamma\) is compact, being a continuous image of a compact set.
Both maps are uniformly continuous. For \(\gamma\) this is Theorem 7.25 applied componentwise to the angle parametrisation on \(\left[0,2\pi\right]\). For \(\gamma^{-1}\), suppose not: there are \(\varepsilon_{0}>0\) and points \(\vect{P}_{k},\vect{P}'_{k}\in\Gamma\) with \(\abs{\vect{P}_{k}-\vect{P}'_{k}}\longrightarrow0\) and \(\abs{\gamma^{-1}\left(\vect{P}_{k}\right) -\gamma^{-1}\left(\vect{P}'_{k}\right)}\ge\varepsilon_{0}\). The circle is compact, hence sequentially compact (Theorem 6.31), so along a subsequence \(\gamma^{-1}\left(\vect{P}_{k}\right)\longrightarrow\zeta\) and \(\gamma^{-1}\left(\vect{P}'_{k}\right)\longrightarrow\zeta'\) with \(\abs{\zeta-\zeta'}\ge\varepsilon_{0}\), so \(\zeta\ne\zeta'\). Applying the continuous \(\gamma\), \(\vect{P}_{k}\longrightarrow\gamma\left(\zeta\right)\) and \(\vect{P}'_{k}\longrightarrow\gamma\left(\zeta'\right)\), and \(\abs{\vect{P}_{k}-\vect{P}'_{k}}\longrightarrow0\) then forces \(\gamma\left(\zeta\right)=\gamma\left(\zeta'\right)\), contradicting injectivity.
Given \(\varepsilon>0\), uniform continuity of \(\gamma\) supplies \(\eta>0\) such that arcs of \(S^{1}\) of length below \(\eta\) have images of diameter below \(\varepsilon\); uniform continuity of \(\gamma^{-1}\) then supplies \(\delta>0\) such that \(\abs{\vect{P}-\vect{P}'}<\delta\) forces \(\gamma^{-1}\left(\vect{P}\right)\) and \(\gamma^{-1}\left(\vect{P}'\right)\) to be at distance below \(\eta/\left(2\pi\right)\) along \(S^{1}\), so that the shorter of the two arcs they determine has length below \(\eta\). Its image is one of the two arcs of \(\Gamma\), of diameter below \(\varepsilon\).
∎Let \(\vect{h}\) be harmonic on \(B\) and continuous on \(\overline{B}\) with \(D\left[\vect{h}\right]\le M\), let \(z_{0}\in S^{1}\), and let \(0<\delta<1\). Write \(C_{\varsigma}=\set{z\in B\mid \abs{z-z_{0}}=\varsigma}\). Then there is a radius \(\varsigma\in\left(\delta,\sqrt{\delta}\right)\) for which the image \(\vect{h}\left(C_{\varsigma}\right)\) has diameter at most
Rests on Definition A.476 and Lemma A.452.
Derives Lemma A.489. Use polar coordinates \(\left(\varsigma,\omega\right)\) centred at \(z_{0}\), in which \(\dd u\,\dd v=\varsigma\,\dd\varsigma\,\dd\omega\) and \(\abs{\nabla\vect{h}}^{2} =\abs{\pp_{\varsigma}\vect{h}}^{2} +\varsigma^{-2}\abs{\pp_{\omega}\vect{h}}^{2}\). Writing \(I_{\varsigma}\) for the set of angles \(\omega\) with \(z_{0}+\varsigma\ee^{\ii\omega}\in B\) and discarding the radial term,
Since \(\int_{\delta}^{\sqrt{\delta}}\dd\varsigma/\varsigma =\tfrac12\log\left(1/\delta\right)\), the mean value of \(\zeta\) against the measure \(\dd\varsigma/\varsigma\) on that range is at most \(4M/\log\left(1/\delta\right)\), so there exists \(\varsigma\in\left(\delta,\sqrt{\delta}\right)\) with \(\zeta\left(\varsigma\right)\le4M/\log\left(1/\delta\right)\) — otherwise the integral would exceed \(2M\). For that radius the length of the image curve is bounded by the Cauchy–Schwarz inequality Equation (A.750),
which is Equation (A.794), and the diameter of a curve is at most its length.
∎Let \(\left(\vect{g}_{k}\right)\subset \mathcal{C}^{\ast}\left(\Gamma\right)\) satisfy \(\mathcal{D}\left[\vect{g}_{k}\right]\le M\) for every \(k\). Then the family is equicontinuous on \(S^{1}\), and a subsequence converges uniformly to some \(\vect{g}\in\mathcal{C}^{\ast}\left(\Gamma\right)\). Rests on Lemmas A.454, A.488 and A.489.
Derives Theorem A.490. Equicontinuity. Let \(\varepsilon>0\) and take \(\delta_{0}\) from Lemma A.488 for this \(\varepsilon\); shrink \(\varepsilon\) if necessary so that \(2\varepsilon\) is smaller than the mutual distances of \(\vect{Q}_{1},\vect{Q}_{2},\vect{Q}_{3}\) and of the arcs between them. Choose \(\delta\in(0,1)\) with \(\left(4\pi M/\log\left(1/\delta\right)\right)^{1/2}<\delta_{0}\) and \(\sqrt{\delta}\) smaller than one third of the least of those arc lengths. Let \(\vect{h}_{k}\) be the harmonic extension of \(\vect{g}_{k}\), which by Equation (A.788) has \(D\left[\vect{h}_{k}\right]=\mathcal{D}\left[\vect{g}_{k}\right] \le M\), and which is continuous on \(\overline{B}\) with boundary values \(\vect{g}_{k}\) by the quoted solvability of the Dirichlet problem (Theorem A.491). For \(z_{0}\in S^{1}\), Lemma A.489 produces a radius \(\varsigma_{k}\in\left(\delta,\sqrt{\delta}\right)\) for which the arc \(C_{\varsigma_{k}}\) has image of diameter below \(\delta_{0}\); its two endpoints lie on \(S^{1}\) and their images \(\vect{P},\vect{P}'\) are values of \(\vect{g}_{k}\) at distance below \(\delta_{0}\). By Lemma A.488 one of the two arcs of \(\Gamma\) they cut off has diameter below \(\varepsilon\); it is the one whose preimage under the monotone map \(\vect{g}_{k}\) is the short boundary arc cut off by \(C_{\varsigma_{k}}\), because the other boundary arc has length above \(2\pi-2\sqrt{\delta}\) and its image therefore contains at least two of \(\vect{Q}_{1},\vect{Q}_{2},\vect{Q}_{3}\) — which by the choice of \(\varepsilon\) cannot lie in a set of diameter \(\varepsilon\). Hence \(\vect{g}_{k}\) oscillates by less than \(\varepsilon\) on the boundary arc cut off at \(z_{0}\), which contains all points of \(S^{1}\) within \(\delta\) of \(z_{0}\); and \(\delta\) was chosen independently of \(k\) and of \(z_{0}\). That is equicontinuity.
Extraction of a limit. Write \(\vect{g}_{k}=\gamma\circ\psi_{k}\) as in Definition A.486. Equicontinuity of \(\vect{g}_{k}\) transfers to \(\psi_{k}\) through the uniform continuity of \(\gamma^{-1}\) (Lemma A.488 and its proof), and the \(\psi_{k}\) are uniformly bounded once normalised by \(\psi_{k}(0)\in\left[0,2\pi\right)\). The Arzelà–Ascoli theorem Lemma A.454 gives a uniformly convergent subsequence \(\psi_{k}\longrightarrow\psi\); the limit is continuous, nondecreasing and satisfies \(\psi\left(2\pi\right)=\psi(0)+2\pi\), these being closed conditions under uniform convergence. Hence \(\vect{g}=\gamma\circ\psi\) is again a monotone parametrisation, it satisfies Equation (A.793) because each \(\vect{g}_{k}\) does and convergence is pointwise, and \(\vect{g}_{k}\longrightarrow\vect{g}\) uniformly by the uniform continuity of \(\gamma\).
∎The two quoted inputs, and the conclusion
For every continuous \(\vect{g}:S^{1}\longrightarrow\R^{3}\) the Poisson integral
is harmonic on \(B\), continuous on \(\overline{B}\), and equal to \(\vect{g}\) on \(S^{1}\); it coincides with the series Equation (A.787), and it minimises \(D\) among all disc-type surfaces with boundary values \(\vect{g}\). This treatise does not prove it: Partial Differential Equations develops the maximum principles and the fundamental solution but not the Poisson kernel of the disc, and Complex Analysis explicitly reserves conformal mapping. See [Courant:1962]. Rests on Definitions 10.66 and 16.67.
Let \(\vect{g}_{\ast}\) minimise \(\mathcal{D}\) over \(\mathcal{C}^{\ast}\left(\Gamma\right)\) and let \(\vect{X}=\vect{h}_{\vect{g}_{\ast}}\) be its harmonic extension. Then \(\vect{X}\) is conformal on \(B\), that is \(E=G\) and \(F=0\), except possibly at isolated branch points where \(\vect{X}_{u}=\vect{X}_{v}=\vect{0}\). Equivalently, the Hopf differential \(\Phi=\abs{\vect{X}_{u}}^{2}-\abs{\vect{X}_{v}}^{2} -2\ii\,\vect{X}_{u}\cdot\vect{X}_{v}\) vanishes identically. The proof is by inner variations — reparametrising the disc rather than moving the surface, which is legitimate precisely because \(\mathcal{C}^{\ast}\) is a full gauge slice by Lemma A.487 — and it is not reproduced here [Douglas:1931]. Rests on Definition A.481 and Lemma A.487.
Let \(\vect{X}\) be harmonic and conformal on \(B\), with \(E=G=\Lambda^{2}>0\) and \(F=0\) at a point. Then the mean curvature \(H=\tfrac12 g^{ij}b_{ij}\) of the surface vanishes there. Conversely a conformal parametrisation of a surface of vanishing mean curvature is harmonic. Rests on Definition 13.25 and Proposition 16.8.
Derives Proposition A.493. Work at a point where \(E=G=\Lambda^{2}>0\) and \(F=0\), with unit normal \(\hat n\) (Equation (13.80)). Differentiating the conformality relations,
using \(F\equiv0\) and \(G\equiv E\). Adding, \(\left(\vect{X}_{uu}+\vect{X}_{vv}\right)\cdot\vect{X}_{u}=0\), and the same computation with \(u\) and \(v\) exchanged gives \(\left(\vect{X}_{uu}+\vect{X}_{vv}\right)\cdot\vect{X}_{v}=0\). So the Laplacian \(\nabla^{2}\vect{X}=\vect{X}_{uu}+\vect{X}_{vv}\) is orthogonal to both tangent vectors, hence parallel to \(\hat n\), and by Equation (13.80) its normal component is
Since \(F=0\) and \(E=G=\Lambda^{2}\), the inverse metric is \(g^{ij}=\Lambda^{-2}\delta^{ij}\) (Definition 13.23), so \(2H=g^{ij}b_{ij}=\Lambda^{-2}\left(b_{11}+b_{22}\right)\) and therefore
If \(\vect{X}\) is harmonic the left side vanishes and, \(\Lambda^{2}\) being positive, \(H=0\); conversely \(H=0\) makes the right side vanish. By Proposition 16.8 the vanishing of the mean curvature is the minimal-surface equation Equation (16.7) wherever the surface is a graph, so the image is a minimal surface in the sense of Section 16.1.3.
∎Proof of Theorem A.477, granting the two quoted theorems. Derives Theorem A.477. Let \(m=\inf\set{\mathcal{D}\left[\vect{g}\right]\mid \vect{g}\in\mathcal{C}^{\ast}\left(\Gamma\right)}\). It is finite: \(\Gamma\) is rectifiable, and a parametrisation proportional to arc length has \(\abs{\vect{g}(\theta)-\vect{g}(\varphi)}\le \ell\,\abs{\theta-\varphi}/\left(2\pi\right)\) with \(\ell\) the length of \(\Gamma\), so the integrand of Equation (A.786) is bounded by \(\ell^{2}\abs{\theta-\varphi}^{2}/ \left(4\pi^{2}\sin^{2}\left(\left(\theta-\varphi\right)/2\right)\right)\), which is bounded on the square, and \(\mathcal{D}<\infty\) there; this is the only place where rectifiability is used, and it is used exactly to make the infimum finite.
Take a minimising sequence \(\left(\vect{g}_{k}\right)\subset\mathcal{C}^{\ast}\), which by discarding finitely many terms satisfies \(\mathcal{D}\left[\vect{g}_{k}\right]\le m+1\). By Theorem A.490 a subsequence converges uniformly to some \(\vect{g}_{\ast}\in\mathcal{C}^{\ast}\left(\Gamma\right)\), and by lower semicontinuity Lemma A.484, \(\mathcal{D}\left[\vect{g}_{\ast}\right]\le \liminf_{k}\mathcal{D}\left[\vect{g}_{k}\right]=m\); the reverse inequality holds because \(\vect{g}_{\ast}\in\mathcal{C}^{\ast}\). So the infimum is attained.
Let \(\vect{X}\) be the harmonic extension of \(\vect{g}_{\ast}\), which exists and is continuous on \(\overline{B}\) by Theorem A.491. By Theorem A.492 it is conformal, so by Proposition A.478 \(A\left[\vect{X}\right]=D\left[\vect{X}\right] =\mathcal{D}\left[\vect{g}_{\ast}\right]=m\), using Equation (A.788) for the last equality.
It remains to see that no disc-type surface spanning \(\Gamma\) has smaller area. Let \(\vect{Y}\) be one, with boundary parametrisation \(\vect{g}\), which by Lemma A.487 may be taken in \(\mathcal{C}^{\ast}\left(\Gamma\right)\). The area is reparametrisation invariant, since \(\sqrt{EG-F^{2}}\) transforms with the Jacobian under any \(C^{1}\) change of variables, and by Proposition A.478 \(A\left[\vect{Y}\right]=A\left[\vect{Y}\circ\tau\right] \le D\left[\vect{Y}\circ\tau\right]\) for every such \(\tau\); the quoted approximation step — item (4) of Remark A.495 — upgrades this to the equality
the infimum being over the reparametrisations of the disc. Each \(\vect{Y}\circ\tau\) has the same boundary values as \(\vect{Y}\) up to a reparametrisation of \(S^{1}\), and by the Dirichlet principle of Theorem A.491 its Dirichlet integral is at least that of the harmonic extension of those boundary values, which by Equation (A.788) and Lemma A.487 is \(\mathcal{D}\left[\vect{g}\right]\ge m\). Hence \(A\left[\vect{Y}\right]\ge m=A\left[\vect{X}\right]\): the surface \(\vect{X}\) minimises area among disc-type surfaces spanning \(\Gamma\). Finally Proposition A.493 makes the image a minimal surface away from branch points.
∎Radó's independent proof [Rado:1930], published a year before Douglas', never forms a boundary functional. It approximates \(\Gamma\) by inscribed polygons \(\Gamma_{k}\); for a polygon the minimal surface is produced by solving finitely many Dirichlet problems on the disc and controlling the resulting sequence with the maximum principle, and the areas \(A_{k}\) converge to the infimum for \(\Gamma\). The compactness that Douglas buys with the three-point normalisation is bought here by a condition on the curve — Radó first assumes \(\Gamma\) bounds some surface of finite area, which for a rectifiable curve is automatic — and the limit surface is obtained from the \(\Gamma_{k}\) by a normal-family argument in the function-theoretic sense. The two proofs meet at the same place: both must produce a conformal parametrisation, and neither can avoid the Dirichlet principle for the disc. What Douglas gains is that his functional is intrinsic to the boundary, which is why it, and not Radó's construction, generalises to higher connectivity and higher genus.
This section does not prove Theorem 16.81. It reduces it to named results, and the reduction is only worth what the honesty of this list is worth.
-
The Dirichlet problem for the disc (Theorem A.491): that the Poisson integral of a continuous boundary function is harmonic inside, continuous up to the boundary, and minimises the Dirichlet integral among all extensions. Used twice — in Theorem A.490 to have a continuous harmonic extension to apply Courant–Lebesgue to, and in the final proof. Not proved anywhere in this treatise.
-
Douglas' conformality theorem (Theorem A.492): that a minimiser of \(\mathcal{D}\) over the normalised class has a conformal harmonic extension. This is the technical heart of the whole subject and the step for which Douglas was awarded one of the first two Fields Medals; it is stated precisely above and not proved [Douglas:1931].
-
Simple transitivity of the conformal group on boundary triples (the fourth assertion of Lemma A.485), which belongs to the conformal mapping theory that Complex Analysis explicitly reserves. Everything else in that lemma — that the maps Equation (A.792) are automorphisms and that the family is non-compact — is proved.
-
The comparison of \(A\) with \(D\) over reparametrisations used at the end of the final proof, i.e. that the infimum of \(D\) over all reparametrisations of a given disc-type surface equals its area. This is Douglas' and Radó's approximation argument and is quoted with them.
-
Lebesgue integration, as everywhere in this appendix: Fatou's lemma in Lemma A.484, and the polar- coordinate rewriting of a double integral in Lemma A.489.
Proved here in full, and used above: the pointwise inequality \(D\ge A\) with its equality case, the conformal invariance of \(D\), the identity \(\mathcal{D}=D\) of the harmonic extension with its constant, lower semicontinuity, the group property and non-compactness of the Möbius family, the arc lemma, the Courant–Lebesgue estimate, the equicontinuity of the normalised class and the extraction of a limit, and the identity Equation (A.796) that turns a harmonic conformal map into a minimal surface. The reader should also carry forward Remark 16.82 of the chapter, which is about what the theorem claims — disc type, minimisation among disc-type maps, possible branch points — as opposed to what this section proves about it.
This section accompanies Theorem 16.81 of Section 16.5.4, replacing the pending derivation there with a proof architecture and an explicit list of imports rather than with a complete derivation — the one place in this appendix where that is so, and it is flagged as such in Remark A.495. The minimal surfaces the theorem produces are the soap films of Section 16.1.3, whose equation Equation (16.7) was derived there for a graph and is recovered here in the conformal form Equation (A.796); the differential geometry is that of Differentiable Manifolds, Tensors, and Curvature, and the physics of the surface tension that realises the area as an energy belongs to Fluid Dynamics.
Change of Variables in a Multiple Integral
This appendix proves Theorem 7.97 of Real Analysis: that a \(C^{1}\) change of coordinates \(\vect{\Phi}\) transports a multiple integral with the modulus of its Jacobian determinant as the weight, Equation (7.113). It is the one property of the multiple integral that the physical parts use on almost every page — every passage to polar, cylindrical or spherical coordinates is an instance — and it is the only one Real Analysis states without proof.
The derivation is elementary throughout: no measure theory, no Lebesgue convergence theorem, no fixed-point theorem, and nothing from outside this treatise. Two ingredients that the chapters state only in a weaker form are proved here rather than quoted — the multiplicativity of the determinant, which Linear Algebra and Representation Theory does not record, and iterated integration in \(\R^{N}\), which Remark 7.96 states for \(N=2\) and \(N=3\). One ingredient is genuinely constructed here and was not foreseen in the plan for this section: a continuous partition of unity, built from distance functions in Lemma A.505. Without it the passage from a local statement to a global one cannot be made for a Riemann integral over a general region, and saying that no partition of unity is used would be false. It costs half a page and quotes nothing; see Remark A.515.
Throughout, points of \(\R^{N}\) are written \(\vect{u}=(u^{1},\ldots,u^{N})\), and when the last coordinate is singled out we write \(\vect{u}=(\vect{u}',t)\) with \(\vect{u}'\in\R^{N-1}\) and \(t\in\R\). A box is a product \([a_{1},b_{1}]\times\cdots\times[a_{N},b_{N}]\) and a cube is a box with all edges equal. The support of a function \(f\) on an open \(V\subseteq\R^{N}\) is the closure of \(\set{\vect{x}\mid f(\vect{x})\neq0}\); we write \(f\in C_{c}(V)\) when \(f\) is continuous and that closure is a compact subset of \(V\), in which case the extension of \(f\) by zero is continuous on the whole of \(\R^{N}\) and vanishes off a box, so \(\int_{\R^{N}}f\) means the integral of that extension over any box containing the support (Definition 7.93), and is independent of the box.
What the derivation uses
From Real Analysis: the Darboux construction of the multiple integral (Definition 7.93), the integrability of a continuous function on a box (Theorem 7.40), uniform continuity on a compact set (Theorem 7.25), the one-variable substitution rule Equation (7.27), the mean value theorem (Theorem 7.35), the chain rule (Proposition 7.72) and the inverse function theorem (Corollary 7.81). From Linear Algebra and Representation Theory: the Leibniz expansion Equation (5.19) and the two properties read off it there — linearity in each column and antisymmetry under exchange of two columns. From Topological and Metric Spaces: compactness (Definition 6.9), the Heine–Borel theorem in \(\R^{N}\) (Theorem 6.12) and the Lebesgue number lemma (Lemma 6.30).
For real \(N\times N\) matrices \(A\) and \(B\), \(\det(AB)=\det A\,\det B\). Rests on Equation (5.19).
Derives Lemma A.497. Write \((AB)^{i}{}_{j}=\sum_{k}A^{i}{}_{k}B^{k}{}_{j}\) and expand Equation (5.19), distributing the product over \(j\):
the interchange being a finite rearrangement. The inner sum is, by Equation (5.19) again, the determinant of the matrix whose \(j\)-th column is the \(k_{j}\)-th column of \(A\). If two of the \(k_{j}\) coincide that matrix has a repeated column and the determinant vanishes, so only the tuples with \(k_{j}=\tau(j)\) for a permutation \(\tau\) survive; and bringing the columns back into their natural order costs \(\sgn(\tau)\), by the antisymmetry recorded after Equation (5.19). Hence the inner sum equals \(\sgn(\tau)\det A\) and
Sets of zero content
The Darboux construction ignores sets that can be covered by finitely many boxes of arbitrarily small total volume. That notion — zero content — is strictly stronger than the measure zero of Section 7.12, which allows countably many covering sets; it is the one a Riemann sum can see, and it is the one used here.
A bounded set \(Z\subset\R^{N}\) has zero content if for every \(\varepsilon>0\) there are finitely many cubes of total volume less than \(\varepsilon\) whose union contains \(Z\). Rests on Definitions 6.9 and 7.93.
Covering by cubes rather than by boxes is no restriction: a box with edges \(a_{1},\ldots,a_{N}\) is covered by cubes of side \(h\) in number at most \(\prod_{i}\left(a_{i}/h+1\right)\), of total volume at most \(\prod_{i}\left(a_{i}+h\right)\), which tends to the volume of the box as \(h\to0\). A finite union of sets of zero content has zero content.
Let \(R\) be a box, \(Z\subset R\) of zero content, and \(f:R\to\R\) bounded with \(\abs{f}\le M\).
-
If \(f\) is continuous at every point of \(R\setminus Z\), then \(f\) is integrable on \(R\).
-
If \(f\) vanishes off \(Z\), then \(f\) is integrable with \(\int_{R}f=0\).
-
If \(E_{1},\ldots,E_{m}\subseteq R\) are such that each characteristic function \(\chi_{E_{i}}\) is integrable, the \(E_{i}\) cover a set \(E\) and the pairwise intersections \(E_{i}\cap E_{j}\) (\(i\neq j\)) have zero content, then \(\int_{E}f=\sum_{i}\int_{E_{i}}f\) for every \(f\) for which the left-hand side is defined.
Rests on Definition A.498, Theorem 7.25 and Definition 7.93.
Derives Lemma A.499. (1) Given \(\varepsilon>0\), cover \(Z\) by finitely many open cubes of total volume less than \(\varepsilon\) and let \(K\) be the complement in \(R\) of their union: \(K\) is closed and bounded, hence compact (Theorem 6.12). Note that \(f\) is continuous, as a function on \(R\), at every point of \(K\).
We claim there is \(\lambda>0\) such that any subset of \(R\) of diameter less than \(\lambda\) which meets \(K\) carries an oscillation of \(f\) of at most \(\varepsilon\). If not, then for each \(j\) there are \(\vect{x}_{j}\in K\) and \(\vect{y}_{j},\vect{z}_{j}\in R\) within \(1/j\) of \(\vect{x}_{j}\) with \(\abs{f(\vect{y}_{j})-f(\vect{z}_{j})}>\varepsilon\); by compactness a subsequence \(\vect{x}_{j}\) converges to some \(\vect{x}\in K\), the corresponding \(\vect{y}_{j}\) and \(\vect{z}_{j}\) converge to \(\vect{x}\) as well, and the continuity of \(f\) at \(\vect{x}\) makes both \(f(\vect{y}_{j})\) and \(f(\vect{z}_{j})\) tend to \(f(\vect{x})\) — a contradiction.
Now partition \(R\) with mesh smaller than \(\lambda/\sqrt{N}\). Writing \(A\) for the union of the covering cubes, every sub-box either lies inside \(A\) or meets \(K=R\setminus A\). Those of the first kind have total volume at most \(\operatorname{vol}(A)<\varepsilon\) and contribute at most \(2M\varepsilon\) to the gap between the upper and the lower sum; on those of the second kind the oscillation of \(f\) is at most \(\varepsilon\), contributing at most \(\varepsilon\operatorname{vol}(R)\). The gap can therefore be made as small as desired, and \(f\) is integrable (Definition 7.93).
(2) By (1), \(f\) is integrable, since it is continuous (with value zero) off \(Z\). Cover \(Z\) by finitely many cubes of total volume less than \(\varepsilon\) and partition \(R\) with a mesh small compared with the smallest of those cubes. Every sub-box on which \(f\) is not identically zero meets \(Z\), hence lies in the union of the cubes enlarged by one mesh on every side, whose volume is then at most \(2\varepsilon\); so \(\abs{\int_{R}f}\le2M\varepsilon\) for every \(\varepsilon>0\).
(3) The function \(\chi_{E}-\sum_{i}\chi_{E_{i}}\) vanishes off \(\bigcup_{i<j}\left(E_{i}\cap E_{j}\right)\), a set of zero content, and is bounded by \(m\); multiplying by \(f\) and applying (2) gives the statement.
∎-
Let \(B\subset\R^{N-1}\) be a box and \(g:B\to\R\) continuous. Then the graph \(\set{(\vect{u}',g(\vect{u}'))\mid \vect{u}'\in B}\) has zero content in \(\R^{N}\); so does its intersection with any coordinate hyperplane, and so does a face of a box.
-
Let \(W\subseteq\R^{N}\) be open, \(\vect{F}:W\to\R^{N}\) of class \(C^{1}\), and \(Z\subset W\) a compact set of zero content. Then \(\vect{F}(Z)\) has zero content.
Rests on Lemma A.499, Theorem 7.25 and Theorem 7.35.
Derives Lemma A.500. (1) Let \(\varepsilon>0\). By uniform continuity (Theorem 7.25) there is \(\delta>0\) such that \(g\) varies by less than \(\varepsilon\) on any subset of \(B\) of diameter less than \(\delta\). Cover \(B\) by cubes of side \(h<\min\set{\delta/\sqrt{N-1},\varepsilon}\), in number at most \(C_{B}h^{-(N-1)}\) with \(C_{B}\) depending on \(B\) alone. Over each such cube the graph lies in a box of base that cube and height \(\varepsilon\), which is covered by \(\lceil\varepsilon/h\rceil\) cubes of side \(h\), of total volume at most \(\left(\varepsilon+h\right)h^{N-1}\le2\varepsilon h^{N-1}\). Summing gives a total volume at most \(2C_{B}\varepsilon\). A face of a box is the graph of a constant function.
(2) Cover \(Z\) by finitely many closed balls contained in \(W\); on each, the entries of \(D\vect{F}\) are bounded, and the mean value theorem (Theorem 7.35) applied to each component along the segment joining two points of the ball — which stays in the ball, the ball being convex — gives a constant \(L\) with \(\abs{\vect{F}(\vect{x})-\vect{F}(\vect{y})}\le L\abs{\vect{x}-\vect{y}}\) there. Take \(L\) to be the largest of the finitely many constants. Now cover \(Z\) by cubes of side \(h\) and total volume less than \(\varepsilon\), refining \(h\) so small that each such cube meeting \(Z\) lies in one of the balls. A cube of side \(h\) has diameter \(\sqrt{N}h\), so its image has diameter at most \(L\sqrt{N}h\) and lies in a cube of that side, of volume \(\left(L\sqrt{N}\right)^{N}h^{N}\). The images therefore cover \(\vect{F}(Z)\) with total volume at most \(\left(L\sqrt{N}\right)^{N}\varepsilon\).
∎Iterated integration in $\R^{N}$
Let \(R=R'\times[a,b]\) be a box in \(\R^{N}=\R^{N-1}\times\R\) and let \(f:R\to\R\) be bounded and integrable, and such that \(t\longmapsto f(\vect{u}',t)\) is integrable on \([a,b]\) for every \(\vect{u}'\in R'\). Then
the inner integral being an integrable function of \(\vect{u}'\). The same holds with any one of the \(N\) coordinates in the role of \(t\). Rests on Definition 7.93, Definition 7.39 and Theorem 7.40.
Derives Lemma A.501. Write \(F(\vect{u}')=\int_{a}^{b}f(\vect{u}',t)\,\dd t\). Let \(P'\) be a partition of \(R'\) into sub-boxes \(S'\) and \(P_{t}\) a partition \(a=t_{0}<\cdots<t_{n}=b\); together they partition \(R\) into the sub-boxes \(S=S'\times[t_{i-1},t_{i}]\). Fix \(\vect{u}'\in S'\). For each \(i\) and every \(t\in[t_{i-1},t_{i}]\) one has \(f(\vect{u}',t)\ge\inf_{S}f\), so
the right-hand side being independent of \(\vect{u}'\in S'\); hence it is a lower bound for \(\inf_{S'}F\). Multiplying by \(\operatorname{vol}(S')\) and summing over \(S'\) gives \(L(F,P')\ge L(f,P)\), where \(L\) denotes the lower Darboux sum Equation (7.22). The mirror argument gives \(U(F,P')\le U(f,P)\). Since \(f\) is integrable, the outer sums squeeze: for every \(\varepsilon>0\) a partition \(P\) with \(U(f,P)-L(f,P)<\varepsilon\) yields \(U(F,P')-L(F,P')<\varepsilon\), so \(F\) is integrable on \(R'\), and both \(\int_{R'}F\) and \(\int_{R}f\) lie between \(L(f,P)\) and \(U(f,P)\); letting \(\varepsilon\to0\) makes them equal. Relabelling the coordinates puts any chosen one in the role of \(t\).
∎Let \(f\in C_{c}(\R^{N})\) and let \(\sigma\) be a permutation of \(\set{1,\ldots,N}\). Then \(\int_{\R^{N}}f\) equals the iterated integral taken in the order \(u^{\sigma(N)},u^{\sigma(N-1)},\ldots,u^{\sigma(1)}\) from innermost to outermost; consequently \(\int_{\R^{N}}f\circ\sigma=\int_{\R^{N}}f\), where \(\sigma\) also denotes the map \(\vect{u}\mapsto\left(u^{\sigma(1)},\ldots,u^{\sigma(N)}\right)\). Rests on Lemma A.501 and Theorem 7.40.
Derives Corollary A.502. Take a box \(R\) containing the support in its interior. A continuous function is integrable on \(R\) and on every slice (Theorem 7.40), so Lemma A.501 applies with any chosen coordinate innermost; the inner integral is again continuous, by the uniform continuity of \(f\) on the compact \(R\) (Theorem 7.25), so the lemma may be applied again to it, and \(N\) applications reduce \(\int_{R}f\) to the iterated integral in the chosen order. All \(N!\) orders therefore give the same number. Finally \(f\circ\sigma\) is continuous with compact support, and its iterated integral in the order \(u^{1},\ldots,u^{N}\) becomes, on renaming the integration variable of the \(j\)-th integration from \(u^{j}\) to \(u^{\sigma(j)}\), the iterated integral of \(f\) in the order \(u^{\sigma(1)},\ldots,u^{\sigma(N)}\), which is \(\int_{\R^{N}}f\).
∎Two reductions
Let \(U,V\subseteq\R^{N}\) be open and \(\vect{\Phi}:U\to V\) a \(C^{1}\) diffeomorphism. We say \(\vect{\Phi}\) has the substitution property if
both integrands being understood as extended by zero; the right-hand one belongs to \(C_{c}(U)\), its support being \(\vect{\Phi}^{-1}(\operatorname{supp}f)\). Rests on Corollary 7.81 and Definition 7.93.
If \(\vect{\Phi}:U\to V\) and \(\vect{\Psi}:V\to W\) are \(C^{1}\) diffeomorphisms with the substitution property, so is \(\vect{\Psi}\circ\vect{\Phi}\). Rests on Definition A.503, Proposition 7.72 and Lemma A.497.
Derives Lemma A.504. Let \(g\in C_{c}(W)\). Then \(f:=\left(g\circ\vect{\Psi}\right)\abs{\det D\vect{\Psi}}\) lies in \(C_{c}(V)\), and applying Equation (A.802) first to \(\vect{\Psi}\) and then to \(\vect{\Phi}\),
By the chain rule (Proposition 7.72), \(D\left(\vect{\Psi}\circ\vect{\Phi}\right)(\vect{u}) =D\vect{\Psi}\left(\vect{\Phi}(\vect{u})\right)D\vect{\Phi}(\vect{u})\), and Lemma A.497 turns the product of the two moduli into \(\abs{\det D\left(\vect{\Psi}\circ\vect{\Phi}\right)}\).
∎Let \(K\subset\R^{N}\) be compact and \(W_{1},\ldots,W_{m}\) open sets with \(K\subseteq W_{1}\cup\cdots\cup W_{m}\). There are continuous functions \(\lambda_{1},\ldots,\lambda_{m}\) on \(\R^{N}\) with \(0\le\lambda_{j}\le1\), with \(\operatorname{supp}\lambda_{j}\) compact and contained in \(W_{j}\), and with \(\sum_{j}\lambda_{j}=1\) on a neighbourhood of \(K\). Rests on Definitions 6.6, 6.9 and 6.24.
Derives Lemma A.505. For each \(\vect{x}\in K\) choose \(j(\vect{x})\) with \(\vect{x}\in W_{j(\vect{x})}\) and \(r(\vect{x})>0\) with the closed ball of radius \(2r(\vect{x})\) about \(\vect{x}\) inside \(W_{j(\vect{x})}\). The open balls of radius \(r(\vect{x})\) cover \(K\); extract a finite subcover (Definition 6.9) with centres \(\vect{x}_{1},\ldots,\vect{x}_{p}\) and let \(K_{j}\) be the union of the closed balls \(\overline{B}_{r(\vect{x}_{i})}(\vect{x}_{i})\) with \(j(\vect{x}_{i})=j\). Each \(K_{j}\) is compact and contained in \(W_{j}\), and \(K\subseteq\bigcup_{j}K_{j}\). Let \(\rho_{j}>0\) be smaller than the distance from \(K_{j}\) to the complement of \(W_{j}\) — positive because \(K_{j}\) is compact and the complement closed — and put
which is continuous, non-negative, strictly positive exactly on the open set \(\set{\operatorname{dist}(\cdot,K_{j})<\rho_{j}}\), and supported in its closure, a compact subset of \(W_{j}\). On the open set \(O=\set{\sum_{i}\psi_{i}>0}\), which contains \(K\), set \(\lambda_{j}=\psi_{j}/\sum_{i}\psi_{i}\); off \(O\) set \(\lambda_{j}=0\). Each \(\lambda_{j}\) is continuous on \(O\) and vanishes on a neighbourhood of every point of \(\R^{N}\setminus\operatorname{supp}\psi_{j}\), so it is continuous everywhere; its support lies in that of \(\psi_{j}\); and \(\sum_{j}\lambda_{j}=1\) on \(O\).
∎Let \(\vect{\Phi}:U\to V\) be a \(C^{1}\) diffeomorphism and suppose every \(\vect{p}\in U\) has an open neighbourhood \(N_{\vect{p}}\subseteq U\) such that the restriction \(\vect{\Phi}|_{N_{\vect{p}}}:N_{\vect{p}}\to\vect{\Phi}(N_{\vect{p}})\) has the substitution property. Then \(\vect{\Phi}\) has it. Rests on Lemma A.505, Definition A.503 and Definition 6.9.
Derives Lemma A.506. Let \(f\in C_{c}(V)\) with \(K=\operatorname{supp}f\). The set \(\vect{\Phi}^{-1}(K)\) is compact, so finitely many \(N_{\vect{p}_{1}},\ldots,N_{\vect{p}_{m}}\) cover it, and the open sets \(W_{j}=\vect{\Phi}(N_{\vect{p}_{j}})\subseteq V\) cover \(K\). Take \(\lambda_{1},\ldots,\lambda_{m}\) as in Lemma A.505. Each \(\lambda_{j}f\) is continuous, vanishes off \(\operatorname{supp}\lambda_{j}\cap K\) — a compact subset of \(W_{j}\) — and hence lies in \(C_{c}(W_{j})\); and \(\sum_{j}\lambda_{j}f=f\), since the two sides agree on the neighbourhood of \(K\) where \(\sum_{j}\lambda_{j}=1\) and both vanish off \(K\). Applying the substitution property of \(\vect{\Phi}|_{N_{\vect{p}_{j}}}\) to \(\lambda_{j}f\) and summing over \(j\),
the first and last steps by the linearity of the integral over a box containing every support involved.
∎Permutations and primitive maps
For a permutation \(\sigma\) of \(\set{1,\ldots,N}\), the linear map \(\vect{u}\longmapsto\left(u^{\sigma(1)},\ldots,u^{\sigma(N)}\right)\) of \(\R^{N}\) onto itself has the substitution property. Rests on Corollary A.502 and Equation (5.19).
Derives Lemma A.507. Its matrix has exactly one \(1\) in each row and column, so Equation (5.19) leaves a single term and the determinant is \(\sgn(\sigma)=\pm1\); hence \(\abs{\det}=1\) and Equation (A.802) reduces to \(\int f=\int f\circ\sigma\), which is Corollary A.502. This is the only step of the whole derivation at which the sign of the determinant is disposed of, and it is why Equation (7.113) carries a modulus while Equation (5.19) does not.
∎A \(C^{1}\) diffeomorphism \(\vect{\Phi}:U\to V\) of open subsets of \(\R^{N}\) is primitive if it changes only the last coordinate:
the determinant being that of a triangular matrix with \(N-1\) unit diagonal entries. Rests on Definition 7.66 and Equation (5.19).
Every primitive \(C^{1}\) diffeomorphism has the substitution property. Rests on Definition A.508, Equation (7.27) and Lemma A.501.
Derives Lemma A.509. Let \(f\in C_{c}(V)\) and write \(g=\left(f\circ\vect{\Phi}\right)\abs{\pp_{t}\Phi^{N}}\in C_{c}(U)\). Both \(f\) and \(g\), extended by zero, are continuous with compact support, so by Corollary A.502 each integral is the iterated integral with \(t\) innermost. It therefore suffices to prove, for every fixed \(\vect{u}'\in\R^{N-1}\),
Fix \(\vect{u}'\) and put \(U_{\vect{u}'}=\set{t\mid(\vect{u}',t)\in U}\) and \(V_{\vect{u}'}=\set{s\mid(\vect{u}',s)\in V}\), both open in \(\R\). Because \(\vect{\Phi}\) leaves \(\vect{u}'\) untouched and is a bijection of \(U\) onto \(V\), the map \(\varphi(t)=\Phi^{N}(\vect{u}',t)\) is a bijection of \(U_{\vect{u}'}\) onto \(V_{\vect{u}'}\); it is \(C^{1}\) with \(\varphi'=\pp_{t}\Phi^{N}\neq0\), hence strictly monotone on each connected component of \(U_{\vect{u}'}\), mapping that component onto a component of \(V_{\vect{u}'}\).
The integrand on the right of Equation (A.807) vanishes off the compact set \(\vect{\Phi}^{-1}(\operatorname{supp}f)\), whose slice at \(\vect{u}'\) is a compact subset \(T\) of \(U_{\vect{u}'}\). The components of \(U_{\vect{u}'}\) are open and disjoint and cover \(T\), so only finitely many of them, \(I_{1},\ldots,I_{q}\), meet \(T\); choose in each a compact interval \([\alpha_{r},\beta_{r}]\subset I_{r}\) containing \(T\cap I_{r}\), with \(\alpha_{r}<\beta_{r}\). On \(I_{r}\) the substitution rule Equation (7.27), applied to the \(C^{1}\) function \(\varphi\) and the continuous \(s\mapsto f(\vect{u}',s)\), gives
If \(\varphi\) increases on \(I_{r}\) then \(\varphi'=\abs{\varphi'}\) and the right-hand side is the integral over \(\left[\varphi(\alpha_{r}),\varphi(\beta_{r})\right]\); if \(\varphi\) decreases then \(\varphi'=-\abs{\varphi'}\) and the right-hand side is minus the integral over \(\left[\varphi(\beta_{r}),\varphi(\alpha_{r})\right]\). In both cases
Outside \(\left[\alpha_{r},\beta_{r}\right]\) the left integrand vanishes, so the left-hand side is the integral over the whole of \(I_{r}\); and \(f(\vect{u}',\cdot)\) vanishes on \(\varphi(I_{r})\setminus \varphi\left(\left[\alpha_{r},\beta_{r}\right]\right)\), since a point \(s\) there has \(\varphi^{-1}(s)\notin T\), so the right-hand side is the integral over the component \(\varphi(I_{r})\) of \(V_{\vect{u}'}\). Summing over \(r=1,\ldots,q\) and noting that \(f(\vect{u}',\cdot)\) vanishes on the components of \(V_{\vect{u}'}\) not among the \(\varphi(I_{r})\), and off \(V_{\vect{u}'}\) altogether, gives Equation (A.807).
∎The local factorisation
Let \(\vect{\Phi}\) be a \(C^{1}\) diffeomorphism of an open subset of \(\R^{N}\) onto an open subset, and let \(\vect{p}\) be a point of its domain. Then some neighbourhood of \(\vect{p}\) carries a factorisation of \(\vect{\Phi}\) as a composition of finitely many primitive maps and coordinate permutations. Rests on Corollary 7.81, Definition A.508 and Proposition 7.72.
Derives Lemma A.510. Say that a \(C^{1}\) diffeomorphism \(\vect{\Theta}\) defined near a point is \(m\)-adapted there, \(1\le m\le N+1\), if \(\Theta^{i}(\vect{u})=u^{i}\) for every \(i<m\) throughout a neighbourhood. Every diffeomorphism is \(1\)-adapted (the condition is empty), and an \((N+1)\)-adapted one is the identity. We show that an \(m\)-adapted \(\vect{\Theta}\) is, near the point, the composition of an \((m+1)\)-adapted diffeomorphism with one primitive map and one transposition of coordinates; \(N\) applications then prove the lemma, since a primitive map in the \(m\)-th coordinate is a transposition conjugate of one in the last coordinate.
Let \(\vect{\Theta}\) be \(m\)-adapted at \(\vect{q}\). Its Jacobian there has the block form
with \(\identity\) of size \(m-1\) and \(C\) the matrix \(\pp\Theta^{i}/\pp u^{j}\) with \(i,j\ge m\); expanding the determinant along the first \(m-1\) rows shows \(\det C=\det D\vect{\Theta}\neq0\), so the rows of \(C\) are independent and in particular the row \(\left(\pp\Theta^{m}/\pp u^{j}\right)_{j\ge m}\) is non-zero: there is \(j\ge m\) with \(\pp\Theta^{m}/\pp u^{j}(\vect{q})\neq0\).
Let \(\tau\) be the transposition of the coordinates \(m\) and \(j\), and set
a map changing only the \(m\)-th coordinate. At \(\vect{q}'=\tau(\vect{q})\) its \(m\)-th partial derivative in the \(m\)-th coordinate is \(\pp\Theta^{m}/\pp u^{j}(\vect{q})\neq0\), so \(\det D\vect{\varsigma}(\vect{q}')\neq0\) and, by the inverse function theorem (Corollary 7.81), \(\vect{\varsigma}\) is a \(C^{1}\) diffeomorphism of a neighbourhood of \(\vect{q}'\) onto an open set; it is a primitive map in the \(m\)-th coordinate.
Put \(\vect{\Xi}=\vect{\Theta}\circ\tau\circ\vect{\varsigma}^{-1}\), so that \(\vect{\Theta}=\vect{\Xi}\circ\vect{\varsigma}\circ\tau\) near \(\vect{q}\) (\(\tau\) being its own inverse). For \(i<m\), neither \(\tau\) nor \(\vect{\varsigma}\) — hence nor \(\vect{\varsigma}^{-1}\) — touches the \(i\)-th coordinate, and \(\vect{\Theta}\) preserves it, so \(\Xi^{i}(\vect{v})=v^{i}\). For \(i=m\): writing \(\vect{u}=\vect{\varsigma}^{-1}(\vect{v})\) one has \(v^{m}=\varsigma^{m}(\vect{u})=\Theta^{m}\left(\tau(\vect{u})\right) =\Xi^{m}(\vect{v})\). Hence \(\vect{\Xi}\) is \((m+1)\)-adapted, as claimed.
∎Every \(C^{1}\) diffeomorphism \(\vect{\Phi}:U\to V\) of open subsets of \(\R^{N}\) has the substitution property: for every \(f\in C_{c}(V)\),
Derives Proposition A.511. By Lemma A.510 each point of \(U\) has a neighbourhood on which \(\vect{\Phi}\) is a composition of primitive maps and permutations; each factor has the substitution property by Lemmas A.507 and A.509, so the composition has it by Lemma A.504. Thus \(\vect{\Phi}\) has the property locally, and Lemma A.506 globalises it.
∎From compact support to the theorem
Proposition A.511 carries no hypothesis on the shape of \(U\) and \(V\). To reach Theorem 7.97 as stated — an integrand merely continuous and bounded on \(V\), not vanishing near the boundary — something must be assumed about that boundary, and about the behaviour of \(\vect{\Phi}\) there; otherwise neither integral need exist. The hypotheses below hold for every use made of the formula in this treatise.
Let \(V\subset\R^{N}\) be open and bounded and suppose its boundary is contained in finitely many graphs of continuous functions over boxes, in the sense of Lemma A.500(1) with the coordinates permuted — the \(N\)-dimensional reading of Definition 7.95. Define the boundary strip
Then the content of \(V_{\delta}\) tends to zero as \(\delta\to0\). Rests on Lemma A.500, Theorem 7.25 and Definition 7.95.
Derives Lemma A.512. It is enough to treat one graph \(\Gamma\), of a continuous \(g\) over a box \(B\subset\R^{N-1}\), and to bound the content of its \(\delta\)-neighbourhood. By uniform continuity (Theorem 7.25) split \(B\) into cubes of side \(h\) on each of which \(g\) varies by less than \(\delta\). A point within distance \(\delta\) of \(\Gamma\) and lying over such a cube has last coordinate within \(2\delta\) of the value of \(g\) at the cube's centre, and lies over the cube enlarged by \(\delta\); so the neighbourhood is covered by boxes of total volume at most \(4\delta\left(\operatorname{vol}(B)+C\delta\right)\) with \(C\) depending on \(B\) alone. Letting first \(h\to0\) and then \(\delta\to0\) gives the claim; the finitely many graphs are summed.
∎Let \(U,V\subseteq\R^{N}\) be open and bounded and let \(\vect{\Phi}:U\to V\) be a bijection of class \(C^{1}\) with \(\det D\vect{\Phi}\neq0\) everywhere. Assume in addition that \(\vect{\Phi}\) is the restriction of a \(C^{1}\) diffeomorphism of an open set containing \(\overline{U}\), and that \(\pp U\) and \(\pp V\) are as in Lemma A.512. Then for every bounded continuous \(f\) on \(V\) both integrals in Equation (7.113) exist and are equal. Rests on Proposition A.511, Lemma A.512 and Lemma A.499.
Derives Theorem A.513. Existence. Extend \(f\) by zero to a box \(R\supseteq\overline{V}\). The extension is bounded and continuous off \(\pp V\), which has zero content by Lemma A.500(1); so it is integrable by Lemma A.499(1). The same argument applies on the \(U\) side: \(\left(f\circ\vect{\Phi}\right) \abs{\det D\vect{\Phi}}\) is bounded there, because \(D\vect{\Phi}\) extends continuously to the compact \(\overline{U}\), and continuous off \(\pp U\).
Equality. Let \(M\) bound \(\abs{f}\) and \(J\) bound \(\abs{\det D\vect{\Phi}}\) on \(\overline{U}\). For \(\delta>0\) put
a continuous function, zero where the distance is at most \(\delta\) and one where it is at least \(2\delta\). The set where \(\eta_{\delta}\neq0\) has closure inside \(\set{\operatorname{dist}(\cdot,\R^{N}\setminus V)\ge\delta}\), which is closed and bounded, hence a compact subset of \(V\); so \(f_{\delta}=\eta_{\delta}f\) lies in \(C_{c}(V)\) and Proposition A.511 gives
On the left, \(f-f_{\delta}\) vanishes outside \(V_{2\delta}\) and is bounded by \(M\), so the difference of the two integrals is at most \(M\) times the content of \(V_{2\delta}\), which tends to zero by Lemma A.512. On the right, the integrands differ only on \(\vect{\Phi}^{-1}(V_{2\delta})\) and by at most \(MJ\). The extension of \(\vect{\Phi}^{-1}\) is \(C^{1}\) on a neighbourhood of the compact \(\overline{V}\), hence Lipschitz there with some constant \(L\) by the argument of Lemma A.500(2); covering \(V_{2\delta}\) by cubes of total volume \(\varepsilon\) and taking images covers \(\vect{\Phi}^{-1}(V_{2\delta})\) with total volume at most \(\left(L\sqrt{N}\right)^{N}\varepsilon\), so its content also tends to zero. Letting \(\delta\to0\) in Equation (A.815) gives Equation (7.113).
∎Let \(U,V\) be open and bounded, let \(Z\subset U\) be closed in \(U\) and contained in finitely many graphs of continuous functions over boxes, and let \(\vect{\Phi}:U\to V\) be \(C^{1}\), mapping \(U\setminus Z\) bijectively onto \(V\setminus\vect{\Phi}(Z)\) with non-vanishing Jacobian there. If \(\vect{\Phi}\) extends \(C^{1}\) to a neighbourhood of \(\overline{U}\), then Equation (7.113) holds for every \(f\in C_{c}(V)\), and for bounded continuous \(f\) under the boundary hypotheses of Theorem A.513 applied to \(U\setminus Z\) and \(V\setminus\vect{\Phi}(Z)\). Rests on Theorem A.513, Lemma A.500 and Lemma A.499.
Derives Corollary A.514. \(U\setminus Z\) is open, and \(\vect{\Phi}|_{U\setminus Z}\) is a \(C^{1}\) diffeomorphism onto the open set \(V\setminus\vect{\Phi}(Z)\); note that \(\vect{\Phi}(Z\cap K)\) has zero content for every compact \(K\subset\overline{U}\), by Lemma A.500. Apply Proposition A.511, or Theorem A.513, to that restriction. Both integrals change by nothing when the sets \(Z\) and \(\vect{\Phi}(Z)\) are restored, because the integrands are bounded and the two sets have zero content (Lemma A.499(2)).
∎The corollary is what licenses the two Jacobians of Example 7.98. The polar map \(\vect{\Phi}(\rho,\varphi)=(\rho\cos\varphi,\rho\sin\varphi)\) of Equation (7.114) has \(\det D\vect{\Phi}=\rho\), which vanishes on the segment \(\rho=0\), and it fails to be injective unless \(\varphi\) is confined to an interval of length \(2\pi\), so that a ray must be removed from the image; segment and ray are each contained in the graph of a continuous function, and neither contributes to either integral. The spherical map is degenerate on the polar axis in the same way.
Reading the result
Nothing outside this treatise, and no result of Real Analysis beyond the ones listed in What the derivation uses. Three points are worth recording, because each is a place where the chapters state slightly less than the derivation needs and the gap is filled here rather than passed over.
First, Linear Algebra and Representation Theory records the Leibniz expansion Equation (5.19) and reads off column linearity and antisymmetry, but never states that the determinant is multiplicative; that is Lemma A.497, proved above from the expansion alone. It is a debt of Part II, small but real: the composition rule for Jacobians is used wherever coordinates are changed twice.
Second, Remark 7.96 states the reduction of a multiple integral to an iterated one for \(N=2\) and \(N=3\) only. The \(N\)-dimensional statement is Lemma A.501, proved here by the same Darboux squeeze. This too is a Part II debt, and it is the reason Theorem 7.97 could not simply be quoted from the chapter's own machinery.
Third — and this corrects the plan under which this section was written — the passage from a local identity to a global one does use a partition of unity. It is a continuous one, constructed in Lemma A.505 out of distance functions and nothing else; no smooth bump function is needed, because \(f\) is only assumed continuous. There is no fixed-point theorem, no convergence theorem of Lebesgue integration and no measure theory anywhere in the section: every set that is discarded is discarded because it has zero content (Definition A.498), which a Darboux sum can see, and not merely measure zero in the sense of Section 7.12.
Proposition A.511, the compactly supported form, is the theorem proper: it holds for arbitrary open \(U\) and \(V\) and is where all the work is. The extra hypotheses of Theorem A.513 — a boundary made of finitely many graphs, and a \(C^{1}\) extension across \(\overline{U}\) — are not artefacts of the method. Without some such assumption the two integrals in Equation (7.113) need not exist at all: a bounded continuous function on a bounded open set with a sufficiently wild boundary is not Riemann integrable, and the statement of Theorem 7.97 that “both integrals exist” is then false rather than unproved. The chapter's phrasing should be read with these hypotheses supplied; they hold in every application made in this book, where the regions are simple in the sense of Definition 7.95 and the maps are restrictions of diffeomorphisms defined on larger open sets. Rests on Theorem A.513, Proposition A.511 and Definition 7.95.
Change of Variables in a Multiple Integral discharges the derivation owed at Theorem 7.97 of Real Analysis. The result is cashed in three places. It is what makes the surface integral Equation (7.109) independent of the parametrization used to define it, as Remark 7.102 states explicitly: two parametrizations of one surface differ by a \(C^{1}\) diffeomorphism of their parameter domains, and the Jacobian factor supplied by Equation (7.113) cancels the corresponding factor in \(\pp_{u}\vect{r}\times\pp_{v}\vect{r}\) — so the flux, and with it Theorem 7.101 and Theorem 7.100, is well defined. It supplies the polar and spherical volume elements of Example 7.98, used throughout the physical parts and, in this appendix alone, in The Helmholtz Decomposition. And it is the statement behind the invariance of phase-space volume, Corollary 22.24 of Part III, where the relevant map is the Hamiltonian flow and the Jacobian determinant is the one Liouville's theorem shows to be unity.
The Helmholtz Decomposition
This appendix proves Theorem 7.105 of Real Analysis: every smooth compactly supported vector field on \(\R^{3}\) splits as in Equation (7.129) into a gradient and a curl, with the potentials given explicitly by Equation (7.130), and the two summands are unique among fields that vanish at infinity. It is the theorem that licenses the routine physical move of treating the irrotational and the solenoidal parts of a field separately — the longitudinal and transverse parts of the electromagnetic field, the P and S waves of a seismogram, the potential and vortical parts of a flow.
The paragraph following the proof link in Section 7.11.5 already gives the mechanism, and it is correct: each Cartesian component of \(\vect{w}\) is a Newtonian potential, so \(\nabla^{2}\vect{w}=\vect{u}\), and Equation (7.105) read backwards turns that Laplacian into the required sum. What the chapter does not supply, and what this section owes, is the analysis: that the singular integral Equation (7.130) converges at all, that it may be differentiated under the integral sign as often as one pleases, that the resulting field decays, and that the decay is enough to force uniqueness. That is the whole content of what follows.
One organisational point should be made at the outset. The argument uses two results of Partial Differential Equations — the fundamental-solution property of the Newtonian potential (Proposition 10.67) and the weak maximum principle (Theorem 10.75) — and Partial Differential Equations comes after Real Analysis. Nothing is circular: neither of those two results uses Theorem 7.105, or anything proved here. But the forward dependence is real and is why the derivation sits in the appendix rather than in the chapter, where it would have to be either postponed or duplicated. See Remark A.526.
Throughout, \(\vect{u}\) is a smooth vector field on \(\R^{3}\) whose support is contained in the ball \(\abs{\vect{y}}\le R_{0}\), and we write
for the Newtonian kernel Equation (10.63). The section is pure vector analysis, but the SI bookkeeping is worth recording once because the physical parts read the decomposition off it: \(\Phi\) has the dimension of a reciprocal length and \(\dd^{3}z\) that of a volume, so \(\vect{w}\) carries the unit of \(\vect{u}\) multiplied by \(\mathrm{m}^{2}\); \(\phi\) and \(\vect{\Psi}\), being first derivatives of \(\vect{w}\), carry that of \(\vect{u}\) multiplied by \(\mathrm{m}\); and \(\nabla\phi\) and \(\nabla\times\vect{\Psi}\) carry the unit of \(\vect{u}\) itself, as Equation (7.129) requires.
What the derivation uses
From Real Analysis: the Leibniz integral rule (Theorem 7.77), both halves of the fundamental theorem of calculus (Theorems 7.42 and 7.43), uniform continuity (Theorem 7.25), the Cauchy criterion (Theorem 7.8), the second-order identities of the nabla calculus (Proposition 7.91), the class \(C^{1}\) (Definition 7.66) and the theorem that a \(C^{1}\) map is differentiable (Theorem 7.68), and the spherical volume element of Example 7.98 — which is itself an instance of Change of Variables in a Multiple Integral. From Partial Differential Equations: Proposition 10.67 and Theorem 10.75, both proved there.
The kernel is integrable, and uniformly so
For \(0\le\delta<\varepsilon\),
so that for every bounded \(g\) on \(\R^{3}\),
Rests on Example 7.98, Theorem 7.97 and Definition 7.93.
Derives Lemma A.518. The shell \(\delta\le\abs{\vect{z}}\le\varepsilon\) is a compact region decomposing into finitely many simple regions (Definition 7.95), and \(\abs{\vect{z}}^{-1}\) is continuous on it, hence integrable. Passing to spherical coordinates, whose Jacobian determinant is \(r^{2}\sin\theta\) (Example 7.98, licensed by Theorem 7.97),
which is Equation (A.817). Bounding the integrand of Equation (A.818) by \(\left(4\pi\abs{\vect{z}}\right)^{-1}\sup\abs{g}\) and dropping \(\delta\) gives the estimate, and the bound does not involve \(\vect{x}\).
∎Let \(g\) be continuous on \(\R^{3}\) with support in \(\abs{\vect{y}}\le R_{0}\). For \(\varepsilon>0\) put
a proper integral over a compact region for each \(\vect{x}\). As \(\varepsilon\to0\) these converge uniformly on \(\R^{3}\) to a continuous limit \(\Phi*g\), and
the right-hand side being understood in the same excised sense. Rests on Lemma A.518, Theorem 7.8 and Theorem 7.97.
Derives Proposition A.519. Fix \(\vect{x}\) in a bounded set. The integrand of Equation (A.820) vanishes unless \(\vect{x}-\vect{z}\) lies in the support of \(g\), i.e. unless \(\vect{z}\) lies in a fixed bounded set; so the integration is over a compact shell, on which the integrand is continuous, and the integral exists. For \(0<\varepsilon'<\varepsilon\) the difference is the integral over the shell \(\varepsilon'\le\abs{\vect{z}}\le\varepsilon\), bounded by \(\tfrac12\left(\sup\abs{g}\right)\varepsilon^{2}\) uniformly in \(\vect{x}\) by Equation (A.818). The family is therefore uniformly Cauchy, so it converges pointwise (Theorem 7.8) and the convergence is uniform.
Each truncation is continuous in \(\vect{x}\): on a compact shell and for \(\vect{x}\) in a compact ball the integrand is uniformly continuous (Theorem 7.25), so the integral is continuous in the parameter — this is the last clause of Theorem 7.77, whose proof needs only that bound. A uniform limit of continuous functions is continuous: given \(\varepsilon>0\) choose \(\varepsilon_{0}\) with \(\abs{\left(\Phi*g\right)-\left(\Phi*g\right)_{\varepsilon_{0}}} <\varepsilon/3\) everywhere and then use the continuity of the truncation at the point in question, the three errors summing to \(\varepsilon\); this is the argument already made for series in Lemma 9.6. Finally Equation (A.821) is the substitution \(\vect{y}=\vect{x}-\vect{z}\), a translation composed with a reflection, whose Jacobian determinant has modulus one (Theorem 7.97).
∎Differentiating under the integral sign
The singular factor in Equation (A.820) does not depend on \(\vect{x}\). That is the whole trick: every derivative in \(\vect{x}\) falls on \(g\), where it is harmless, and the singularity is never differentiated.
Let \(O\subseteq\R^{n}\) be open and let \(f_{j}:O\to\R\) be of class \(C^{1}\), converging pointwise on \(O\) to \(f\), with each \(\pp_{i}f_{j}\) converging uniformly on \(O\) to a function \(h_{i}\). Then \(f\) is of class \(C^{1}\) on \(O\) and \(\pp_{i}f=h_{i}\). Rests on Theorem 7.42, Theorem 7.43 and Definition 7.66.
Derives Lemma A.520. Each \(h_{i}\) is continuous, being a uniform limit of continuous functions (the \(\varepsilon/3\) argument recalled in the proof of Proposition A.519). Fix \(\vect{x}\in O\) and \(\sigma>0\) with the segment \(\vect{x}+s\vect{e}_{i}\), \(\abs{s}\le\sigma\), inside \(O\). By the fundamental theorem of calculus (Theorem 7.43) applied to the one-variable \(C^{1}\) function \(s\mapsto f_{j}(\vect{x}+s\vect{e}_{i})\),
The left-hand side tends to \(f(\vect{x}+t\vect{e}_{i})-f(\vect{x})\); on the right the error is at most \(\abs{t}\sup_{O}\abs{\pp_{i}f_{j}-h_{i}}\), which tends to zero. Hence \(f(\vect{x}+t\vect{e}_{i})-f(\vect{x}) =\int_{0}^{t}h_{i}(\vect{x}+s\vect{e}_{i})\,\dd s\), and since \(h_{i}\) is continuous, differentiating at \(t=0\) (Theorem 7.42) gives \(\pp_{i}f(\vect{x})=h_{i}(\vect{x})\). All partial derivatives of \(f\) exist and are continuous, so \(f\) is of class \(C^{1}\) (Definition 7.66) and differentiable (Theorem 7.68).
∎Let \(g\) be of class \(C^{k}\), \(k\ge1\), with compact support. Then \(\Phi*g\) is of class \(C^{k}\) and, for every multi-index of order at most \(k\),
In particular \(\Phi*g\) is smooth when \(g\) is. Rests on Lemma A.520, Theorem 7.77 and Lemma A.518.
Derives Proposition A.521. It suffices to prove the case \(m=1\) and to iterate, since \(\pp_{i}g\) is again \(C^{k-1}\) with compact support.
Fix a ball \(\abs{\vect{x}}<\rho\). For \(\vect{x}\) in it, the integrand of Equation (A.820) vanishes unless \(\vect{z}\) lies in the compact shell \(D_{\varepsilon}=\set{\varepsilon\le\abs{\vect{z}}\le \rho+R_{0}}\), which decomposes into finitely many simple regions. On \(D_{\varepsilon}\times\left(-\rho,\rho\right)\) the function \((\vect{z},x^{i})\mapsto\Phi(\vect{z})g(\vect{x}-\vect{z})\) is continuous with continuous derivative \(-\Phi(\vect{z})\left(\pp_{i}g\right)(\vect{x}-\vect{z})\) in the parameter \(x^{i}\) — the kernel is bounded away from its singularity on \(D_{\varepsilon}\) — so the Leibniz integral rule (Theorem 7.77) applies in each variable \(x^{i}\) separately and gives
the sign of the chain rule being absorbed because the derivative is taken with respect to \(\vect{x}\) and not \(\vect{z}\): writing \(g(\vect{x}-\vect{z})\) and differentiating in \(x^{i}\) produces \(\left(\pp_{i}g\right)(\vect{x}-\vect{z})\) with no sign at all.
By Proposition A.519 applied to \(g\) and to \(\pp_{i}g\), both sides of Equation (A.824) converge uniformly as \(\varepsilon\to0\), the first to \(\Phi*g\) and the second to \(\Phi*\pp_{i}g\); each truncation is \(C^{1}\) on the ball by the same Leibniz rule. Lemma A.520 therefore gives \(\pp_{i}\left(\Phi*g\right)=\Phi*\pp_{i}g\) there, and \(\rho\) was arbitrary.
∎Poisson's equation
Let \(\vect{u}\) be smooth with compact support and let \(\vect{w}=\Phi*\vect{u}\), meaning the convolution taken componentwise. Then \(\vect{w}\) is smooth and
Rests on Proposition A.521, Proposition 10.67 and Definition 17.38.
Derives Proposition A.522. Smoothness and \(\nabla^{2}w_{i}=\Phi*\left(\nabla^{2}u_{i}\right)\) are Proposition A.521. Fix \(\vect{x}\) and put \(\varphi(\vect{z})=u_{i}(\vect{x}-\vect{z})\), a smooth function of compact support and hence a test function in the sense of Definition 17.38. Its Laplacian in \(\vect{z}\) is \(\left(\nabla^{2}u_{i}\right)(\vect{x}-\vect{z})\), the two sign changes of the chain rule cancelling in a second derivative. By Proposition 10.67, which says precisely that \(\int_{\R^{3}}\Phi\,\nabla^{2}\varphi\,\dd^{3}z=\varphi(\vect{0})\) for every test function,
which is Equation (A.825) componentwise. Note that the integral on the left is the excised one of Proposition A.519 and the one in Proposition 10.67 is the same limit, the kernel being absolutely integrable near the origin by Equation (A.817).
∎Existence of the decomposition
Let \(\vect{u}\) be smooth on \(\R^{3}\) with compact support and let \(\vect{w}=\Phi*\vect{u}\). Then \(\phi=\nabla\cdot\vect{w}\) and \(\vect{\Psi}=-\nabla\times\vect{w}\) are smooth, satisfy Equation (7.130), and
Rests on Proposition A.522, Proposition 7.91 and Equation (7.105).
Derives Theorem A.523. \(\vect{w}\) is smooth by Proposition A.521, so \(\phi\) and \(\vect{\Psi}\) are, and Equation (A.821) identifies \(\vect{w}\) with the field displayed in Equation (7.130). The identity Equation (7.105) of Proposition 7.91, applied to the \(C^{2}\) field \(\vect{w}\) and rearranged, reads
and the left-hand side is \(\vect{u}\) by Equation (A.825). Finally \(\nabla\cdot\vect{\Psi} =-\nabla\cdot\left(\nabla\times\vect{w}\right)=0\) by Equation (7.104), and \(\nabla\times\nabla\phi=\vect{0}\) by Equation (7.103), so the first summand is irrotational and the second solenoidal, as Theorem 7.105 asserts.
∎Decay at infinity
The uniqueness clause of Theorem 7.105 speaks of fields that tend to zero at infinity. The potentials just constructed do, and with a definite rate.
With \(\vect{u}\) supported in \(\abs{\vect{y}}\le R_{0}\) and \(\vect{w}=\Phi*\vect{u}\) as above, there is a constant \(C\), depending only on \(R_{0}\) and \(\sup\abs{\vect{u}}\), such that for \(\abs{\vect{x}}\ge2R_{0}\)
In particular \(\nabla\phi\) and \(\nabla\times\vect{\Psi}\) tend to zero at infinity, and so do \(\phi\) and \(\vect{\Psi}\) themselves. Rests on Proposition A.519, Theorem 7.77 and Proposition A.521.
Derives Proposition A.524. Use the form Equation (A.821), in which the integration runs over the fixed compact ball \(\abs{\vect{y}}\le R_{0}\) and the singular factor now depends on \(\vect{x}\). For \(\abs{\vect{x}}\ge2R_{0}\) and \(\vect{y}\) in that ball,
so the integrand is bounded and no singularity is present: the integral is proper, and the Leibniz rule (Theorem 7.77) may be applied directly in each \(x^{i}\), differentiating the kernel. Doing so,
and similarly the second derivatives of the kernel are bounded by \(3\abs{\vect{x}-\vect{y}}^{-3}\). Inserting Equation (A.830) and bounding the integral by the volume \(\tfrac43\pi R_{0}^{3}\) of the ball containing the support times the supremum of the integrand gives all three estimates of Equation (A.829) with the single constant \(C=8R_{0}^{3}\sup\abs{\vect{u}}\): the three bounds so obtained are \(\tfrac23R_{0}^{3}\sup\abs{\vect{u}}\), \(\tfrac43R_{0}^{3}\sup\abs{\vect{u}}\) and \(8R_{0}^{3}\sup\abs{\vect{u}}\) respectively, the factor \(\left(4\pi\right)^{-1}\) cancelling against \(\tfrac43\pi\) and the remaining powers of \(2\) coming from Equation (A.830). The bound is not sharp and is not meant to be. Since \(\nabla\phi\) and \(\nabla\times\vect{\Psi}\) are built from second derivatives of \(\vect{w}\), they are \(O\left(\abs{\vect{x}}^{-3}\right)\); \(\phi\) and \(\vect{\Psi}\) are built from first derivatives and are \(O\left(\abs{\vect{x}}^{-2}\right)\).
∎Uniqueness
Let \(\phi_{1},\phi_{2}\) be \(C^{3}\) scalar fields and \(\vect{\Psi}_{1},\vect{\Psi}_{2}\) be \(C^{3}\) vector fields on \(\R^{3}\) with
and suppose each of the four fields \(\nabla\phi_{1},\nabla\times\vect{\Psi}_{1}, \nabla\phi_{2},\nabla\times\vect{\Psi}_{2}\) tends to zero as \(\abs{\vect{x}}\to\infty\). Then \(\nabla\phi_{1}=\nabla\phi_{2}\) and \(\nabla\times\vect{\Psi}_{1}=\nabla\times\vect{\Psi}_{2}\). Rests on Proposition 7.91, Theorem 10.75 and Equation (7.105).
Derives Theorem A.525. Put \(\vect{h}=\nabla\phi_{1}-\nabla\phi_{2} =\nabla\times\vect{\Psi}_{2}-\nabla\times\vect{\Psi}_{1}\), a \(C^{2}\) field tending to zero at infinity. Read from the left it is a gradient, so \(\nabla\times\vect{h}=\vect{0}\) by Equation (7.103); read from the right it is a curl, so \(\nabla\cdot\vect{h}=0\) by Equation (7.104). Hence Equation (7.105) gives
so each Cartesian component \(h_{i}\) is harmonic on the whole of \(\R^{3}\) (Definition 10.66).
Fix \(\vect{x}_{0}\) and let \(R>\abs{\vect{x}_{0}}\). On the open ball \(B_{R}\) the function \(h_{i}\) is harmonic and continuous on the closure, so by the weak maximum principle (Theorem 10.75) it attains both its extremes on the sphere \(\abs{\vect{x}}=R\):
Both bounds tend to zero as \(R\to\infty\), because \(\vect{h}\to\vect{0}\) at infinity: given \(\varepsilon>0\) there is \(R_{1}\) with \(\abs{\vect{h}}<\varepsilon\) outside \(B_{R_{1}}\), so for \(R>R_{1}\) the two extremes lie in \(\left(-\varepsilon,\varepsilon\right)\). Hence \(\abs{h_{i}(\vect{x}_{0})}<\varepsilon\) for every \(\varepsilon>0\), and \(\vect{h}=\vect{0}\) identically.
∎Proof of Theorem 7.105. Derives Theorem 7.105. Existence, with the potentials Equation (7.130), is Theorem A.523; the two summands are respectively curl-free and divergence-free by the same theorem, which is what Proposition 7.91 contributes; and the uniqueness clause is Theorem A.525. That the potentials constructed here do satisfy the decay hypothesis of the uniqueness clause is Proposition A.524.
∎Reading the result
Nothing from outside this treatise. Two results are taken from a later chapter, and that is worth stating plainly rather than letting a reader discover it: Proposition 10.67, which identifies Equation (A.816) as the fundamental solution of the Laplacian, and Theorem 10.75, the weak maximum principle. Both are proved in Partial Differential Equations, both are proved there without any appeal to Theorem 7.105 or to anything in this section, and neither proof uses the Helmholtz decomposition anywhere in the book. So the forward reference is a matter of exposition, not of logic. It is also the reason the derivation belongs in an appendix: placed in Real Analysis it would have to import a chapter that has not yet been written, and placed in Partial Differential Equations it would be far from the vector calculus it completes.
The remaining ingredients are all from Real Analysis and are listed in What the derivation uses. Three things are deliberately not used: no Lebesgue convergence theorem — the passage to the limit \(\varepsilon\to0\) is uniform, by Lemma A.518, and needs only the Cauchy criterion; no Fourier transform, although the decomposition has a one-line proof in Fourier variables that hides every question of convergence; and no theory of distributions beyond the single statement of Proposition 10.67, which is quoted as an identity between numbers, tested on one explicit test function.
Compact support is a convenience, not a necessity. The same argument runs whenever \(\vect{u}\) decays faster than \(\abs{\vect{x}}^{-2}\) with first derivatives decaying faster than \(\abs{\vect{x}}^{-3}\): the excision estimate Equation (A.818) is untouched, and the only new work is to check that the integral over the far region converges, which those rates supply.
Some decay is indispensable, and the failure is not subtle. The constant field \(\vect{u}=\hat{\vect{z}}\) is a gradient, \(\vect{u}=\nabla z\); it is also a curl, \(\vect{u}=\nabla\times\left(x\,\hat{\vect{y}}\right)\), as one checks from Equation (7.87) by evaluating the three components of the curl of \(\left(0,x,0\right)\). So the split Equation (7.129) exists for it in two ways — everything in the gradient, or everything in the curl — and nothing in the local differential identities can choose between them. The decay condition in the uniqueness clause of Theorem 7.105 is what excludes such a field, and Theorem A.525 shows it is exactly enough.
The regularity stated in Theorem A.525 is a small strengthening of the chapter's phrasing, which leaves it implicit: the argument differentiates the potentials three times, so \(C^{3}\) potentials — equivalently \(C^{2}\) summands — are what Equation (A.833) needs. For the potentials constructed here the point is moot, since they are smooth. Rests on Theorem A.525, Theorem 7.105 and Equation (7.87).
The Helmholtz Decomposition discharges the derivation owed at Theorem 7.105 of Real Analysis. It also settles a debt recorded in Part III: Remark 30.46 of Continuum Mechanics and Elasticity states that the textbook route to the separation of P and S waves needs exactly this theorem — existence and uniqueness of the split, with the decay conditions that make it unique — and that Part II did not carry it, which is why the chapter re-routed the proof of Phenomenon 30.43 onto the acoustic tensor and plane waves instead. That re-routing stands as an independent derivation and is not superseded; what changes is that the alternative route is now available. The second site named there, Proposition 30.48, could not be re-routed: its displacement potentials are introduced as an ansatz sufficient for an existence claim, and the chapter says so. With Theorem A.523 in hand that ansatz can be upgraded to a decomposition, and the Rayleigh wave shown to be the general plane-strain surface motion rather than one exhibited example.
The Quotient Manifold Theorem
This appendix proves Theorem 13.67 of Differentiable Manifolds, Tensors, and Curvature: when a Lie group \(G\) of dimension \(d\) acts smoothly, freely and properly on a manifold \(M_{n}\), the set of orbits carries exactly one smooth structure making the projection Equation (13.156) a submersion, its dimension is \(n-d\), its orbits are embedded copies of \(G\), and a function on it is smooth exactly when its pullback is. It is the second of the two standard ways a manifold is produced — the first, Theorem 13.62, cuts a manifold out with equations; this one makes one by identifying points — and it is the one physics reaches for whenever a symmetry is quotiented away.
The consumer that makes the theorem urgent is Theorem 24.47 of Symplectic Geometry of Phase Space: the Marsden–Weinstein reduced phase space is a level set of a momentum map divided by the isotropy group of that level, and the entire smooth structure of the reduced space comes from the statement proved here. The same theorem is what makes a configuration space of shapes out of a configuration space of frames, and what turns a gauge orbit space into a manifold wherever the action is free.
The proof runs in five movements: the orbit space is Hausdorff and second countable (The orbit space is Hausdorff and second countable); orbits are embedded submanifolds diffeomorphic to \(G\) (Orbits are embedded submanifolds); through every point there is a slice, a transversal meeting nearby orbits at most once (The slice); the slices are the charts (The charts and the smooth structure); and the universal property follows, with it the uniqueness of the structure (The universal property and uniqueness). The heart is the slice, and it is built with the constant rank theorem Theorem 13.63 and the inverse function theorem Corollary A.292 and nothing else. In particular no Lie algebra, no exponential map and no fundamental vector field is used: only that \(G\) is a manifold on which multiplication and inversion are smooth.
What the derivation uses, and what it must state itself
From Differentiable Manifolds, Tensors, and Curvature: the definition of the action and of freeness and properness (Definition 13.66), smooth maps and their differentials (Definitions 13.49 and 13.52), immersions, submersions and embeddings (Definition 13.53), embedded submanifolds (Definition 13.54), the differentiable structure (Definition 13.47) and the constant rank theorem (Theorem 13.63) together with the observation of Remark 13.65 that it transfers to manifolds by being read in charts. From The Implicit Function Theorem: the inverse function theorem Corollary A.292. From Topological and Metric Spaces: compactness (Definition 6.9), the continuous image of a compact set (Proposition 6.10) and homeomorphisms (Definition 6.7). From Lie Groups, Lie Algebras, and Fibre Bundles: nothing beyond the definition of a Lie group as a manifold with smooth multiplication and inversion.
Two separation properties are named in the statement of Theorem 13.67 and are defined nowhere in Part II; Topological and Metric Spaces carries neither. They are stated here, and the gap is recorded as a Part II debt in Remark A.541.
A topological space \(X\) is Hausdorff if any two distinct points have disjoint open neighbourhoods; second countable if its topology has a countable basis, i.e. a countable family of open sets of which every open set is a union; and locally compact if every point has a compact neighbourhood. Rests on Definitions 6.1, 6.2 and 6.9.
Every manifold in the sense of Definition 13.48 is assumed Hausdorff and second countable, as is standard and as the statement of Theorem 13.67 presupposes when it asserts those properties of the quotient. Every manifold is also locally compact and first countable, both because it is locally homeomorphic to \(\R^{n}\): a closed coordinate ball is a compact neighbourhood, and the coordinate balls of rational radius about a point form a countable neighbourhood basis. First countability is what licenses the sequential arguments used below, in which a property is established by testing it on convergent sequences.
Let \(X\) be Hausdorff.
-
Every compact \(C\subseteq X\) is closed.
-
If \(X\) is in addition locally compact and \(F:Y\longrightarrow X\) is continuous and proper — the preimage of every compact set is compact — then \(F\) carries closed sets to closed sets.
Rests on Definition A.529, Definition 6.9 and Proposition 6.10.
Derives Lemma A.530. (1) Let \(q\notin C\). For each \(c\in C\) choose disjoint open \(V_{c}\ni q\) and \(W_{c}\ni c\); the \(W_{c}\) cover \(C\), so finitely many \(W_{c_{1}},\ldots,W_{c_{m}}\) do, and \(V_{c_{1}}\cap\cdots\cap V_{c_{m}}\) is an open neighbourhood of \(q\) disjoint from \(C\). Hence the complement of \(C\) is open.
(2) Let \(A\subseteq Y\) be closed and let \(q\) lie in the closure of \(F(A)\). Choose a compact neighbourhood \(K\) of \(q\). Then \(F^{-1}(K)\) is compact by properness, so \(A\cap F^{-1}(K)\) is compact — a closed subset of a compact set — and \(F\left(A\cap F^{-1}(K)\right)=F(A)\cap K\) is compact by Proposition 6.10, hence closed by (1). Every neighbourhood of \(q\) contained in the interior of \(K\) meets \(F(A)\), and the intersection lies in \(K\); so \(q\) lies in the closure of \(F(A)\cap K\), which is \(F(A)\cap K\) itself. Thus \(q\in F(A)\).
∎Throughout the rest of the section, \(G\) is a Lie group of dimension \(d\), \(M=M_{n}\) a manifold, and the action is smooth, free and proper in the sense of Definition 13.66. We write \(L_{g}:M\to M\) for the diffeomorphism \(P\mapsto g\cdot P\), with inverse \(L_{g^{-1}}\), and
for the map whose properness is the hypothesis. We write \(\theta_{P}:G\to M\), \(\theta_{P}(g)=g\cdot P\), for the orbit map, and \(k=n-d\).
The orbit space is Hausdorff and second countable
For every open \(U\subseteq M\) the set \(\pi(U)\) is open in \(M/G\), and \(\pi\) is continuous and surjective. Rests on Definitions 6.2, 6.6 and 13.66.
Derives Lemma A.531. Continuity and surjectivity are the definition of the quotient topology. For openness,
a union of open sets, since each \(L_{g}\) is a diffeomorphism. A set whose preimage is open is open in the quotient topology.
∎The orbit space \(M/G\) is Hausdorff and second countable, and every orbit is closed in \(M\). Rests on Lemma A.530, Lemma A.531 and Definition 13.66.
Derives Proposition A.532. The orbit relation is closed. \(M\times M\) is Hausdorff and locally compact, being a manifold, and \(\Theta\) of Equation (A.835) is continuous and proper by hypothesis. Write
By Lemma A.530(2) this set is closed — \(G\times M\) being closed in itself. Fixing \(P\) and intersecting with \(M\times\set{P}\) shows each orbit closed in \(M\).
Hausdorff. Let \(\pi(P)\neq\pi(Q)\), so \((Q,P)\notin\mathcal{R}\). Since \(\mathcal{R}\) is closed and the products of open sets form a basis of \(M\times M\), there are open \(U\ni P\) and \(V\ni Q\) with \(\left(V\times U\right)\cap\mathcal{R}=\varnothing\): no point of \(V\) lies on the orbit of a point of \(U\). Then \(\pi(U)\) and \(\pi(V)\) are open by Lemma A.531, contain \(\pi(P)\) and \(\pi(Q)\), and are disjoint — a common point would be an orbit meeting both \(U\) and \(V\), which is what was excluded.
Second countable. Let \(\set{B_{i}}\) be a countable basis of \(M\). Each \(\pi(B_{i})\) is open. If \(W\subseteq M/G\) is open and \(\pi(P)\in W\), then \(P\in\pi^{-1}(W)\), which is open, so \(P\in B_{i}\subseteq\pi^{-1}(W)\) for some \(i\), whence \(\pi(P)\in\pi(B_{i})\subseteq\pi\left(\pi^{-1}(W)\right)=W\), the last equality by surjectivity of \(\pi\). So \(\set{\pi(B_{i})}\) is a countable basis.
∎Orbits are embedded submanifolds
For every \(P\in M\) the map \(\theta_{P}:G\to M\) has the same rank at every point of \(G\), and that rank is \(d\); that is, \(\theta_{P}\) is an injective immersion. Rests on Theorem 13.63, Definition 13.66 and Definition 13.53.
Derives Lemma A.533. Constancy. Write \(\ell_{g}:G\to G\) for left translation \(h\mapsto gh\), a diffeomorphism. The associativity of the action gives \(\theta_{P}\left(gh\right)=g\cdot\left(h\cdot P\right)\), that is
Differentiating at the identity \(e\) and using the chain rule for differentials of smooth maps (Definition 13.52), \(\left(\theta_{P}\right)_{*g}\circ\left(\ell_{g}\right)_{*e} =\left(L_{g}\right)_{*P}\circ\left(\theta_{P}\right)_{*e}\). Both \(\left(\ell_{g}\right)_{*e}\) and \(\left(L_{g}\right)_{*P}\) are isomorphisms, so the ranks of \(\left(\theta_{P}\right)_{*g}\) and \(\left(\theta_{P}\right)_{*e}\) agree.
The rank is \(d\). Suppose it were \(\rho<d\). Read in charts, \(\theta_{P}\) is a \(C^{\infty}\) map of an open subset of \(\R^{d}\) into \(\R^{n}\) of constant rank \(\rho\), so by the constant rank theorem (Theorem 13.63, transferred to manifolds as in Remark 13.65) there are coordinates about \(e\) and about \(P\) in which it reads \(\left(u^{1},\ldots,u^{d}\right)\mapsto \left(u^{1},\ldots,u^{\rho},0,\ldots,0\right)\), which is Equation (13.152). Two distinct nearby points differing only in the coordinate \(u^{\rho+1}\) then have the same image, so \(\theta_{P}\) is not injective on any neighbourhood of \(e\) — and \(\theta_{P}\) is injective, because \(g\cdot P=h\cdot P\) gives \(\left(h^{-1}g\right)\cdot P=P\) and hence \(g=h\) by freeness (Definition 13.66). The rank is therefore \(d\), and \(\left(\theta_{P}\right)_{*g}\) is injective for every \(g\).
∎For every \(P\in M\) the orbit map \(\theta_{P}\) is an embedding, so \(G\cdot P\) is an embedded submanifold of \(M\) of dimension \(d\), diffeomorphic to \(G\). Rests on Lemma A.533, Definition 13.53 and Definition 13.54.
Derives Proposition A.534. By Lemma A.533 the map is an injective immersion; by Definition 13.53 what remains is that it be a homeomorphism onto its image with the topology inherited from \(M\). Equivalently, since \(G\) is first countable, that \(g_{j}\cdot P\to g\cdot P\) in \(M\) implies \(g_{j}\to g\) in \(G\).
Let \(g_{j}\cdot P\to Q=g\cdot P\). The set \(C=\set{g_{j}\cdot P\mid j}\cup\set{Q}\) is compact — a convergent sequence with its limit — and so is \(C\times\set{P}\); properness makes \(\Theta^{-1}\left(C\times\set{P}\right)\) compact, and it contains every \(\left(g_{j},P\right)\). Hence the \(g_{j}\) lie in a compact subset \(K\) of \(G\). Since \(G\) is first countable, a compact subset of it is sequentially compact, so every subsequence of \(\left(g_{j}\right)\) has a further subsequence converging in \(K\); and if \(g_{j_{m}}\to g'\), then continuity of the action gives \(g'\cdot P=\lim g_{j_{m}}\cdot P=Q=g\cdot P\), whence \(g'=g\) by freeness. A sequence in a compact set all of whose convergent subsequences have the limit \(g\) converges to \(g\): otherwise some neighbourhood of \(g\) would be left by infinitely many terms, and those terms would have a convergent subsequence whose limit, lying in the complement of that neighbourhood, could not be \(g\). So \(g_{j}\to g\).
The image of an embedding is an embedded submanifold and the map is a diffeomorphism onto it, as recorded after Definition 13.54.
∎The slice
Let \(P\in M\). There are an open neighbourhood \(O\subseteq G\) of \(e\), an embedded \(k\)-dimensional submanifold \(S\subseteq M\) with \(P\in S\) carried by a single chart \(\varsigma:S\to B\subseteq\R^{k}\), and an open neighbourhood \(\Omega=O\cdot S\) of \(P\) in \(M\), such that
-
the map \(F:O\times S\to\Omega\), \(F(g,Q)=g\cdot Q\), is a diffeomorphism; and
-
\(S\) meets every orbit at most once: if \(Q,Q'\in S\) and \(Q'=g\cdot Q\) for some \(g\in G\), then \(g=e\) and \(Q=Q'\).
Rests on Lemma A.533, Corollary A.292 and Definition 13.66.
Derives Proposition A.535. Choice of a transversal. By Lemma A.533 the image \(V_{P}=\left(\theta_{P}\right)_{*e}\left(T_{e}G\right)\) is a \(d\)-dimensional subspace of \(T_{P}M\). Choose a chart \((\phi,W)\) about \(P\) with \(\phi(P)=0\); composing \(\phi\) with a linear automorphism of \(\R^{n}\) — itself a change of chart — we may assume \(\phi_{*P}\left(V_{P}\right)=\R^{d}\times\set{0}\). For a small ball \(B\subseteq\R^{k}\) put
an embedded \(k\)-submanifold through \(P\) with \(T_{P}S_{0}\oplus V_{P}=T_{P}M\), since \(\phi_{*P}\) carries the two summands onto the complementary \(\set{0}\times\R^{k}\) and \(\R^{d}\times\set{0}\).
The map \(F\) is a local diffeomorphism. On the \(n\)-manifold \(G\times S_{0}\) define \(F(g,Q)=g\cdot Q\), smooth as a restriction of the action. Its differential at \((e,P)\) acts on \(T_{e}G\oplus T_{P}S_{0}\) by
because \(F(\cdot,P)=\theta_{P}\) and \(F(e,\cdot)\) is the inclusion of \(S_{0}\). By the choice of \(S_{0}\) this is a linear isomorphism onto \(T_{P}M\). Read in charts, \(F\) is a smooth map between open subsets of \(\R^{n}\) with invertible derivative at the point, so the inverse function theorem (Corollary A.292) makes it a diffeomorphism of a neighbourhood of \((e,P)\) onto an open neighbourhood of \(P\). Shrinking, that neighbourhood contains a product \(O\times S\) with \(O\ni e\) open in \(G\) and \(S\ni P\) open in \(S_{0}\); put \(\Omega=F\left(O\times S\right)=O\cdot S\), open in \(M\). This is (1).
Shrinking to a genuine slice. Suppose (2) failed for every choice of \(S\). Take \(S_{j}=\phi^{-1}\left(\set{0}\times B_{1/j}\right)\), \(j\in\N\), a basis of neighbourhoods of \(P\) in \(S_{0}\). For each \(j\) there would then be \(Q_{j},Q_{j}'\in S_{j}\) and \(g_{j}\neq e\) with \(g_{j}\cdot Q_{j}=Q_{j}'\). Both sequences converge to \(P\), so \(C=\set{Q_{j}}\cup\set{Q_{j}'}\cup\set{P}\) is compact and \(\Theta^{-1}\left(C\times C\right)\) is compact by properness; it contains every \(\left(g_{j},Q_{j}\right)\), so the \(g_{j}\) lie in a compact subset of \(G\) and a subsequence converges, \(g_{j_{m}}\to g\). Continuity of the action gives \(g\cdot P=\lim g_{j_{m}}\cdot Q_{j_{m}}=\lim Q_{j_{m}}'=P\), so \(g=e\) by freeness. Hence \(g_{j_{m}}\in O\) and \(Q_{j_{m}},Q_{j_{m}}'\in S\) for \(m\) large. But then
and the injectivity of \(F\) on \(O\times S\) forces \(g_{j_{m}}=e\), contradicting \(g_{j}\neq e\). So some \(S_{j}\) satisfies (2); replace \(S\) by it and shrink \(O\) and \(\Omega\) accordingly, which preserves (1).
∎The charts and the smooth structure
With \(O,S,\Omega\) as in Proposition A.535, the set \(\pi(S)=\pi(\Omega)\) is open in \(M/G\) and \(\pi|_{S}:S\to\pi(S)\) is a homeomorphism. Writing \(\varrho:\Omega\to S\) for the second component of \(F^{-1}\), a smooth map, one has \(\varrho=\left(\pi|_{S}\right)^{-1}\circ\pi\) on \(\Omega\). Rests on Proposition A.535, Lemma A.531 and Definition 6.7.
Derives Lemma A.536. Every point of \(\Omega\) is \(g\cdot Q\) with \(Q\in S\), so \(\pi(\Omega)=\pi(S)\), and \(\pi(\Omega)\) is open by Lemma A.531. The map \(\pi|_{S}\) is continuous, surjective onto \(\pi(S)\) by definition, and injective by Proposition A.535(2).
For \(R\in\Omega\) write \(R=F(g,Q)=g\cdot Q\) with \((g,Q)\in O\times S\) unique; then \(\varrho(R)=Q\) lies on the orbit of \(R\) and in \(S\), and by (2) it is the only point of \(S\) on that orbit. Hence \(\varrho=\left(\pi|_{S}\right)^{-1}\circ\pi\) on \(\Omega\), and \(\varrho\) is smooth because \(F^{-1}\) is.
It remains to see that \(\left(\pi|_{S}\right)^{-1}\) is continuous. The restriction \(\pi|_{\Omega}:\Omega\to\pi(S)\) is continuous, surjective and open (Lemma A.531 applied to open subsets of \(\Omega\)), hence a quotient map: a subset of \(\pi(S)\) is open exactly when its preimage is. The map \(\varrho\) is continuous and constant on the fibres of \(\pi|_{\Omega}\), by the previous paragraph; so for \(U\subseteq S\) open the set \(\left(\left(\pi|_{S}\right)^{-1}\right)^{-1}(U)=\pi\left( \varrho^{-1}(U)\right)\) has preimage \(\varrho^{-1}(U)\), which is open, and is therefore open. Thus \(\left(\pi|_{S}\right)^{-1}\) is continuous and \(\pi|_{S}\) is a homeomorphism onto the open set \(\pi(S)\).
∎The maps
one for each slice of Proposition A.535, form a \(k\)-dimensional atlas on \(M/G\). With the differentiable structure it generates (Definition 13.47), \(M/G\) is a smooth manifold of dimension \(k=n-d\), which is Equation (13.157), and \(\pi\) is a smooth surjective submersion. Rests on Lemma A.536, Proposition A.532 and Definition 13.53.
Derives Theorem A.537. The charts cover and are homeomorphisms. Each \(u\) is a homeomorphism of the open set \(\pi(S)\) onto the open \(B\subseteq\R^{k}\), being the composition of the homeomorphism \(\left(\pi|_{S}\right)^{-1}\) of Lemma A.536 with the chart \(\varsigma\) of \(S\); and every orbit meets some slice, namely one built at any of its points.
Smooth compatibility. Let \(S\) and \(S'\) be two slices with \(\pi(S)\cap\pi(S')\neq\varnothing\) and let \(\pi(Q_{0})\) lie in the intersection, \(Q_{0}\in S\). The transition map is \(u'\circ u^{-1}=\varsigma'\circ\tau\circ\varsigma^{-1}\), where \(\tau=\left(\pi|_{S'}\right)^{-1}\circ\pi|_{S}\) sends a point of \(S\) to the unique point of \(S'\) on its orbit. Put \(Q_{0}'=\tau(Q_{0})\in S'\) and let \(g_{0}\in G\) with \(Q_{0}'=g_{0}\cdot Q_{0}\). The set \(L_{g_{0}}^{-1}\left(\Omega'\right)\) is an open neighbourhood of \(Q_{0}\) in \(M\), and on it the map \(\varrho'\circ L_{g_{0}}\) is smooth and takes each point to the unique point of \(S'\) on its orbit — which is \(\tau\) where both are defined. Restricted to the embedded submanifold \(S\), a smooth map remains smooth, so \(\tau\) is smooth near \(Q_{0}\) as a map \(S\to S'\), and the transition map is smooth as a map between open subsets of \(\R^{k}\). Exchanging the roles of \(S\) and \(S'\) gives smoothness of the inverse. Together with Proposition A.532 — Hausdorff and second countable — this makes \(M/G\) a smooth \(k\)-manifold.
\(\pi\) is a submersion. Fix \(P\in M\), take the slice at \(P\) and use \(F^{-1}:\Omega\to O\times S\) to give \(M\) the chart \(\left(\kappa,\varsigma\right)\circ F^{-1}\) near \(P\), where \(\kappa\) is any chart of \(G\) about \(e\). In these coordinates, and with the chart \(u\) downstairs, \(\pi\) reads
because \(\pi\left(F(g,Q)\right)=\pi(Q)\) and \(u\left(\pi(Q)\right)=\varsigma(Q)\). A projection has surjective differential, so \(\pi\) is a submersion at \(P\) (Definition 13.53); and \(P\) was arbitrary.
∎The universal property and uniqueness
Let \(p:M\to X\) be a smooth submersion between manifolds and let \(P\in M\). There are an open \(W\ni p(P)\) in \(X\) and a smooth \(s:W\to M\) with \(p\circ s=\id_{W}\) and \(s\left(p(P)\right)=P\). Rests on Theorem 13.63, Definition 13.53 and Definition 13.49.
Derives Lemma A.538. A submersion has constant rank \(\dim X\), so Theorem 13.63, read in charts as in Remark 13.65, supplies coordinates \(\left(u^{1},\ldots,u^{\dim M}\right)\) about \(P\) and \(\left(y^{1},\ldots,y^{\dim X}\right)\) about \(p(P)\), both vanishing at the point, in which \(p\) reads \(\left(u^{1},\ldots\right)\mapsto\left(u^{1},\ldots,u^{\dim X}\right)\). In those coordinates set \(s(y)=(y,0)\), which is smooth and satisfies both requirements on the coordinate neighbourhood.
∎Let \(p:M\to X\) be a smooth surjective submersion and \(N\) a manifold. A map \(f:X\to N\) is smooth if and only if \(f\circ p\) is smooth. Rests on Lemma A.538 and Definition 13.49.
Derives Proposition A.539. If \(f\) is smooth then so is \(f\circ p\), a composition of smooth maps. Conversely let \(f\circ p\) be smooth and let \(x\in X\); pick \(P\) with \(p(P)=x\), which exists by surjectivity, and a smooth local section \(s\) on \(W\ni x\) from Lemma A.538. On \(W\), \(f=f\circ p\circ s\) is a composition of smooth maps, hence smooth; and \(f\) is continuous there for the same reason. Smoothness is local, so \(f\) is smooth.
∎There is at most one differentiable structure on the topological space \(M/G\) for which \(\pi\) is a smooth submersion. Rests on Proposition A.539 and Definition 13.47.
Derives Corollary A.540. Let \(X\) and \(X'\) denote the same topological space \(M/G\) carrying two such structures. Since \(\pi:M\to X'\) is smooth and \(\pi:M\to X\) is a smooth surjective submersion, Proposition A.539 applied to \(f=\id\) shows \(\id:X\to X'\) is smooth. Exchanging the two gives \(\id:X'\to X\) smooth. The identity is therefore a diffeomorphism, and two differentiable structures related by the identity diffeomorphism have the same smooth functions and hence the same maximal atlas.
∎Proof of Theorem 13.67. Derives Theorem 13.67. Assembling. The orbit space is Hausdorff and second countable by Proposition A.532; it carries a smooth structure of dimension \(n-d\) for which Equation (13.156) is a submersion, by Theorem A.537, and Equation (13.157) is that dimension count; the structure is the only one with that property, by Corollary A.540; every orbit is an embedded submanifold diffeomorphic to \(G\), by Proposition A.534; and a map \(f\) on the quotient is smooth exactly when \(f\circ\pi\) is, by Proposition A.539 applied to the surjective submersion \(\pi\).
∎Reading the result
No theorem from outside this treatise is used, and the two analytic inputs — the inverse function theorem Corollary A.292 and the constant rank theorem Theorem 13.63 — are proved in The Implicit Function Theorem and in Differentiable Manifolds, Tensors, and Curvature respectively. What is missing from Part II, and is supplied here rather than cited, is point-set vocabulary: Topological and Metric Spaces defines neither the Hausdorff property nor second countability nor local compactness, although Theorem 13.67 asserts the first two of the quotient and Definition 13.48 tacitly assumes all of them of a manifold. They are stated in Definition A.529, and the two facts about them that the proof consumes — a compact subset of a Hausdorff space is closed, and a proper map into a locally compact Hausdorff space is closed — are proved in Lemma A.530. That is a Part II debt of about a page, and it is the only one this section incurs.
Two things are deliberately not used, and the omission is what keeps the section self-contained. There is no Lie algebra anywhere: the standard proof shows the orbit map to be an immersion by computing its differential as the fundamental vector field of an element of \(\mathfrak{g}\) and appealing to freeness, which would import the machinery of Lie Groups, Lie Algebras, and Fibre Bundles; instead Lemma A.533 gets the same conclusion from the constant rank theorem alone, by observing that a constant-rank map of rank less than \(d\) cannot be injective near a point. And there is no partition of unity and no smooth-structure transport theorem: the charts are the slices, and the transitions are computed from the action itself.
Freeness is used twice and properness three times, and it is worth saying where, because Example 13.68 shows what happens when either fails but not which step breaks.
Freeness makes the orbit map injective, which is what forces its rank to be full in Lemma A.533 — and hence what gives every orbit the same dimension \(d\), without which no dimension count is possible. It is used again in Proposition A.535 to identify the limiting group element as \(e\). When it fails, orbits through different points have different dimensions, and the quotient is a stratified space: this is the \(\SO(2)\) action of Example 13.68, whose orbit space is a half-line with a bad endpoint.
Properness gives the closedness of the orbit relation (Proposition A.532), and with it the Hausdorff property of the quotient; it gives the embedding property of the orbit map (Proposition A.534); and it gives the second shrinking in Proposition A.535, which is what turns a local transversal into a genuine slice. The irrational flow on the torus in Example 13.68 fails at the first of these, and every later step fails with it: no slice exists, because every orbit returns arbitrarily close to every point. Rests on Proposition A.535, Proposition A.532 and Example 13.68.
The Quotient Manifold Theorem discharges the derivation owed at Theorem 13.67 of Differentiable Manifolds, Tensors, and Curvature. The statement is used at once in the same chapter, where it is the second of the two manufacturing theorems beside Theorem 13.62 and Corollary 13.64, and it is what makes Example 13.68 a statement about manifolds rather than about sets. Its principal consumer in the physical parts is Theorem 24.47 of Symplectic Geometry of Phase Space: symplectic reduction divides a level set of a momentum map — a manifold by Theorem 13.62 — by the isotropy group of that level, and every smooth statement about the reduced phase space, from the existence of its symplectic form to the descent of the dynamics, rests on the projection being a submersion with the universal property Proposition A.539. The one hypothesis to watch there is the same one watched here: the action must be free, and where it is not the reduced space is singular and none of this applies.
The Liouville–Arnold Theorem
This appendix proves Theorem 23.31 of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy: a system of \(f\) degrees of freedom possessing \(f\) functionally independent integrals \(F_{1}=\Ham,F_{2},\ldots,F_{f}\) in involution has each compact connected component of a regular level set diffeomorphic to the torus \(T^{f}\), carries quasi-periodic motion on it, and admits action–angle variables in a neighbourhood of it. The last clause is what makes the theorem worth its length: it establishes Theorem 23.30 without the hypothesis of separability under which that theorem was proved, so that the action–angle description of a bounded integrable motion is a consequence of the involution relations alone and not of a lucky choice of coordinates.
Exactly one ingredient is imported, and Remark 23.32 already tells the reader which: that a discrete subgroup of \(\R^{f}\) with compact quotient is a lattice of full rank. It is stated below as Theorem A.550 and used once, in The translation action and the torus. Everything else is proved here from the symplectic material of Symplectic Geometry of Phase Space and the manifold material of Differentiable Manifolds, Tensors, and Curvature; in particular Frobenius' theorem, which the argument needs twice, is Theorem 13.133 and is not reproved.
Setting and statement
Let \((M,\omega)\) be a symplectic manifold of dimension \(2f\) (Definition 24.8), let \(F_{1},\ldots,F_{f}\) be smooth real functions on \(M\) in involution,
which is Equation (23.54), and write \(\vect{F}=(F_{1},\ldots,F_{f}):M\longrightarrow\R^{f}\). Fix \(\vect{c}\in\R^{f}\) and suppose that \(\vect{c}\) is a regular value, so that the differentials \(\dd F_{1},\ldots,\dd F_{f}\) are linearly independent at every point of \(M_{\vect{c}}=\vect{F}^{-1}(\vect{c})\). Denote by \(X_{a}=X_{F_{a}}\) the Hamiltonian vector field of \(F_{a}\), defined by \(\iota_{X_{a}}\omega=\dd F_{a}\) as in Equation (24.5); by Equation (24.7) it acts on functions as \(X_{a}[g]=\pb{g}{F_{a}}\).
Nothing here fixes the SI dimension of \(F_{2},\ldots,F_{f}\): each is whatever conserved quantity it happens to be, and by Remark 24.11 the vector field \(X_{a}\) carries \([F_{a}]/\mathrm{J}\,\mathrm{s}\), so its flow parameter carries \(\mathrm{J}\,\mathrm{s}/[F_{a}]\) — seconds when \(F_{a}\) is the energy, radians when it is an angular momentum. The action variables of The action variables carry \(\mathrm{J}\,\mathrm{s}\) without exception, because \(\oint p_{a}\dd q^{a}\) does; the angles are pure numbers. That is the whole of the dimensional bookkeeping, and it is worth stating because the lattice \(\Gamma\) of The translation action and the torus lives in a copy of \(\R^{f}\) whose \(f\) axes carry \(f\) different units.
Under the hypotheses above, let \(N\) be a compact connected component of \(M_{\vect{c}}\). Then
-
\(N\) is an \(f\)-dimensional embedded submanifold of \(M\), diffeomorphic to the torus \(T^{f}=\R^{f}/\Z^{f}\), and it is Lagrangian: \(\omega\) vanishes on \(T_{P}N\) for every \(P\in N\);
-
in the angular coordinates supplied by that diffeomorphism the flow of \(\Ham=F_{1}\) is linear, so the motion is quasi-periodic with \(f\) frequencies;
-
on a neighbourhood of \(N\) there exist action–angle variables \((I_{a},\theta^{a})\), canonical, with \(\Ham=\Ham(I)\) and with the equations of motion Equation (23.51) and Equation (23.52).
Rests on Equation (23.54), Theorem 13.133 and Equation (24.8).
The commuting fields and the level set
On the open set \(U\subseteq M\) where \(\dd F_{1},\ldots,\dd F_{f}\) are linearly independent:
-
\(X_{a}[F_{b}]=0\) for all \(a,b\), so every \(X_{a}\) is tangent to every level set of \(\vect{F}\);
-
\(\comm{X_{a}}{X_{b}}=0\) as vector fields;
-
\(X_{1},\ldots,X_{f}\) are linearly independent at every point of \(U\), and \(\omega(X_{a},X_{b})=0\).
Derives Lemma A.546. (1) By Equation (24.7), \(X_{a}[F_{b}]=\pb{F_{b}}{F_{a}}\), which vanishes by Equation (A.844). A vector annihilating every \(\dd F_{b}\) is tangent to the level set through its base point.
(2) Equation (24.8) states that \(X_{\pb{u}{v}}=-\comm{X_{u}}{X_{v}}\), so \(\comm{X_{a}}{X_{b}}=-X_{\pb{F_{a}}{F_{b}}}=-X_{0}=0\). This is the one place where involution does its real work, and it is worth saying what it buys: the flows of the \(f\) integrals commute, which is what will make them into an action of an abelian group.
(3) Suppose \(\sum_{a}\lambda^{a}X_{a}(P)=0\) at some \(P\in U\). Contracting with \(\omega\) and using Equation (24.5) gives \(\sum_{a}\lambda^{a}\,\dd F_{a}(P)=0\), whence every \(\lambda^{a}\) vanishes by the independence of the differentials. Finally \(\omega(X_{a},X_{b})=\pb{F_{a}}{F_{b}}=0\) by Equation (24.7) and Equation (A.844).
∎Let \(D_{P}=\operatorname{span}\set{X_{1}(P),\ldots,X_{f}(P)}\) for \(P\in U\). Then \(D\) is a smooth distribution of constant rank \(f\) (Definition 13.130), it is involutive, and it is integrable. Every connected component \(N\) of \(M_{\vect{c}}\cap U\) is an embedded \(f\)-dimensional submanifold with \(T_{P}N=D_{P}\), hence an integral manifold of \(D\), and it is Lagrangian. Rests on Lemma A.546, Theorem 13.133 and Theorem 13.59.
Derives Proposition A.547. Smoothness and constant rank are Lemma A.546(3); involutivity is Lemma A.546(2), since the bracket of two spanning fields vanishes and so lies in \(D\) trivially. Integrability then follows from Theorem 13.133, which also supplies, around each point of \(U\), a chart in which \(D\) is spanned by the first \(f\) coordinate fields Equation (13.281), so that \(U\) is foliated by \(f\)-dimensional integral manifolds.
Since \(\vect{c}\) is a regular value, \(\vect{F}\) restricted to \(U\) is a submersion onto a neighbourhood of \(\vect{c}\), and by the regular value theorem Theorem 13.59 the set \(M_{\vect{c}}\cap U\) is an embedded submanifold of dimension \(2f-f=f\) with \(T_{P}(M_{\vect{c}}\cap U)=\ker\dd\vect{F}_{P}\). By Lemma A.546(1) every \(X_{a}(P)\) lies in that kernel, and by (3) the \(f\) of them are independent; a subspace of dimension \(f\) inside a space of dimension \(f\) is the whole of it, so \(T_{P}(M_{\vect{c}}\cap U)=D_{P}\). Each connected component is therefore a connected integral manifold of \(D\). That it is Lagrangian is Lemma A.546(3) again: \(\omega\) evaluated on two vectors of \(T_{P}N=D_{P}\) is a combination of the numbers \(\omega(X_{a},X_{b})\), all zero.
∎The translation action and the torus
From here on \(N\) is a compact connected component of \(M_{\vect{c}}\).
Each \(X_{a}\) restricts to a complete vector field on \(N\), and
where \(\phi^{a}\) is the flow of \(X_{a}\), is independent of the order of the factors and defines a smooth action of the additive group \(\R^{f}\) on \(N\). Rests on Lemma A.546, Proposition 13.131 and Theorem 13.125.
Derives Lemma A.548. By Lemma A.546(1) each \(X_{a}\) is tangent to \(N\), so its integral curves through points of \(N\) stay in \(N\) as long as they exist. A smooth vector field on a compact manifold is complete: by Theorem 13.125 each point has a neighbourhood on which the flow exists for a time at least \(\varepsilon\) depending on the neighbourhood, compactness extracts a finite subcover and hence a single \(\varepsilon>0\) that works everywhere on \(N\), and the flow is then extended to all times by composing it with itself. Commutativity of the flows is Proposition 13.131 applied to \(\comm{X_{a}}{X_{b}}=0\), which makes the composition Equation (A.845) independent of the order and gives \(\Phi(\vect{t})\Phi(\vect{t}')=\Phi(\vect{t}+\vect{t}')\); \(\Phi(\vect{0})\) is the identity, and smoothness in \((\vect{t},P)\) jointly is that of each flow.
∎For every \(P\in N\) the orbit map \(\vect{t}\longmapsto\Phi(\vect{t})P\) is a local diffeomorphism of \(\R^{f}\) onto its image; the orbit is open in \(N\); and since \(N\) is connected there is exactly one orbit, so the action is transitive. Rests on Lemma A.548, Corollary A.292 and Proposition A.547.
Derives Lemma A.549. Differentiating Equation (A.845) at \(\vect{t}=\vect{0}\) sends the \(a\)-th basis vector of \(\R^{f}\) to \(X_{a}(P)\), and by Lemma A.546(3) together with Proposition A.547 those \(f\) vectors are a basis of \(T_{P}N\). The differential of the orbit map at the origin is therefore an isomorphism onto \(T_{P}N\), and the inverse function theorem (Corollary A.292, read in a chart of \(N\)) makes the orbit map a diffeomorphism of a neighbourhood of \(\vect{0}\) onto a neighbourhood of \(P\) in \(N\). Applying this at an arbitrary point \(\Phi(\vect{t}_{0})P\) of the orbit — which is legitimate because \(\Phi(\vect{t}_{0})\) is a diffeomorphism of \(N\) carrying the orbit map based at \(P\) to the one based at \(\Phi(\vect{t}_{0})P\) — shows that the orbit contains a neighbourhood of each of its points, i.e. is open. Distinct orbits are disjoint, so the orbits form a partition of \(N\) into open sets; connectedness leaves only one.
∎Let \(\Gamma\subseteq\R^{f}\) be a subgroup that is discrete, in the sense that some neighbourhood of \(\vect{0}\) meets \(\Gamma\) only in \(\vect{0}\). Then there are linearly independent vectors \(\vect{\lambda}_{1},\ldots,\vect{\lambda}_{k}\in\R^{f}\), \(k\leq f\), with
and the quotient group \(\R^{f}/\Gamma\) is compact if and only if \(k=f\), in which case it is diffeomorphic to the torus \(T^{f}\). Rests on Definitions 4.20 and 6.9.
Theorem A.550 is the only statement in this section that is not proved, and Remark 23.32 of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy names it as the debt in the same words. It is elementary — an induction on the dimension of the linear span of \(\Gamma\), at each step choosing a shortest non-zero element of the part of \(\Gamma\) lying in a new direction and showing that discreteness makes it generate that direction's contribution — but it is a theorem about subgroups of \(\R^{f}\), which is algebra and topology rather than mechanics, and it belongs with the material of Algebraic Structures and Topological and Metric Spaces, which do not carry it. The treatise's bibliography holds no source for it either, so the attribution is made in words: it is standard lattice theory, and the form used here is the one Arnold states in Mathematical Methods of Classical Mechanics (2nd edition, 1989), a work for which Symplectic Geometry of Phase Space also records that there is no key. Nothing else below is quoted.
Fix \(P\in N\) and let \(\Gamma=\set{\vect{t}\in\R^{f}\mid\Phi(\vect{t})P=P}\). Then \(\Gamma\) is a discrete subgroup of \(\R^{f}\) of rank exactly \(f\), the induced map
is a diffeomorphism, and \(N\) is diffeomorphic to \(T^{f}\). Rests on Lemma A.549, Theorem A.550 and Lemma A.548.
Derives Proposition A.552. \(\Gamma\) is a subgroup because \(\Phi\) is an action, and it is discrete because Lemma A.549 makes the orbit map injective on some neighbourhood \(B\) of \(\vect{0}\), so \(\Gamma\cap B=\set{\vect{0}}\). The orbit map is constant on cosets of \(\Gamma\) and, by transitivity, descends to a bijection Equation (A.847) of \(\R^{f}/\Gamma\) onto \(N\); it is a local diffeomorphism by the same lemma, hence a diffeomorphism. Since \(N\) is compact so is \(\R^{f}/\Gamma\), and Theorem A.550 then gives \(\Gamma\) of rank \(f\) and \(\R^{f}/\Gamma\cong T^{f}\).
∎Under the identification Equation (A.847), the Hamiltonian flow of \(\Ham=F_{1}\) starting at \(P\) is \(t\longmapsto t\,\vect{e}_{1}+\Gamma\): a straight line in \(\R^{f}\) read modulo the lattice. The motion is therefore quasi-periodic, with \(f\) frequencies fixed by the position of \(\vect{e}_{1}\) relative to the lattice basis, and it is periodic exactly when the line closes, that is when \(T\vect{e}_{1}\in\Gamma\) for some \(T>0\). Rests on Proposition A.552 and Lemma A.548.
Derives Corollary A.553. The Hamiltonian vector field of \(F_{1}\) is \(X_{1}\), whose flow is \(\Phi(t\vect{e}_{1})\) by Equation (A.845). The identification Equation (A.847) carries it to translation by \(t\vect{e}_{1}\) on \(\R^{f}/\Gamma\), and a translation flow on a torus closes if and only if its generator is commensurable with the lattice. Remark 23.33 records what happens to these tori under perturbation, and Theorem 32.51 is the statement that most of them survive.
∎The action variables
Let \(\vect{\lambda}_{1}(\vect{c}),\ldots,\vect{\lambda}_{f}(\vect{c})\) be a basis of the lattice \(\Gamma\) supplied by Proposition A.552, chosen to depend smoothly on \(\vect{c}\) — which is possible because a lattice basis is locally rigid: it is determined by finitely many periods, and a nearby torus has nearby periods, so the basis is continued uniquely. Write \(\gamma_{a}(\vect{c})\) for the closed curve \(t\in[0,1]\mapsto\Phi(t\vect{\lambda}_{a})P\) on the torus \(N(\vect{c})\); its homology class is independent of the base point \(P\), and \(\gamma_{1},\ldots,\gamma_{f}\) generate the first homology of \(N\) because they are the images of the lattice generators under Equation (A.847).
With \(\theta\) the tautological one-form of Equation (24.3), that is \(\theta=p_{a}\dd q^{a}\) in any canonical chart, put
Rests on Proposition A.552, Equation (24.3) and Equation (23.49).
The integral Equation (A.848) depends only on the homology class of \(\gamma_{a}\) in \(N\), not on the representing curve, and it depends smoothly on \(\vect{c}\). Rests on Proposition A.547, Definition A.554 and Theorem A.311.
Derives Lemma A.555. \(N\) is Lagrangian by Proposition A.547, so the restriction of \(\omega=-\dd\theta\) to \(N\) vanishes, that is the restriction of \(\theta\) to \(N\) is a closed one-form. Two homologous closed curves on \(N\) bound a two-chain there, and Stokes' theorem (Theorem A.311) turns the difference of the two circuit integrals into the integral of \(\dd\theta=-\omega\) over that chain, which is zero. Smoothness in \(\vect{c}\) follows because the curves \(\gamma_{a}(\vect{c})\) were chosen to depend smoothly on \(\vect{c}\) and \(\theta\) is smooth.
∎The next statement is the one the chapter's own proof of Theorem 23.30 assumed without proving: that the actions may replace the constants \(\vect{c}\) as labels of the tori. Here it is a theorem, and its content is an exact identity between the Jacobian and the lattice.
With the conventions above,
so \(\vect{c}\longmapsto\vect{I}(\vect{c})\) is a local diffeomorphism and the \(I_{a}\) are independent functions on a neighbourhood of \(N\). Rests on Definition A.554, Proposition A.552 and Proposition 22.18.
Derives Proposition A.556. Let \(S\) be the union of the tori \(N(\vect{c}')\) for \(\vect{c}'\) near \(\vect{c}\), a neighbourhood of \(N\) in \(M\), and on it define
the integral taken along a path from a fixed reference point. Since the restriction of \(\theta\) to each torus is closed (Lemma A.555), \(W\) is well defined up to the periods, and by Equation (A.848) carrying the path once around \(\gamma_{a}\) increases it by
Read \(W\) as a generating function of the second kind (Proposition 22.18) with the constants \(c_{b}\) as new momenta; the conjugate new coordinates are then \(\tau^{b}=\pp W/\pp c_{b}\), and by the second of Hamilton's equations in the new variables (Theorem 22.6) each \(\tau^{b}\) advances at unit rate along the flow of \(F_{b}\) and is constant along the flow of every other \(F_{a}\). In other words \(\tau^{b}\) is the \(b\)-th time coordinate of the action Equation (A.845). Carrying the path once around \(\gamma_{a}\) means applying \(\Phi(\vect{\lambda}_{a})\), so
On the other hand \(\tau^{b}=\pp W/\pp c_{b}\), and the circuit \(\gamma_{a}\) lies inside a single torus, so the constants \(c_{b}\) are fixed along it and the two differentiations may be exchanged under the circuit integral, exactly as in Equation (23.53):
Comparing Equation (A.852) with Equation (A.853) gives Equation (A.849). The determinant of the lattice generators is non-zero because they are linearly independent (Theorem A.550).
∎For one degree of freedom and a bounded orbit of energy \(E\), \(f=1\) and \(\vect{\lambda}_{1}\) is the period \(T\) of the motion, so Equation (A.849) reads \(\dd I/\dd E=T/2\pi=1/\omega\), which is exactly Equation (23.57), derived there from Hamilton's equations alone. For the harmonic oscillator, \(I=E/\omega\) with \(\omega\) independent of \(E\), and both sides are \(1/\omega\). Rests on Proposition A.556 and Equation (23.57).
The angles, and the proof of the theorem
Proof of Theorem A.545. Derives Theorem A.545. Part (1) is Proposition A.547 together with Proposition A.552; part (2) is Corollary A.553. It remains to build the action–angle variables.
By Proposition A.556 the map \(\vect{c}\mapsto\vect{I}\) may be inverted on a neighbourhood, so \(W\) of Equation (A.850) may be regarded as a function \(W(q,I)\) of the position coordinates and the actions. It is again a generating function of the second kind, now with the \(I_{a}\) as new momenta, so
which is Equation (23.50), and the transformation \((q,p)\mapsto(\theta,I)\) is canonical by Proposition 22.18. Repeating the exchange of differentiations that produced Equation (A.853), now with \(I_{a}\) in place of \(c_{b}\),
so each \(\theta^{a}\) is well defined modulo \(2\pi\) and the \(f\) of them are angular coordinates on the torus. The \(I_{a}\) are functions of the \(c_{b}\) alone and are therefore constants of the motion; conversely \(\Ham=F_{1}=c_{1}\) is a function of the \(c_{b}\), hence of the \(I_{a}\), and contains no angle. Theorem 22.6 in the new variables then gives \(\dot{I}_{a}=-\pp\Ham/\pp\theta^{a}=0\), which is Equation (23.51), and \(\dot{\theta}^{a}=\pp\Ham/\pp I_{a}=\omega^{a}(I)\), constant along the motion, which integrates to Equation (23.52). That is part (3), and with it Theorem 23.30 freed of the separability hypothesis under which it was originally proved.
∎One step above deserves to be stated with its caveat rather than glossed. Writing \(W=W(q,I)\) presumes that the \(f\) position coordinates \(q^{a}\) are good coordinates on the torus, i.e. that \(N\to\R^{f}\), \(P\mapsto q(P)\), is a local diffeomorphism. A Lagrangian torus need not project that way everywhere: the projection degenerates on the caustics, which for the one-dimensional bounded orbit are the two turning points, where \(\pp p/\pp q\) blows up and the orbit folds back. The repair is local and is the one Proposition 22.18 already provides: near a point where some subset of the \(q^{a}\) fails, Legendre-transform in those variables and use a generating function of a different kind, whose arguments are the surviving \(q\)'s and the corresponding \(p\)'s. The circuit integrals Equation (A.848) and Equation (A.855) are unaffected, because \(\oint\theta\) is defined on \(N\) without reference to any projection. What the caustics do affect is the semiclassical use of \(W\): each fold crossed contributes a phase, and the count of them around \(\gamma_{a}\) is the Maslov index \(\mu_{a}\) of Equation (23.62). The classical theorem is indifferent to it; the quantization condition is not.
The Liouville–Arnold Theorem discharges the derivation owed at Theorem 23.31 of Hamilton–Jacobi Theory and the Optical–Mechanical Analogy, and Remark 23.32 there states the single imported ingredient, restated above as Theorem A.550. Three connections back into the book are worth making explicit. First, the chapter observes that nothing in it depends on the theorem, because Proposition 23.20 establishes on its own that a separable system is integrable; what the appendix adds is the converse direction of usefulness — action–angle variables now exist for an integrable system whether or not anyone can separate its Hamilton–Jacobi equation, which is Theorem 23.30 with its hypothesis removed. Second, the tori built here are the objects whose fate under perturbation is the subject of Theorem 32.51, and Remark 23.33 is the honest statement that the hypotheses of Theorem 23.31 are met by almost no system. Third, the Lagrangian property established in Proposition A.547 is what Remark 24.13 of Symplectic Geometry of Phase Space has in mind when it lists the Liouville–Arnold tori among the statements of symplectic geometry that carry physical content: they are global objects, and a symplectic manifold has nothing local to say.
Marsden–Weinstein Reduction
This appendix proves Theorem 24.47 of Symplectic Geometry of Phase Space: when a Lie group acts freely and properly on a symplectic manifold by symplectomorphisms with an equivariant momentum map, and \(\mu\) is a regular value of that map, the quotient Equation (24.31) of the level set by the isotropy group of \(\mu\) is again a symplectic manifold, of the dimension Equation (24.32), and every invariant Hamiltonian descends to it with its dynamics intact.
The chapter states the geometric core of the argument and does not prove it: the tangent space to the level set of the momentum map is the symplectic orthogonal of the group orbit (Lemma 24.46, which is proved there), and the degenerate directions of the restricted form turn out to be exactly the orbit directions of the isotropy group \(G_{\mu}\), so that quotienting by that group removes the degeneracy and removes nothing else. That sentence is The degenerate directions of the restricted form below. Everything around it — that the level set is a manifold, that the quotient is one, that the descended form is well defined, closed and nondegenerate — is what makes the theorem long, and it is written out here.
The paragraph following the proof link in Symplectic Geometry of Phase Space names one input as quoted rather than proved: the quotient-manifold theorem, that a free and proper action of a Lie group on a manifold has a smooth quotient with the projection a submersion, described there as differential topology belonging in Differentiable Manifolds, Tensors, and Curvature, “which does not yet carry it”. That sentence is now out of date, and the correction is worth making rather than leaving as a discrepancy between two pages of the same book: Differentiable Manifolds, Tensors, and Curvature carries the statement as Theorem 13.67, with its own proof, and this section therefore cites it rather than quoting it. Nothing in Marsden–Weinstein Reduction is imported from outside the treatise. The role of that theorem is unchanged and remains confined to producing the smooth structure on \(M_{\mu}\): the symplectic content — existence, uniqueness, closedness and nondegeneracy of the reduced form, and the dimension count — rests on Lemma 24.46 and on Two facts of symplectic linear algebra alone.
Setting and statement
Throughout, \((M,\omega)\) is a symplectic manifold (Definition 24.8), \(G\) is a Lie group with Lie algebra \(\mathfrak{g}\) acting smoothly on \(M\) by symplectomorphisms (Definition 24.17), and \(\xi_{M}\) denotes the vector field generating the action of \(\xi\in\mathfrak{g}\), so that \(\xi_{M}(x)=\frac{\dd}{\dd t}\bigl|_{t=0}\exp(t\xi)\cdot x\). The action is assumed free and proper in the sense of Definition 13.66. The momentum map \(\vect{J}:M\longrightarrow\mathfrak{g}^{*}\) is the one of Definition 24.43, defined by Equation (24.28), and the interior product is written \(\iota_{X}\), as in Symplectic Geometry of Phase Space; the same operation is written \(i_{X}\) in Differentiable Manifolds, Tensors, and Curvature.
The coadjoint action of \(G\) on \(\mathfrak{g}^{*}\) is
a left action, and its infinitesimal form is \(\ad^{*}_{\xi}=\frac{\dd}{\dd t}\bigl|_{t=0} \mathrm{Ad}^{*}_{\exp(t\xi)}\), so that
A momentum map is equivariant if \(\vect{J}(g\cdot x)=\mathrm{Ad}^{*}_{g}\vect{J}(x)\) for every \(g\) and \(x\). For \(\mu\in\mathfrak{g}^{*}\) write \(G_{\mu}=\set{g\in G\mid\mathrm{Ad}^{*}_{g}\mu=\mu}\) for its isotropy group, a closed subgroup, and \(\mathfrak{g}_{\mu}\) for the Lie algebra of \(G_{\mu}\). Rests on Definitions 4.20, 14.2 and 24.43.
Two sign conventions circulate for Equation (A.856), differing by \(g\leftrightarrow g^{-1}\); with the one chosen here the coadjoint action is a left action and equivariance reads \(\vect{J}\circ g=\mathrm{Ad}^{*}_{g}\circ\vect{J}\) without an inverse. Sources that define \(\mathrm{Ad}^{*}\) by \(\langle\mathrm{Ad}^{*}_{g}\mu,\eta\rangle=\langle\mu, \mathrm{Ad}_{g}\eta\rangle\) write the same statement with \(g^{-1}\) on the right, and the two agree.
This section is pure geometry and the only dimensional statement it needs is the one Remark 24.11 already fixes: \(\omega\) carries \(\mathrm{J}\,\mathrm{s}\), and a component \(J_{\xi}\) of the momentum map carries the SI unit of the conserved quantity it generates — so \(\vect{J}\) is not a single dimensioned object but a covector whose \(f\) components may carry \(f\) different units, exactly as the lattice of The Liouville–Arnold Theorem does. The reduced form \(\omega_{\mu}\) inherits \(\mathrm{J}\,\mathrm{s}\) from \(\omega\), because it is defined by pullback.
Let \(G\) act freely and properly on \((M,\omega)\) by symplectomorphisms with an equivariant momentum map \(\vect{J}\), and let \(\mu\in\mathfrak{g}^{*}\). Write \(N=\vect{J}^{-1}(\mu)\), let \(i:N\hookrightarrow M\) be the inclusion and \(\pi:N\longrightarrow M_{\mu}=N/G_{\mu}\) the projection. Then
-
every value of \(\vect{J}\) is regular, and \(N\) is a closed embedded submanifold of \(M\) of dimension \(\dim M-\dim G\) with \(T_{x}N=\ker\dd\vect{J}_{x}\);
-
\(G_{\mu}\) acts freely and properly on \(N\), and \(M_{\mu}\) is a smooth manifold with \(\pi\) a surjective submersion and
\begin{equation}\tag{A.858} \dim M_{\mu}=\dim M-\dim G-\dim G_{\mu}\ec \end{equation}which is Equation (24.32);
-
there is exactly one two-form \(\omega_{\mu}\) on \(M_{\mu}\) with \(\pi^{*}\omega_{\mu}=i^{*}\omega\), and it is symplectic;
-
every \(G\)-invariant \(H\in C^{\infty}(M)\) descends to \(H_{\mu}\in C^{\infty}(M_{\mu})\) with \(H_{\mu}\circ\pi=H\circ i\); the flow of \(X_{H}\) preserves \(N\); and \(\pi\) carries it to the flow of \(X_{H_{\mu}}\) on \((M_{\mu},\omega_{\mu})\).
Rests on Lemma 24.46, Theorem 13.67 and Definition A.561.
Two facts of symplectic linear algebra
Let \((V,\Omega)\) be a finite-dimensional real vector space with a nondegenerate antisymmetric bilinear form, and for a subspace \(W\subseteq V\) let \(W^{\Omega}=\set{v\in V\mid\Omega(v,w)=0\ \text{for all}\ w\in W}\). Then
Rests on Definition 24.2 and Theorem 5.38.
Derives Lemma A.564. Nondegeneracy makes \(\flat:v\longmapsto\Omega(v,\cdot)\) an injective, hence bijective, linear map \(V\longrightarrow V^{*}\). By definition \(W^{\Omega}=\flat^{-1}\bigl(W^{\circ}\bigr)\), where \(W^{\circ}\subseteq V^{*}\) is the annihilator of \(W\); and \(\dim W^{\circ}=\dim V-\dim W\) by rank–nullity (Theorem 5.38) applied to the restriction map \(V^{*}\longrightarrow W^{*}\), which is surjective. That is the first identity. For the second, \(W\subseteq(W^{\Omega})^{\Omega}\) is immediate from antisymmetry, and applying the first identity twice gives \(\dim(W^{\Omega})^{\Omega}=\dim V-\dim W^{\Omega}=\dim W\), so the inclusion is an equality.
∎The level set is a manifold
If the action is free then \(\dd\vect{J}_{x}\) is surjective at every \(x\in M\). Consequently every \(\mu\in\mathfrak{g}^{*}\) is a regular value, \(N=\vect{J}^{-1}(\mu)\) is a closed embedded submanifold of dimension \(\dim M-\dim G\), and \(T_{x}N=\ker\dd\vect{J}_{x} =\bigl(T_{x}(G\cdot x)\bigr)^{\omega}\). Rests on Lemma 24.46, Lemma A.564 and Theorem 13.59.
Derives Proposition A.565. Fix \(x\) and write \(W=T_{x}(G\cdot x)\), the span of the values \(\xi_{M}(x)\), \(\xi\in\mathfrak{g}\). Freeness makes the map \(\xi\mapsto\xi_{M}(x)\) injective: if \(\xi_{M}(x)=0\) then the curve \(t\mapsto\exp(t\xi)\cdot x\) has vanishing velocity and, being an integral curve of \(\xi_{M}\) through \(x\), is constant by the uniqueness clause of Theorem 13.125; so \(\exp(t\xi)\cdot x=x\) for all \(t\), whence \(\exp(t\xi)=e\) by freeness and \(\xi=0\). Hence \(\dim W=\dim\mathfrak{g}=\dim G\).
By Lemma 24.46, \(\ker\dd\vect{J}_{x}=W^{\omega}\), so Equation (A.859) gives \(\dim\ker\dd\vect{J}_{x}=\dim M-\dim W=\dim M-\dim G\), and by rank–nullity the rank of \(\dd\vect{J}_{x}\) is \(\dim G =\dim\mathfrak{g}^{*}\): the differential is surjective. The regular value theorem Theorem 13.59 then makes \(N=\vect{J}^{-1}(\mu)\) an embedded submanifold of dimension \(\dim M-\dim G\), closed because it is the preimage of a point under a continuous map, with tangent space the kernel.
∎\(G_{\mu}\) maps \(N\) into itself, and its action on \(N\) is free and proper. Hence \(M_{\mu}=N/G_{\mu}\) carries exactly one smooth structure making \(\pi:N\longrightarrow M_{\mu}\) a surjective submersion; it is Hausdorff and second countable, the fibres of \(\pi\) are the \(G_{\mu}\)-orbits, and Equation (A.858) holds. Rests on Definition A.561, Theorem 13.67 and Proposition A.565.
Derives Proposition A.566. For \(g\in G_{\mu}\) and \(x\in N\), equivariance gives \(\vect{J}(g\cdot x)=\mathrm{Ad}^{*}_{g}\vect{J}(x) =\mathrm{Ad}^{*}_{g}\mu=\mu\), so \(g\cdot x\in N\). The restricted action is free because the \(G\)-action is, and it is proper: \(G_{\mu}\) is a closed subgroup of \(G\) and \(N\) is a closed subset of \(M\), so the preimage of a compact set under \((g,x)\mapsto(g\cdot x,x)\) on \(G_{\mu}\times N\) is a closed subset of the corresponding preimage on \(G\times M\), hence compact. Theorem 13.67 applied to this action supplies the smooth structure, the submersion, the separation properties and the fibres, together with the dimension \(\dim M_{\mu}=\dim N-\dim G_{\mu}\), which with Proposition A.565 is Equation (A.858).
∎The degenerate directions of the restricted form
Write \(\omega_{N}=i^{*}\omega\) for the restriction of \(\omega\) to \(N\). It is a closed two-form, but it is not symplectic: it has a radical, and identifying that radical is the heart of the theorem.
Let \(x\in N\), so that \(\vect{J}(x)=\mu\). Then for every \(\xi\in\mathfrak{g}\)
and consequently \(\xi_{M}(x)\in\ker\dd\vect{J}_{x}\) if and only if \(\xi\in\mathfrak{g}_{\mu}\). Rests on Definition A.561, Proposition A.565 and Theorem 13.125.
Derives Lemma A.567. Apply equivariance along the one-parameter subgroup \(g(t)=\exp(t\xi)\) and differentiate at \(t=0\). The left-hand side of \(\vect{J}(g(t)\cdot x)=\mathrm{Ad}^{*}_{g(t)}\mu\) differentiates by the chain rule to \(\dd\vect{J}_{x}(\xi_{M}(x))\), since \(\frac{\dd}{\dd t}\bigl|_{0}g(t)\cdot x=\xi_{M}(x)\) by definition; the right-hand side differentiates to \(\ad^{*}_{\xi}\mu\) by Definition A.561. That is Equation (A.860).
For the second claim, \(\ad^{*}_{\xi}\mu=0\) certainly holds when \(\xi\in\mathfrak{g}_{\mu}\), since then \(\mathrm{Ad}^{*}_{\exp(t\xi)}\mu=\mu\) for all \(t\). Conversely suppose \(\ad^{*}_{\xi}\mu=0\) and set \(\nu(t)=\mathrm{Ad}^{*}_{\exp(t\xi)}\mu\). Differentiating Equation (A.856) in \(t\) and using the group law gives the linear differential equation \(\dot{\nu}(t)=\ad^{*}_{\xi}\nu(t)\) with \(\nu(0)=\mu\). The constant function \(\nu\equiv\mu\) solves it, because \(\ad^{*}_{\xi}\mu=0\), and the solution of a linear equation with given initial value is unique (Theorem 9.8); so \(\nu(t)=\mu\) for all \(t\), that is \(\exp(t\xi)\in G_{\mu}\) for all \(t\), that is \(\xi\in\mathfrak{g}_{\mu}\).
∎For every \(x\in N\),
Derives Theorem A.568. Write \(W=T_{x}(G\cdot x)\) as before. By Proposition A.565 and Equation (24.30), \(T_{x}N=\ker\dd\vect{J}_{x}=W^{\omega}\). Since \(\omega_{N}\) is the restriction of \(\omega\), the radical of \(\omega_{N}\) at \(x\) is \(T_{x}N\cap(T_{x}N)^{\omega}\), the second symplectic orthogonal being taken in the symplectic vector space \((T_{x}M,\omega_{x})\). Hence, using the second identity of Equation (A.859),
A vector of \(T_{x}(G\cdot x)\) is \(\xi_{M}(x)\) for exactly one \(\xi\in\mathfrak{g}\), by the injectivity established in Proposition A.565, and it lies in \(\ker\dd\vect{J}_{x}\) precisely when \(\xi\in\mathfrak{g}_{\mu}\), by Lemma A.567. The intersection is therefore \(\set{\xi_{M}(x)\mid\xi\in\mathfrak{g}_{\mu}} =T_{x}(G_{\mu}\cdot x)\).
∎The radical of a closed two-form of constant rank is always an involutive distribution, so it is integrable by Theorem 13.133 and the manifold is foliated by its leaves. The check is two lines with Cartan's magic formula: if \(\iota_{X}\omega_{N}=0\) and \(\iota_{Y}\omega_{N}=0\) then \(\mathcal{L}_{X}\omega_{N} =\dd\,\iota_{X}\omega_{N}+\iota_{X}\dd\omega_{N}=0\) by Equation (13.276) and \(\dd\omega_{N}=i^{*}\dd\omega=0\), and then Equation (13.277) gives \(\iota_{\comm{X}{Y}}\omega_{N} =\mathcal{L}_{X}\iota_{Y}\omega_{N} -\iota_{Y}\mathcal{L}_{X}\omega_{N}=0\), so \(\comm{X}{Y}\) lies in the radical too. What Theorem A.568 adds is that the leaves of that foliation are exactly the \(G_{\mu}\)-orbits. Without it one could still divide \(N\) by the characteristic foliation and get a symplectic quotient, but the quotient would be a leaf space with no reason to be Hausdorff or to be a manifold at all. Identifying the leaves as orbits of a group acting freely and properly is what lets Theorem 13.67 supply the smooth structure — and it is also what makes the reduced space computable in practice, since a physicist knows the symmetry group and does not know the foliation.
Descent of the form
There is exactly one two-form \(\omega_{\mu}\) on \(M_{\mu}\) with \(\pi^{*}\omega_{\mu}=\omega_{N}\), it is smooth, and it is closed and nondegenerate. Rests on Theorem A.568, Proposition A.566 and Theorem 24.16.
Derives Proposition A.570. Definition. Let \([x]\in M_{\mu}\) and \(u,v\in T_{[x]}M_{\mu}\). Since \(\pi\) is a submersion, \(\dd\pi_{x}\) is onto, so there are \(\tilde{u},\tilde{v}\in T_{x}N\) with \(\dd\pi_{x}\tilde{u}=u\) and \(\dd\pi_{x}\tilde{v}=v\); set
This is independent of the lifts. The fibres of \(\pi\) are the \(G_{\mu}\)-orbits (Proposition A.566), so \(\ker\dd\pi_{x}=T_{x}(G_{\mu}\cdot x)\), which is exactly the radical of \((\omega_{N})_{x}\) by Theorem A.568; changing a lift changes it by an element of that radical, which \(\omega_{N}\) does not see.
Independence of the point in the fibre. Let \(g\in G_{\mu}\) and replace \(x\) by \(g\cdot x\), lifting through \(\dd(g\cdot)_{x}\tilde{u}\) and \(\dd(g\cdot)_{x}\tilde{v}\), which are legitimate lifts because \(\pi\circ g=\pi\). Since \(G\) acts by symplectomorphisms, \(g^{*}\omega=\omega\), and \(g\) maps \(N\) to \(N\), so \(g^{*}\omega_{N}=\omega_{N}\) and the two evaluations agree.
Smoothness. A submersion admits local smooth sections: near any \([x]\) choose \(s\) with \(\pi\circ s=\id\); then \(\omega_{\mu}=s^{*}\omega_{N}\) on that neighbourhood by Equation (A.863), a pullback of a smooth form along a smooth map.
Uniqueness. If \(\pi^{*}\omega_{\mu}=\pi^{*}\omega_{\mu}'\) then, \(\dd\pi_{x}\) being surjective at every \(x\) and \(\pi\) being surjective, the two forms agree on every pair of tangent vectors at every point.
Closedness. \(\pi^{*}\dd\omega_{\mu}=\dd\pi^{*}\omega_{\mu} =\dd\,i^{*}\omega=i^{*}\dd\omega=0\), the exterior derivative commuting with pullback; and \(\pi^{*}\) is injective on forms because \(\dd\pi\) is surjective. Hence \(\dd\omega_{\mu}=0\).
Nondegeneracy. Suppose \((\omega_{\mu})_{[x]}(u,\cdot)=0\). Lift \(u\) to \(\tilde{u}\in T_{x}N\); then \((\omega_{N})_{x}(\tilde{u},\tilde{v})=0\) for every \(\tilde{v}\in T_{x}N\), because every \(v\) is \(\dd\pi_{x}\tilde{v}\). So \(\tilde{u}\) lies in the radical, which is \(\ker\dd\pi_{x}\), and therefore \(u=\dd\pi_{x}\tilde{u}=0\).
∎Descent of the dynamics
Let \(H\in C^{\infty}(M)\) be \(G\)-invariant. Then \(H\circ i\) is \(G_{\mu}\)-invariant and descends to a smooth \(H_{\mu}\) on \(M_{\mu}\); the flow of \(X_{H}\) preserves \(N\); and \(\dd\pi_{x}\bigl(X_{H}(x)\bigr)=X_{H_{\mu}}([x])\) for every \(x\in N\). Rests on Theorem 24.45, Proposition A.570 and Equation (24.5).
Derives Proposition A.571. Invariance under \(G\) implies invariance under the subgroup \(G_{\mu}\), and by the last clause of Theorem 13.67 a function on \(M_{\mu}\) is smooth exactly when its composition with \(\pi\) is; the \(G_{\mu}\)-invariant function \(H\circ i\) therefore defines a smooth \(H_{\mu}\) with \(H_{\mu}\circ\pi=H\circ i\).
That the flow of \(X_{H}\) preserves \(N\) is Noether's theorem in the form Theorem 24.45: \(G\) acts by symplectomorphisms and preserves \(H\), so \(\vect{J}\) is constant along the flow, and a trajectory starting on \(\vect{J}^{-1}(\mu)\) stays there. In particular \(X_{H}(x)\in T_{x}N\) for \(x\in N\).
For the last claim, let \(v\in T_{[x]}M_{\mu}\) and lift it to \(\tilde{v}\in T_{x}N\). Using Equation (A.863), then Equation (24.5), then \(H_{\mu}\circ\pi=H\circ i\):
Since \(v\) was arbitrary and \(\omega_{\mu}\) is nondegenerate, the vector \(\dd\pi_{x}X_{H}(x)\) is the unique one satisfying Equation (24.5) for \(H_{\mu}\), that is \(X_{H_{\mu}}([x])\).
∎Proof of Theorem A.563. Derives Theorem A.563. Part (1) is Proposition A.565, part (2) is Proposition A.566, part (3) is Proposition A.570 and part (4) is Proposition A.571. Since \(M_{\mu}=\vect{J}^{-1}(\mu)/G_{\mu}\) is Equation (24.31) and Equation (A.858) is Equation (24.32), this is Theorem 24.47.
∎The example the chapter names
Take \(M=T^{*}\R^{3}\) with \(\omega=\dd q^{i}\wedge\dd p_{i}\) and \(G=\SO(3)\) acting by \((\vect{q},\vect{p})\mapsto(R\vect{q},R\vect{p})\). This is an action by symplectomorphisms, and Example 24.44 computes its momentum map: \(\vect{J}=\vect{q}\times\vect{p}=\vect{L}\), the angular momentum, in \(\mathrm{J}\,\mathrm{s}\).
The action is not free on all of \(M\) — a rotation about \(\vect{q}\) fixes any point with \(\vect{p}\) parallel to \(\vect{q}\), and every rotation fixes the origin — so the hypotheses hold only on the open, \(\SO(3)\)-invariant subset \(M^{\times}=\set{(\vect{q},\vect{p})\mid\vect{q}\times\vect{p} \neq\vect{0}}\), where a rotation fixing two independent vectors is the identity. That restriction is not a technicality to be waved through: it is exactly the collinear and radial motions, which the reduced picture cannot describe and which Section 27.4 treats separately.
On \(M^{\times}\) fix \(\mu=\vect{L}\) with \(\ell=\abs{\vect{L}}>0\). The isotropy group of \(\mu\) under the coadjoint action of \(\SO(3)\) — which for \(\SO(3)\) is the ordinary rotation action on \(\R^{3}\), the algebra carrying an invariant inner product — is the group \(\SO(2)\) of rotations about the axis of \(\vect{L}\), of dimension \(1\). So Equation (A.858) gives
and the reduced space is the two-dimensional radial phase space \((r,p_{r})\). For \(H=\abs{\vect{p}}^{2}/2m+V(r)\), resolving \(\vect{p}\) into its radial and transverse parts on \(\vect{J}^{-1}(\mu)\) gives \(\abs{\vect{p}}^{2}=p_{r}^{2}+\ell^{2}/r^{2}\), so
which is the effective potential of Equation (27.32), and the centrifugal term \(\ell^{2}/2mr^{2}\) carries \(\mathrm{J}\) as it must. This is the identification Remark 24.48 asserts: the effective potential is the reduced Hamiltonian at the fixed value of the angular momentum, and the disappearance of the two angular coordinates is not a trick of the coordinate system but the reduction of six dimensions to two. Rests on Theorem A.563, Example 24.44 and Equation (27.32).
If \(G\) is abelian the coadjoint action is trivial, so \(G_{\mu}=G\) for every \(\mu\) and Equation (A.858) reads \(\dim M_{\mu}=\dim M-2\dim G\): two dimensions are lost per generator. For \(G=\R\) acting by translation in a cyclic coordinate \(q^{1}\), the momentum map is \(p_{1}\), fixing \(\mu\) freezes that momentum, and dividing by the group deletes \(q^{1}\) — the pair \((q^{1},p_{1})\) disappears together, which is precisely the elementary manoeuvre Remark 24.48 says every physicist performs without naming it. Rests on Theorem A.563 and Remark 24.48.
What is quoted here
Nothing. The inputs are Lemma 24.46, proved in Symplectic Geometry of Phase Space; the regular value theorem Theorem 13.59, the quotient-manifold theorem Theorem 13.67, Frobenius' theorem Theorem 13.133 and the Cartan identities Equation (13.276) and Equation (13.277), all proved in Differentiable Manifolds, Tensors, and Curvature; rank–nullity Theorem 5.38 from Linear Algebra and Representation Theory; and the uniqueness of solutions of an ordinary differential equation, Theorem 9.8. Remark A.560 records that the chapter's own prose still describes the quotient-manifold theorem as owed by Part II, which it no longer is.
The three hypotheses divide the work cleanly, and it is worth recording which of them fails first in practice.
Freeness is used twice: in Proposition A.565, to make every value of \(\vect{J}\) regular and \(N\) a manifold, and in Proposition A.566, as a hypothesis of Theorem 13.67. Properness is used once, also through Theorem 13.67, and what it buys is that the quotient topology is Hausdorff — Example 13.68 exhibits a free but improper action whose orbit space is not. Equivariance is used once, in Lemma A.567, and it is what makes \(\mathfrak{g}_{\mu}\) rather than \(\mathfrak{g}\) appear in Equation (A.861); without it the radical would not be the tangent space to an orbit of anything.
When freeness fails — which, as Example A.572 shows, is the normal state of affairs at the symmetric configurations — \(M_{\mu}\) is not a manifold but a stratified space, a union of manifolds of different dimensions glued along their boundaries. That is singular reduction, and it is not treated in this treatise. The honest summary is that the theorem proved here describes the generic stratum and says nothing about the symmetric configurations, which for the central-force problem are exactly the collinear orbits.
The theorem is Marsden and Weinstein's, from Reduction of symplectic manifolds with symmetry (Reports on Mathematical Physics 5, 1974, 121–130). That paper has no entry in this treatise's bibliography — Symplectic Geometry of Phase Space records the omission in its own header, together with those of Abraham–Marsden and Arnold — so the attribution is made in words, on the footing of Darboux's memoir in Remark A.81. Nothing above rests on it: the derivation is carried out in full from results proved in this book.
Marsden–Weinstein Reduction discharges the derivation owed at Theorem 24.47 of Symplectic Geometry of Phase Space. Two threads run back into the book from here. Remark 24.48 lists four familiar manoeuvres — dropping a cyclic pair, passing to the centre-of-mass frame, reducing the two-body problem to the radial one, and using the effective potential — and says that each is this theorem applied to a particular group; Example A.572 and Example A.573 above verify two of them, including the identification of Equation (27.32) as the reduced Hamiltonian. And the same remark points at the Lie–Poisson bracket Definition 24.40 as reduction with the whole group divided out, which is why the free rigid body of Proposition 24.41 lives on three variables rather than the six a cotangent bundle would demand: the dimension count is Equation (A.858) once more, with \(M=T^{*}\SO(3)\) of dimension six, \(G=\SO(3)\) of dimension three, and a generic \(G_{\mu}=\SO(2)\) removing one further dimension to leave the two-dimensional coadjoint orbit — a sphere of constant \(\abs{\vect{L}}\) — on which the Euler equations Equation (24.27) run.
The Stone–von Neumann Theorem
This appendix proves Theorem 25.34 of The Poisson Algebra and the Canonical Bridge to Quantum Mechanics: for finitely many degrees of freedom an irreducible representation of the canonical commutation relations in Weyl's exponentiated form is unitarily equivalent to the Schrödinger representation on \(L^{2}(\R^{f})\), so that the quantum kinematics of a system with finitely many canonical pairs is fixed, up to equivalence, by the classical bracket algebra alone.
Nothing in the argument is physics, and the statement it proves stands on the books twice. Theorem 12.114 of Hilbert Spaces states it with the care about domains and irreducibility that the subject needs, and Remarks 12.115 and 25.35 both record that the proof was owed and that it belongs beside the Hilbert-space material rather than in a chapter of mechanics. This section discharges both debts at once. It is therefore written throughout in the notation of Definition 12.109 — dimensionful parameters, generators \(\vect{Q}\) and \(\vect{P}\), the relation in the form Equation (12.71) — so that moving it into Hilbert Spaces would change nothing but its address. Two by-products are worth naming in advance: the argument produces the vacuum vector explicitly, and it proves that the Schrödinger system is irreducible, a fact that Example 12.110 currently takes from [Reed:1972] without proof.
Statement, conventions and units
Throughout, \(\mathcal{H}\) is a separable complex Hilbert space (Definition 12.2), \(f\) is a finite positive integer, and \((U,V)\) is a Weyl system of \(f\) degrees of freedom in the sense of Definition 12.109: two families of strongly continuous unitary groups
whose members belonging to different degrees of freedom commute, and which satisfy
The components of \(\vect{Q}\) carry \(\mathrm{m}\) and those of \(\vect{P}\) carry \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), so the parameter \(\vect{\alpha}\) carries \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) and \(\vect{\beta}\) carries \(\mathrm{m}\); both \(\vect{\alpha}\cdot\vect{Q}/\hbar\) and \(\vect{\alpha}\cdot\vect{\beta}/\hbar\) are then pure numbers, since \(\hbar\) carries \(\mathrm{J}\,\mathrm{s}\). Every exponent written below is dimensionless as it stands, and \(\hbar\) is never set to unity. Equation (25.39) of The Poisson Algebra and the Canonical Bridge to Quantum Mechanics writes the same relation with parameters of reciprocal dimension, \(a\) in \(/\mathrm{m}\) and \(b\) in \(\mathrm{s}/\mathrm{kg}/\mathrm{m}\); the substitution \(a=\alpha/\hbar\), \(b=-\beta/\hbar\) turns Equation (A.868) into Equation (25.39) and back, and no statement below depends on which of the two is used. Finally, a fixed length \(s>0\), in \(\mathrm{m}\), enters in The Gaussian average as the width of a Gaussian weight; it is arbitrary, and Remark A.588 records what changes when it is changed.
Let \((U,V)\) be a Weyl system of \(f<\infty\) degrees of freedom on a separable Hilbert space \(\mathcal{H}\neq\set{0}\).
-
If the system acts irreducibly (Definition 12.90) there is a unitary \(T:\mathcal{H}\longrightarrow L^{2}(\R^{f})\) carrying it to the Schrödinger system Equation (12.72), and \(T\) is unique up to a factor of modulus one.
-
Without the irreducibility hypothesis, \(\mathcal{H}\) is an orthogonal direct sum (Definition 12.84) of at most countably many closed subspaces, each invariant under the system and each carrying a copy of the Schrödinger system.
Rests on Definition 12.109, Definition 12.90 and Example 12.110.
The proof occupies the rest of this section: The Weyl operators and their composition law assembles the Weyl operators and their composition law, Two analytic tools the two analytic tools, The Gaussian average the Gaussian average and the single identity on which everything turns, The vacuum vector the vacuum vector, The intertwining unitary the unitary, and The Schrödinger system the Schrödinger system itself.
The Weyl operators and their composition law
For \(z=(\vect{\alpha},\vect{\beta})\in\R^{2f}\) put
and let
be the standard symplectic form on \(\R^{2f}\), of dimension \(\mathrm{J}\,\mathrm{s}\). Rests on Definition 12.109, Equation (12.71) and Definition 12.64.
Every \(W(z)\) is unitary, \(W(0)=\identity\), and
Consequently \(W(z)W(z')=\ee^{\ii\sigma(z,z')/\hbar}W(z')W(z)\), and the family \(\set{W(z)}_{z\in\R^{2f}}\) is self-adjoint in the sense of Definition 12.90. Moreover \(W(\vect{\alpha},\vect{0})=U(\vect{\alpha})\) and \(W(\vect{0},\vect{\beta})=V(\vect{\beta})\), so a closed subspace is invariant under every \(W(z)\) if and only if it is invariant under every \(U(\vect{\alpha})\) and every \(V(\vect{\beta})\). Rests on Definition A.580, Equation (12.71) and Definition 12.64.
Derives Lemma A.581. Unitarity and \(W(0)=\identity\) are immediate from Equation (A.869), a phase times a product of unitaries. Rearranging Equation (A.868) gives the form in which the relation will be used,
Now compute, using Equation (A.872) once to move \(U(\vect{\alpha}')\) to the left of \(V(\vect{\beta})\), and then the group laws of \(U\) and of \(V\) separately:
On the other hand \(W(z+z')=\ee^{-\ii(\vect{\alpha}+\vect{\alpha}')\cdot (\vect{\beta}+\vect{\beta}')/2\hbar} U(\vect{\alpha}+\vect{\alpha}')V(\vect{\beta}+\vect{\beta}')\), so the ratio of Equation (A.873) to \(W(z+z')\) is a phase whose exponent, multiplied by \(\hbar/\ii\), is
every term in \(\vect{\alpha}\cdot\vect{\beta}\) and in \(\vect{\alpha}'\cdot\vect{\beta}'\) cancelling. That is Equation (A.871). Since \(\sigma(z,-z)=0\) we get \(W(z)W(-z)=W(0)=\identity\), so \(W(z)^{-1}=W(-z)\), and unitarity turns the inverse into the adjoint. Interchanging \(z\) and \(z'\) in Equation (A.871) and using \(\sigma(z',z)=-\sigma(z,z')\) gives the commutation form. Setting \(\vect{\beta}=\vect{0}\) or \(\vect{\alpha}=\vect{0}\) in Equation (A.869) recovers \(U\) and \(V\); conversely \(W(z)\) is a phase times a product of one \(U\) and one \(V\), so the two families have the same invariant subspaces.
∎It is convenient to remove the units once and for all. Fix a length \(s>0\) and put
where \(\sigma_{0}(\zeta,\zeta') =\vect{a}\cdot\vect{b}'-\vect{a}'\cdot\vect{b}\) is the same form built from the dimensionless coordinates, and \(\dd^{f}\alpha\,\dd^{f}\beta=\hbar^{f}\,\dd^{2f}\zeta\). Writing \(W(\zeta)\) for \(W(z)\) under this substitution, Equation (A.871) becomes
It is also useful to write \(\sigma_{0}(\zeta,\zeta') =\zeta\cdot J\zeta'\), where \(J(\vect{a},\vect{b}) =(\vect{b},-\vect{a})\) is a real orthogonal map of \(\R^{2f}\) with \(J^{2}=-\identity\); then \(\zeta\cdot J\zeta=0\) and \(\abs{J\zeta}=\abs{\zeta}\) for every \(\zeta\).
For every \(x\in\mathcal{H}\) the map \(\zeta\longmapsto W(\zeta)x\) is continuous from \(\R^{2f}\) to \(\mathcal{H}\), and \(\norm{W(\zeta)x}=\norm{x}\). Rests on Lemma A.581 and Definition 12.64.
Derives Lemma A.582. The norm statement is unitarity. For continuity it is enough, by Equation (A.869) and continuity of the phase, to show that \(\vect{\alpha}\mapsto U(\vect{\alpha})x\) and \(\vect{\beta}\mapsto V(\vect{\beta})x\) are continuous. Take \(U\); the argument for \(V\) is identical. By hypothesis the \(f\) one-parameter groups \(\alpha_{j}\mapsto U(\alpha_{j}\vect{e}_{j})\) commute and each is strongly continuous, and \(U(\vect{\alpha})=U(\alpha_{1}\vect{e}_{1})\cdots U(\alpha_{f}\vect{e}_{f})\). Telescoping the difference of two such products and using that every factor is an isometry,
with \(y_{j}=U(\alpha_{j+1}'\vect{e}_{j+1})\cdots U(\alpha_{f}'\vect{e}_{f})x\) a fixed vector for each \(j\) once \(\vect{\alpha}'\) is fixed. Each summand tends to \(0\) as \(\alpha_{j}\to\alpha_{j}'\) by strong continuity of the \(j\)-th group, which is continuity at \(\vect{\alpha}'\).
∎Two analytic tools
Let \(G:\R^{2f}\longrightarrow\C\) be continuous with \(\int_{\R^{2f}}\abs{G}<\infty\). Then for every \(x\in\mathcal{H}\) the \(\mathcal{H}\)-valued integral
exists as the limit of the integrals over the balls \(\abs{\zeta}\leq R\); \(\Pi_{G}\) is a bounded operator with \(\norm{\Pi_{G}}\leq\int\abs{G}\); and
If \(G\) and \(G'\) both satisfy the hypotheses then
Finally, a closed subspace invariant under every \(W(\zeta)\) is invariant under \(\Pi_{G}\). Rests on Lemma A.582, Lemma A.254 and Proposition 12.4.
Derives Lemma A.583. On each ball the integrand is a continuous \(\mathcal{H}\)-valued function of compact support, so the Riemann construction of Lemma A.254 applies verbatim in \(2f\) variables — it uses only uniform continuity on a compact set and completeness of \(\mathcal{H}\), neither of which cares how many variables there are. For \(R'>R\) the difference of the two truncated integrals is bounded in norm by \(\norm{x}\int_{R<\abs{\zeta}\leq R'}\abs{G}\), which tends to zero as \(R\to\infty\) because \(\abs{G}\) is integrable; the truncations therefore form a Cauchy net and converge. The bound on \(\norm{\Pi_{G}}\) and the first identity in Equation (A.879) pass to the limit from Equation (A.471). For the adjoint, use that identity twice together with \(W(\zeta)^{\dagger}=W(-\zeta)\): \(\braket{\Pi_{G}w}{x}=\overline{\braket{x}{\Pi_{G}w}} =\int\overline{G(\zeta)}\braket{W(\zeta)w}{x}\dd^{2f}\zeta =\int\overline{G(\zeta)}\braket{w}{W(-\zeta)x}\dd^{2f}\zeta\), and substituting \(\zeta\to-\zeta\) identifies the result as \(\braket{w}{\Pi_{G^{\sharp}}x}\). Invariance of a closed invariant subspace is inherited from the truncated integrals, whose Riemann sums lie in it, the subspace being closed.
For Equation (A.880), take \(w,x\in\mathcal{H}\) and expand \(\braket{w}{\Pi_{G}\Pi_{G'}x}\) by the first identity of Equation (A.879), applied once to \(\Pi_{G}\) acting on the vector \(\Pi_{G'}x\) and once inside. The result is the scalar double integral
absolutely convergent because the matrix element is bounded by \(\norm{w}\norm{x}\) and \(G\), \(G'\) are integrable. Substituting Equation (A.876) and changing variables to \(u=\zeta+\zeta'\) at fixed \(\zeta\) — a translation, which preserves \(\dd^{2f}\zeta'\) — and then exchanging the order of integration, which absolute convergence and Equation (7.112) on an exhausting sequence of boxes permit, turns Equation (A.881) into \(\int G''(u)\braket{w}{W(u)x}\dd^{2f}u\), using \(\sigma_{0}(\zeta,u-\zeta)=\sigma_{0}(\zeta,u)\). Since \(w\) was arbitrary this is \(\braket{w}{\Pi_{G''}x}\), and \(G''\) is continuous and integrable because \(\abs{G''}\leq\abs{G}\ast\abs{G'}\).
∎The second tool is the injectivity of the Fourier transform in several variables. Corollary 17.33 states it on the line; the proof of Theorem 17.32 carries over unchanged, and it is worth writing out, because Hudson's Theorem: the Pure States of Non-Negative Wigner Function needs the same statement.
Let \(h:\R^{n}\longrightarrow\C\) be continuous with \(\int_{\R^{n}}\abs{h}<\infty\), and suppose \(\int_{\R^{n}}h(\xi)\,\ee^{\ii\eta\cdot\xi}\,\dd^{n}\xi=0\) for every \(\eta\in\R^{n}\). Then \(h\equiv0\). Rests on Theorem 17.32, Equation (17.42) and Definition 17.29.
Derives Lemma A.584. Write \(\hat h(\eta)=(2\pi)^{-n/2}\int h(\xi)\ee^{-\ii\eta\cdot\xi} \dd^{n}\xi\), which vanishes identically by hypothesis, with \(\eta\) replaced by \(-\eta\). For \(\varepsilon>0\) put
Inserting the definition of \(\hat h\) and exchanging the order of integration — legitimate because \(\abs{h(\xi')}\ee^{-\varepsilon\abs{\eta}^{2}/2}\) is integrable over \(\R^{n}\times\R^{n}\), so Equation (7.112) applies on an exhausting sequence of boxes — gives \(I_{\varepsilon}(\xi)=\int h(\xi')g_{\varepsilon}(\xi-\xi') \dd^{n}\xi'\) with
the inner \(\eta\)-integral factorising into \(n\) one-dimensional Gaussian integrals, each evaluated by Equation (17.42). The family \(g_{\varepsilon}\) is an approximate identity exactly as in the proof of Theorem 17.32: non-negative, of unit mass, and concentrating on \(\abs{u}<\delta\) as \(\varepsilon\to0\). Since \(h\) is continuous at the fixed point \(\xi\), bounded near it and integrable at infinity, \(I_{\varepsilon}(\xi)\to h(\xi)\). But \(I_{\varepsilon}\equiv0\), so \(h(\xi)=0\).
∎The Gaussian average
With \(s>0\) fixed and \(\zeta\) as in Equation (A.875), put
a bounded operator by Lemma A.583. The two expressions agree by Equation (A.875), and the second exhibits the weight as a Gaussian of width \(\hbar/s\) in momentum and \(s\) in position. Rests on Lemma A.583 and Definition A.580.
Everything rests on one identity, which is a Gaussian integral and nothing else.
For every \(n\geq1\) and every \(v\in\C^{n}\),
where \(\zeta\cdot v=\sum_{k}\zeta_{k}v_{k}\) and \(v\cdot v=\sum_{k}v_{k}^{2}\) are bilinear, not Hermitian. Rests on Equations (7.112), (17.35) and (17.42).
Derives Lemma A.586. Split \(v=p+\ii q\) with \(p,q\in\R^{n}\) and complete the square in the real part, \(-\abs{\zeta}^{2}/2+\zeta\cdot p =\abs{p}^{2}/2-\abs{\zeta-p}^{2}/2\). Substituting \(u=\zeta-p\),
The remaining integral factorises into \(n\) one-dimensional integrals by Equation (7.112), and each is Equation (17.42) with \(a=1\), giving \((2\pi)^{n/2}\ee^{-\abs{q}^{2}/2}\). Collecting the three exponents, \(\abs{p}^{2}/2+\ii p\cdot q-\abs{q}^{2}/2=v\cdot v/2\), which is Equation (A.885).
∎\(\Pi\) is self-adjoint, \(\Pi^{2}=\Pi\), and for every \(\eta\in\R^{2f}\)
Moreover \(\Pi\neq0\) whenever \(\mathcal{H}\neq\set{0}\). Rests on Definition A.585, Lemma A.586 and Lemma A.584.
Derives Theorem A.587. Self-adjointness. The weight \(G(\zeta)=(2\pi)^{-f}\ee^{-\abs{\zeta}^{2}/4}\) is real and even, so \(G^{\sharp}=G\) and the second identity of Equation (A.879) gives \(\Pi^{\dagger}=\Pi\).
Idempotence. By Equation (A.880) it suffices to show that
Expand the product of the two Gaussians,
and write the phase as \(\ii\sigma_{0}(\zeta,u)/2=\ii\,\zeta\cdot Ju/2\). The integral in Equation (A.888) is therefore \((2\pi)^{-2f}\ee^{-\abs{u}^{2}/4}\) times the integral Equation (A.885) in \(n=2f\) variables with
because \(u\cdot Ju=\sigma_{0}(u,u)=0\) and \(Ju\cdot Ju=\abs{u}^{2}\) by orthogonality of \(J\). Hence the integral equals \((2\pi)^{-2f}\ee^{-\abs{u}^{2}/4}(2\pi)^{f} =(2\pi)^{-f}\ee^{-\abs{u}^{2}/4}=G(u)\), which is Equation (A.888). The exponent \(-\abs{\zeta}^{2}/4\) and the coefficient \((2\pi)^{-f}\) are fixed by this requirement and by nothing else; that is where they come from.
The absorption identity. Expanding as in Equation (A.881), and with the same justification,
and two applications of Equation (A.876) give
Substitute \(u=\zeta+\zeta'\) at fixed \(\zeta\). Using \(\sigma_{0}(\zeta,u-\zeta)=\sigma_{0}(\zeta,u)\) and \(\sigma_{0}(\eta,u-\zeta)=\sigma_{0}(\eta,u)+\sigma_{0}(\zeta,\eta)\), the bracket becomes \(2\sigma_{0}(\zeta,\eta)+\sigma_{0}(\zeta,u)+\sigma_{0}(\eta,u)\), so
The inner integral is Equation (A.885) again, with the same Gaussian Equation (A.889) but with \(v=\bigl[u+\ii J(u+2\eta)\bigr]/2=p+\ii q\), where \(p=u/2\) and \(q=J(u/2+\eta)\). Now
using \(\abs{J\cdot}=\abs{\cdot}\) and \(p\cdot q=(u/2)\cdot J(u/2+\eta)=\sigma_{0}(u/2,\eta)\). Hence the inner integral equals \(G(u)\,\ee^{-u\cdot\eta/2-\abs{\eta}^{2}/2}\, \ee^{\ii\sigma_{0}(u,\eta)/2}\), and its phase cancels the prefactor \(\ee^{\ii\sigma_{0}(\eta,u)/2}\) of Equation (A.893) exactly, because \(\sigma_{0}(\eta,u)+\sigma_{0}(u,\eta)=0\). What remains is a pure Gaussian in \(u\),
so that, substituting \(w=u+\eta\), \(\Pi W(\eta)\Pi=\ee^{-\abs{\eta}^{2}/4} (2\pi)^{-f}\int\ee^{-\abs{w}^{2}/4}W(w)\dd^{2f}w =\ee^{-\abs{\eta}^{2}/4}\Pi\). That is Equation (A.887), and at \(\eta=0\) it reproduces \(\Pi^{2}=\Pi\), a check on every sign above.
Non-vanishing. Suppose \(\Pi=0\). Conjugating a single Weyl operator by another, Equation (A.876) twice gives \(W(\eta)W(\zeta)W(\eta)^{\dagger} =\ee^{\ii\sigma_{0}(\eta,\zeta)}W(\zeta)\), hence
which vanishes for every \(\eta\) if \(\Pi\) does. Fix \(x,y\in\mathcal{H}\) and set \(h(\zeta)=G(\zeta)\braket{y}{W(\zeta)x}\), continuous by Lemma A.582 and absolutely integrable because \(\abs{h}\leq G\norm{x}\norm{y}\). Taking the matrix element of Equation (A.896) gives \(\int h(\zeta)\ee^{\ii\eta\cdot J\zeta}\dd^{2f}\zeta=0\) for every \(\eta\); since \(\eta\cdot J\zeta=-(J\eta)\cdot\zeta\) and \(J\) is a bijection of \(\R^{2f}\), that is the hypothesis of Lemma A.584, so \(h\equiv0\). At \(\zeta=0\), \(h(0)=(2\pi)^{-f}\braket{y}{x}\), so \(\braket{y}{x}=0\) for all \(x\) and \(y\) — impossible unless \(\mathcal{H}=\set{0}\).
∎Nothing above fixes \(s\), and no statement below depends on it: the dimensionless variable Equation (A.875) absorbs it completely, so \(\Pi\) and every identity satisfied by it are the same for every \(s\). What does change with \(s\) is the vector \(\Pi\) projects onto, which in the Schrödinger picture is the Gaussian of width \(s\), Equation (A.902). The theorem is a statement about a representation, so it cannot depend on \(s\); the vacuum is a statement about a vector, so it must.
The vacuum vector
Let \(\Omega\in\operatorname{ran}\Pi\) with \(\norm{\Omega}=1\), and let \(M_{\Omega}\) be the closed linear span of \(\set{W(\zeta)\Omega:\zeta\in\R^{2f}}\). Then
-
\(M_{\Omega}\) reduces the Weyl system, and the restriction of the system to \(M_{\Omega}\) is again a Weyl system whose own Gaussian average is the restriction of \(\Pi\);
-
that restricted average has range exactly \(\C\,\Omega\), and the restricted system is irreducible;
-
if the original system is irreducible then \(M_{\Omega}=\mathcal{H}\) and \(\operatorname{ran}\Pi=\C\,\Omega\).
Rests on Theorem A.587, Definition 12.90 and Proposition 12.89.
Derives Proposition A.589. (1) \(M_{\Omega}\) is invariant under every \(W(\eta)\), since \(W(\eta)W(\zeta)\Omega\) is a multiple of \(W(\eta+\zeta)\Omega\) by Equation (A.876), and the family is self-adjoint by Lemma A.581; so \(M_{\Omega}^{\perp}\) is invariant too and \(M_{\Omega}\) reduces the system in the sense of Proposition 12.89. The restricted groups are again strongly continuous unitary groups on \(M_{\Omega}\) satisfying Equation (A.868), and the restricted average is given by the same formula Equation (A.884), so it is \(\Pi\) restricted to \(M_{\Omega}\), which by the last clause of Lemma A.583 maps \(M_{\Omega}\) into itself.
(2) \(\Omega=\Pi\Omega\in M_{\Omega}\) lies in the range of the restricted average. Let \(u\in M_{\Omega}\) satisfy \(\Pi u=u\) and \(\braket{\Omega}{u}=0\). Then, for every \(\eta\),
by Equation (A.887) and \(\Pi=\Pi^{\dagger}\). So \(u\) is orthogonal to every generator of \(M_{\Omega}\), hence to \(M_{\Omega}\), hence to itself: \(u=0\). The range is therefore the line \(\C\Omega\). For irreducibility, let \(N\subseteq M_{\Omega}\) be a closed subspace reducing the restricted system, and let \(N'\) be its orthogonal complement inside \(M_{\Omega}\). By the last clause of Lemma A.583 both \(N\) and \(N'\) are invariant under \(\Pi\), and their images are orthogonal subspaces of the line \(\C\Omega\), so at most one of them is non-zero; and at least one is, because \(\Pi\Omega=\Omega\neq0\). Say \(\Pi N'=\set{0}\) and write \(\Omega=m+m'\) with \(m\in N\), \(m'\in N'\); then \(\Omega=\Pi\Omega=\Pi m\in N\), so \(N\) contains every \(W(\zeta)\Omega\) and equals \(M_{\Omega}\). In the other case the same argument gives \(N'=M_{\Omega}\), i.e. \(N=\set{0}\).
(3) \(M_{\Omega}\neq\set{0}\) contains \(\Omega\), so irreducibility of the original system forces \(M_{\Omega}=\mathcal{H}\), and then (2) is the statement about \(\Pi\) itself.
∎Let the system act irreducibly and let \(\Omega\) be a unit vector spanning \(\operatorname{ran}\Pi\). Then for all \(\zeta,\zeta'\)
The right-hand side involves the Weyl algebra alone: it is the same number in every irreducible Weyl system of \(f\) degrees of freedom. Rests on Proposition A.589, Theorem A.587 and Lemma A.581.
Derives Lemma A.590. First, \(\Pi\Omega=\Omega\) and Equation (A.887) give \(\braket{\Omega}{W(\eta)\Omega} =\braket{\Omega}{\Pi W(\eta)\Pi\Omega} =\ee^{-\abs{\eta}^{2}/4}\). Then, by Equation (A.876) and \(W(\zeta)^{\dagger}=W(-\zeta)\),
which is Equation (A.898).
∎The intertwining unitary
Let \((U,V)\) on \(\mathcal{H}\) and \((U',V')\) on \(\mathcal{H}'\) be irreducible Weyl systems of the same finite number \(f\) of degrees of freedom. Then there is a unitary \(T:\mathcal{H}\longrightarrow\mathcal{H}'\) with \(TU(\vect{\alpha})=U'(\vect{\alpha})T\) and \(TV(\vect{\beta})=V'(\vect{\beta})T\) for all \(\vect{\alpha},\vect{\beta}\), and \(T\) is unique up to a factor of modulus one. Rests on Lemma A.590, Proposition A.589 and Theorem 12.91.
Derives Proposition A.591. Choose unit vectors \(\Omega\in\operatorname{ran}\Pi\) and \(\Omega'\in\operatorname{ran}\Pi'\), which exist and span those ranges by Theorem A.587 and Proposition A.589. On the dense subspace \(D\subset\mathcal{H}\) of finite linear combinations \(x=\sum_{k=1}^{N}c_{k}W(\zeta_{k})\Omega\) — dense because \(M_{\Omega}=\mathcal{H}\) — define
By Lemma A.590 applied in each space,
the two Gram matrices being equal entry by entry. In particular the right-hand side of Equation (A.900) vanishes whenever \(x\) does, so \(T\) is well defined on \(D\), and it is isometric. Its range contains every \(W'(\zeta)\Omega'\) and is therefore dense in \(\mathcal{H}'\), again by Proposition A.589. An isometry with dense domain and dense range extends by continuity to a unitary of \(\mathcal{H}\) onto \(\mathcal{H}'\). Intertwining holds on \(D\) by construction — \(TW(\eta)W(\zeta_{k})\Omega\) and \(W'(\eta)TW(\zeta_{k})\Omega\) are both \(\ee^{\ii\sigma_{0}(\eta,\zeta_{k})/2}W'(\eta+\zeta_{k})\Omega'\) — and extends by continuity; by the last clause of Lemma A.581 intertwining the \(W\)'s is the same as intertwining the \(U\)'s and the \(V\)'s.
For uniqueness, let \(T_{1}\) and \(T_{2}\) both intertwine. Then \(S=T_{2}T_{1}^{-1}\) is a unitary of \(\mathcal{H}'\) commuting with every \(W'(\zeta)\). The family \(\set{W'(\zeta)}\) is self-adjoint and acts irreducibly, so Theorem 12.91 gives \(S=c\,\identity\) with \(c\in\C\), and \(\abs{c}=1\) because \(S\) is unitary.
∎The Schrödinger system
It remains to exhibit one irreducible Weyl system and to identify its vacuum. Here the Gaussian average can be computed in closed form, and the computation is the shortest route to both facts at once.
On \(\mathcal{H}=L^{2}(\R^{f})\) take the Schrödinger system Equation (12.72), read componentwise for \(f\) degrees of freedom, and build \(\Pi\) from it by Equation (A.884). Then, with
the average is the rank-one orthogonal projector onto that vector,
Consequently \(\set{W(\zeta)\Omega_{s}}\) is total in \(L^{2}(\R^{f})\) and the Schrödinger system acts irreducibly. Rests on Example 12.110, Theorem A.587 and Lemma A.586.
Derives Proposition A.592. The kernel. By Equation (12.72) and Equation (A.869),
In the dimensionless variables Equation (A.875), \(\vect{\alpha}\cdot\vect{\beta}/2\hbar=\vect{a}\cdot\vect{b}/2\), \(\vect{\alpha}\cdot\vect{x}/\hbar=\vect{a}\cdot\vect{x}/s\) and \(\vect{\beta}=s\vect{b}\), so
Do the \(\vect{a}\) integral first, at fixed \(\vect{b}\). It is Equation (A.885) in \(n=f\) variables, after the rescaling \(\vect{a}=\sqrt{2}\,\vect{u}\), and gives
Substituting \(\vect{x}'=\vect{x}-s\vect{b}\), so that \(\vect{b}=(\vect{x}-\vect{x}')/s\) and \(\dd^{f}b=s^{-f}\dd^{f}x'\), the two exponents combine:
because \(\vect{x}/s-\vect{b}/2=(\vect{x}+\vect{x}')/2s\) and the parallelogram identity turns the sum of the two squares into twice the sum of the squares of \(\vect{x}\) and \(\vect{x}'\). Hence
and the constant is \((4\pi)^{f/2}(2\pi)^{-f}s^{-f}=(\pi s^{2})^{-f/2}\), which is exactly what turns the two bare Gaussians into the normalised Equation (A.902) twice over. That is Equation (A.903), and it is manifestly an orthogonal projector of rank one, in agreement with Theorem A.587.
Totality. Let \(K\) be the closed linear span of \(\set{W(\zeta)\Omega_{s}}\); it reduces the system by Proposition A.589(1), so \(K^{\perp}\) is invariant and carries a Weyl system of its own, whose Gaussian average is the restriction of \(\Pi\) to \(K^{\perp}\). But that restriction is zero: for \(x\in K^{\perp}\), Equation (A.903) gives \(\Pi x=\braket{\Omega_{s}}{x}\Omega_{s}=0\), because \(\Omega_{s}=W(0)\Omega_{s}\in K\). By the last clause of Theorem A.587 a Weyl system on a non-zero space has a non-zero average, so \(K^{\perp}=\set{0}\) and \(K=L^{2}(\R^{f})\).
Irreducibility. \(\Omega_{s}\) spans \(\operatorname{ran}\Pi\) and \(M_{\Omega_{s}}=K=L^{2}(\R^{f})\), so Proposition A.589(2) applied to \(M_{\Omega_{s}}\) says that the system is irreducible.
∎Proof of Theorem A.579. Derives Theorem A.579. Part (1) is Proposition A.591 applied to the given system and to the Schrödinger system, which is an irreducible Weyl system by Proposition A.592.
Part (2). Let \(\mathcal{H}\neq\set{0}\) carry a Weyl system, not assumed irreducible. By Theorem A.587 the operator \(\Pi\) is a non-zero orthogonal projector; choose a unit vector \(\Omega_{1}\) in its range and put \(M_{1}=M_{\Omega_{1}}\). By Proposition A.589, \(M_{1}\) reduces the system and the restriction to it is an irreducible Weyl system, hence carries a copy of the Schrödinger system by part (1). Now \(M_{1}^{\perp}\) is invariant and carries a Weyl system of its own; if it is not \(\set{0}\) its average is again non-zero, and the construction produces \(M_{2}\subseteq M_{1}^{\perp}\), and so on. The subspaces so produced are mutually orthogonal and each is infinite-dimensional — it is a copy of \(L^{2}(\R^{f})\) — so separability of \(\mathcal{H}\) (Definition 12.2) allows at most countably many of them, and the process exhausts \(\mathcal{H}\): the orthogonal complement of their closed sum is invariant and carries a vanishing average, hence is \(\set{0}\). That is the direct-sum statement, the sum being the one of Definition 12.84.
∎What is quoted here
Nothing, beyond the standing declaration of Remark 12.1. Every step above is a Gaussian integral, an application of Equation (17.42) through Lemma A.586 and Lemma A.584, or an application of Definition 12.90 and Theorem 12.91 — all proved in Hilbert Spaces or in Fourier Analysis and Integral Transforms. Two points deserve naming rather than leaving implicit.
First, the passage between the exponentiated form and the unbounded generators \(\vect{Q}\) and \(\vect{P}\) is Theorem 12.66, proved in Stone's Theorem on One-Parameter Unitary Groups; it is what makes Theorem A.579 a statement about the canonical commutation relations at all, and Remark 25.36 explains why the unexponentiated relation would not do. The vector-valued integral of Lemma A.583 is the one built there, in Lemma A.254, extended from a compact interval to \(\R^{2f}\) by absolute convergence.
Second, the irreducibility of the Schrödinger system is here proved, in Proposition A.592, by computing the Gaussian average in closed form. Example 12.110 takes it instead from [Reed:1972], through the maximal-abelian property of \(L^{\infty}\) acting on \(L^{2}\); that route is shorter but rests on a fact about operator algebras which this treatise does not build, so the argument given here is the one on which Theorem A.579 stands.
The exponentiated form of the commutation relations is Weyl's, from Quantenmechanik und Gruppentheorie (Zeitschrift für Physik 46, 1927, 1–46), and the uniqueness theorem is von Neumann's, from Die Eindeutigkeit der Schrödingerschen Operatoren (Mathematische Annalen 104, 1931, 570–578); Stone had announced the result, and the theorem carries both names. Neither of those two papers has an entry in this treatise's bibliography — and the existing key for Weyl in 1929 is a different work of his, on the electron and gravitation, which must not be used for it — so that part of the attribution is made in words, on the footing of Darboux's memoir in Remark A.81. Nothing above rests on it, the proof being carried out in full. Stone's own paper on one-parameter unitary groups [Stone:1932] is in the bibliography and is cited where it is used, in Stone's Theorem on One-Parameter Unitary Groups; the modern treatment against which every step here can be checked is [Reed:1972], theorem VIII.14.
The Stone–von Neumann Theorem discharges the derivation owed at Theorem 25.34 of The Poisson Algebra and the Canonical Bridge to Quantum Mechanics, and with it the debt recorded in Remark 12.115 against Theorem 12.114 of Hilbert Spaces — the two are the same statement, and Remark 25.35 says that the proof belongs to Part II. Where the result is cashed is worth repeating. In The Poisson Algebra and the Canonical Bridge to Quantum Mechanics it is hypothesis (3) of Theorem 25.38: the theorem is what makes “the” Schrödinger representation a legitimate phrase, so that the Groenewold–van Hove obstruction obstructs a quantization rule and not merely one choice of Hilbert space. Its failure for infinitely many degrees of freedom, which Remark 25.37 and Remark 12.116 both record, is visible in the proof at exactly one place: the Gaussian weight of Equation (A.884) is a product over the \(f\) degrees of freedom, and for \(f=\infty\) that product is not a measure on any space on which the argument could be run. Nothing here can be repaired to cover that case, and the physics of inequivalent vacua is the reason it should not be.
Hudson's Theorem: the Pure States of Non-Negative Wigner Function
This appendix proves Theorem 25.55 of The Poisson Algebra and the Canonical Bridge to Quantum Mechanics: a pure state has a Wigner function that is non-negative everywhere if and only if its wavefunction is Gaussian. Together with Corollary 25.54 it is what makes negativity of the Wigner function an exact witness of departure from the Gaussian family, and the chapter uses it in that role.
Remark 25.56 states that the converse half of the theorem “reduces it to one imported theorem — the Hadamard factorization of an entire function of order at most two” and records the absence of any order-and-genus theory in Complex Analysis as a debt of Part II. That is the standard route, and it is more than is needed. The only consequence of Hadamard's theorem this proof uses is that an entire function of one variable whose real part is bounded above by \(A+B\abs{\lambda}^{2}\) is a polynomial of degree at most two, and that statement follows in a dozen lines from the Taylor expansion Theorem 8.20 and the coefficient formula Equation (8.13), both of which Complex Analysis proves. It is Lemma A.608 below. Consequently this section quotes nothing, Part II is not in arrears on this account, and Remark 25.56 overstates the debt — which is worth saying plainly, because an unnecessary debt on the books is as misleading as a hidden one.
The other two statements of Remark 25.56 stand unchanged and are repeated in Remark A.614: the theorem is about pure states only, and Hudson's 1974 paper has no entry in this bibliography, so the attribution is made without a source the reader can check — which is exactly why the proof is written out.
Statement and conventions
Throughout, the system is one particle in three-dimensional space, as in Definition 25.52, and \(\psi\in L^{2}(\R^{3})\) is normalised. Its Wigner function is Equation (25.53) with \(\rho=\psi\psi^{\ast}\), that is
carrying \(/\mathrm{J}^{3}/\mathrm{s}^{3}\). For a complex symmetric \(3\times3\) matrix \(A\) write \(A_{\mathrm{r}}=\operatorname{Re}A\) and \(A_{\mathrm{i}}=\operatorname{Im}A\), both real symmetric, and \(A_{\mathrm{r}}>0\) for positive definiteness. Bilinear, never Hermitian, products are meant throughout: \(\vect{u}\cdot A\vect{v} =\sum_{jk}u_{j}A_{jk}v_{k}\) and \(\vect{u}\cdot\vect{v}=\sum_{j}u_{j}v_{j}\), even for complex arguments.
In the Gaussian form \(\psi(\vect{x})=\exp(-\vect{x}\cdot A\vect{x} +\vect{b}\cdot\vect{x}+c)\) the exponent must be dimensionless, so \(A\) carries \(/\mathrm{m}^{2}\) and \(\vect{b}\) carries \(/\mathrm{m}\); \(\ee^{c}\) is then whatever makes \(\abs{\psi}^{2}\) carry \(/\mathrm{m}^{3}\), so \(c\) is fixed only once a unit of length is chosen and it is cleaner to read the Gaussian family as \(\psi=\mathcal{N}\exp(-\vect{x}\cdot A\vect{x}+\vect{b}\cdot\vect{x})\) with \(\mathcal{N}\) the normalisation. Nothing below depends on which of the two readings is used, and \(\hbar\) appears explicitly at every step where it appears at all: in the kernel of Equation (A.909), and in the auxiliary length \(s\) of Coherent states, and where the Wigner function cannot vanish, which enters the combination \(s^{2}\vect{p}/\hbar\) — a length, as it must be.
Let \(\psi\in L^{2}(\R^{3})\) be normalised. Then \(W\geq0\) everywhere if and only if there are a complex symmetric matrix \(A\) with \(A_{\mathrm{r}}>0\), a vector \(\vect{b}\in\C^{3}\) and a constant \(c\in\C\) with
for almost every \(\vect{x}\). When that holds, \(W\) is a strictly positive Gaussian on phase space. Rests on Equation (25.53), Equation (25.56) and Theorem 8.20.
Gaussian integrals
Let \(a\in\C\) with \(\operatorname{Re}a>0\) and \(w\in\C\). Then
the square root being the principal branch, which is defined and holomorphic on the half-plane \(\operatorname{Re}a>0\). Rests on Equation (17.42), Theorem 7.77 and Theorem 8.6.
Derives Lemma A.599. Both integrals converge absolutely, since \(\abs{\ee^{-az^{2}+wz}} =\ee^{-(\operatorname{Re}a)z^{2}+(\operatorname{Re}w)z}\).
The case \(w=0\). Put \(I(a)=\int_{\R}\ee^{-az^{2}}\dd z\). Differentiating under the integral sign, which Theorem 7.77 licenses on any compact subset of the half-plane because the differentiated integrand is dominated there by \(z^{2}\ee^{-\varepsilon z^{2}}\) for some \(\varepsilon>0\), and checking the Cauchy–Riemann equations Equation (8.5) on the integrand, \(I\) is holomorphic with \(I'(a)=-\int z^{2}\ee^{-az^{2}}\dd z\). Integrating that by parts, \(\int z\cdot\bigl(z\ee^{-az^{2}}\bigr)\dd z =\int z\cdot\bigl(-\tfrac{1}{2a}\bigr) \tfrac{\dd}{\dd z}\ee^{-az^{2}}\dd z =\tfrac{1}{2a}\int\ee^{-az^{2}}\dd z\), the boundary terms vanishing. Hence
on the half-plane, which is connected; and \(\sqrt{1}\,I(1)=\sqrt{\pi}\) by the Gauss integral, itself the case \(\omega=0\) of Equation (17.42). So \(I(a)=\sqrt{\pi/a}\).
General \(w\). Put \(J(w)=\int_{\R}\ee^{-az^{2}+wz}\dd z\) at fixed \(a\). The same argument makes \(J\) entire with \(J'(w)=\int z\,\ee^{-az^{2}+wz}\dd z\), and the same integration by parts, now carrying the factor \(\ee^{wz}\), gives \(J'(w)=\tfrac{w}{2a}J(w)\). With \(J(0)=\sqrt{\pi/a}\) this integrates to Equation (A.911).
∎Let \(M\) and \(S\) be real symmetric \(n\times n\) matrices with \(M>0\). Then \(M+\ii S\) is invertible and
Derives Lemma A.600. Look for the inverse in the form \(X+\ii Y\) with \(X,Y\) real. The two equations \(MX-SY=\identity\) and \(MY+SX=0\) give \(Y=-M^{-1}SX\) and then \(\left(M+SM^{-1}S\right)X=\identity\). The matrix \(M+SM^{-1}S\) is real symmetric and positive definite — it is a sum of a positive definite matrix and the positive semi-definite \(SM^{-1}S\), since \(v\cdot SM^{-1}Sv=(Sv)\cdot M^{-1}(Sv)\geq0\) and \(M^{-1}>0\) — hence invertible, and \(X=(M+SM^{-1}S)^{-1}\) solves the pair. That exhibits an inverse, which is unique, and its real part is \(X\).
∎Let \(N\) be a complex symmetric \(n\times n\) matrix with \(N_{\mathrm{r}}>0\), and let \(\vect{w}\in\C^{n}\). Then \(N\) is invertible, \(\operatorname{Re}\bigl(N^{-1}\bigr)>0\), and
with \(C_{N}\neq0\) depending on \(N\) alone; if \(N\) is real then \(C_{N}=\pi^{n/2}\left(\det N\right)^{-1/2}>0\). Rests on Lemma A.599, Lemma A.600 and Theorem 5.80.
Derives Lemma A.601. Invertibility and \(\operatorname{Re}(N^{-1})>0\) are Lemma A.600 with \(M=N_{\mathrm{r}}\) and \(S=N_{\mathrm{i}}\). For the integral, let \(R\) be the real symmetric positive-definite square root of \(N_{\mathrm{r}}\), which exists by Theorem 5.80 — diagonalise \(N_{\mathrm{r}}\) orthogonally and take the positive square roots of its eigenvalues — and substitute \(\vect{x}=R^{-1}\vect{y}\). Then \(\vect{x}\cdot N\vect{x}=\vect{y}\cdot(\identity+\ii T)\vect{y}\) with \(T=R^{-1}N_{\mathrm{i}}R^{-1}\) real symmetric. Diagonalise \(T\) orthogonally, \(T=O\transpose DO\) with \(D=\diag(d_{1},\ldots,d_{n})\), and substitute \(\vect{y}=O\transpose\vect{z}\); the quadratic form becomes \(\sum_{j}(1+\ii d_{j})z_{j}^{2}\) and the linear form becomes \(\vect{w}'\cdot\vect{z}\) with \(\vect{w}'=OR^{-1}\vect{w}\). The integral factorises into \(n\) copies of Equation (A.911) with \(a_{j}=1+\ii d_{j}\), each of which has \(\operatorname{Re}a_{j}=1>0\):
the Jacobian being \(\abs{\det(R^{-1}O\transpose)}=1/\det R\). The exponent reassembles as \(\vect{w}'\cdot(\identity+\ii D)^{-1}\vect{w}'/4 =\vect{w}\cdot R^{-1}O\transpose(\identity+\ii D)^{-1}OR^{-1}\vect{w}/4 =\vect{w}\cdot N^{-1}\vect{w}/4\), because \(N=R(\identity+\ii T)R\). The prefactor is a non-zero constant determined by \(N\), and for real \(N\) every \(d_{j}\) is zero and it is \(\pi^{n/2}(\det N)^{-1/2}\).
∎The easy direction
Let \(\psi\) be of the form Equation (A.910) with \(A_{\mathrm{r}}>0\). Then, with \(\vect{w}=-2A_{\mathrm{i}}\vect{x}+\operatorname{Im}\vect{b} -\vect{p}/\hbar\),
which is real, strictly positive at every point of phase space, and Gaussian in \((\vect{x},\vect{p})\) jointly. Rests on Equation (25.53), Lemma A.601 and Equation (A.910).
Derives Proposition A.602. Put \(\vect{u}=\vect{x}+\vect{y}/2\) and \(\vect{v}=\vect{x}-\vect{y}/2\), so that the integrand of Equation (A.909) is \(\exp\bigl(-\vect{u}\cdot A\vect{u}+\vect{b}\cdot\vect{u}+c -\vect{v}\cdot\overline{A}\vect{v} +\overline{\vect{b}}\cdot\vect{v}+\overline{c}\bigr)\) times \(\ee^{-\ii\vect{p}\cdot\vect{y}/\hbar}\). Expanding by symmetry of \(A\),
while \(\vect{b}\cdot\vect{u}+\overline{\vect{b}}\cdot\vect{v} =2\operatorname{Re}\vect{b}\cdot\vect{x} +\ii\operatorname{Im}\vect{b}\cdot\vect{y}\) and \(c+\overline{c}=2\operatorname{Re}c\). Collecting the terms carrying \(\vect{y}\) and adding the Fourier kernel, the whole \(\vect{y}\) dependence is \(-\tfrac{1}{2}\vect{y}\cdot A_{\mathrm{r}}\vect{y} +\ii\,\vect{y}\cdot\vect{w}\) with \(\vect{w}\) as stated, a real vector. The remaining integral is Equation (A.914) with the real positive-definite \(N=A_{\mathrm{r}}/2\) and \(\vect{w}\to\ii\vect{w}\), giving \((2\pi)^{3/2}(\det A_{\mathrm{r}})^{-1/2} \exp(-\tfrac{1}{2}\vect{w}\cdot A_{\mathrm{r}}^{-1}\vect{w})\). That is Equation (A.916). Every factor is real and positive: the exponential of a real number, and a positive constant. Since \(\vect{w}\) is an affine function of \((\vect{x},\vect{p})\), the exponent is a real quadratic in \((\vect{x},\vect{p})\) and \(W\) is a Gaussian on phase space.
∎Coherent states, and where the Wigner function cannot vanish
Fix once and for all a length \(s>0\) and put, for \(\vect{x}_{0},\vect{p}_{0}\in\R^{3}\),
a normalised element of \(L^{2}(\R^{3})\) of the form Equation (A.910), with \(A=\identity/2s^{2}\), \(\vect{b}=\vect{x}_{0}/s^{2} +\ii\vect{p}_{0}/\hbar\) and the appropriate \(c\). It is exactly the vector \(W(\zeta)\Omega_{s}\) of Proposition A.592, up to a phase: by Equation (A.904), \(\bigl(W(\vect{\alpha},\vect{\beta})\Omega_{s}\bigr)(\vect{x}) =\ee^{-\ii\vect{\alpha}\cdot\vect{\beta}/2\hbar} \varphi_{\vect{\beta},\vect{\alpha}}(\vect{x})\). That identification is used once, at the end of Recovering the wavefunction.
For any normalised \(\psi\),
and if \(W_{\psi}\geq0\) everywhere then
Rests on Equation (25.56), Theorem 25.53 and Equation (A.918).
Derives Lemma A.603. For pure states \(\hat{\rho}_{1}=\ketbra{\psi_{1}}{\psi_{1}}\) and \(\hat{\rho}_{2}=\ketbra{\psi_{2}}{\psi_{2}}\) one has \(\tr(\hat{\rho}_{1}\hat{\rho}_{2}) =\braket{\psi_{1}}{\psi_{2}}\braket{\psi_{2}}{\psi_{1}} =\abs{\braket{\psi_{1}}{\psi_{2}}}^{2}\), so Equation (25.56) is Equation (A.919).
For the second statement, note first that Equation (A.918) is \(\varphi_{\vect{x}_{0},\vect{p}_{0}}(\vect{x}) =\ee^{\ii\vect{p}_{0}\cdot\vect{x}/\hbar}\, \varphi_{\vect{0},\vect{0}}(\vect{x}-\vect{x}_{0})\), and that substituting this into Equation (A.909) makes the two exponential factors combine as \(\ee^{\ii\vect{p}_{0}\cdot\vect{y}/\hbar}\) and shifts the argument, so that
By Proposition A.602 that function is a strictly positive Gaussian, and by Theorem 25.53 it integrates to \(1\) over phase space; hence \(\iint W_{\varphi_{\vect{x}_{0},\vect{p}_{0}}}(\vect{x},\vect{p}) \dd^{3}x_{0}\dd^{3}p_{0}=1\) for every \((\vect{x},\vect{p})\). Integrating Equation (A.919) over \((\vect{x}_{0},\vect{p}_{0})\) and exchanging the order of integration — legitimate because every integrand in sight is non-negative, which is where the hypothesis \(W_{\psi}\geq0\) is used — gives \((2\pi\hbar)^{3}\int W_{\psi}\dd^{3}x\,\dd^{3}p\), which is \((2\pi\hbar)^{3}\) by the normalisation clause of Theorem 25.53.
∎If \(W_{\psi}\geq0\) everywhere then \(\braket{\varphi_{\vect{x}_{0},\vect{p}_{0}}}{\psi}\neq0\) for every \(\vect{x}_{0}\) and every \(\vect{p}_{0}\). Rests on Lemma A.603, Proposition A.602 and Equation (25.54).
Derives Proposition A.604. Suppose the overlap vanishes for some \((\vect{x}_{0},\vect{p}_{0})\). By Equation (A.919) the integral of \(W_{\varphi}W_{\psi}\) over phase space is then zero. Both factors are non-negative — the first strictly positive everywhere by Proposition A.602, the second by hypothesis — so the product vanishes identically, and therefore \(W_{\psi}\) does. Integrating \(W_{\psi}\) over \(\vect{p}\) and using Equation (25.54) gives \(\abs{\psi(\vect{x})}^{2}=0\) for every \(\vect{x}\), contradicting \(\norm{\psi}=1\).
∎An entire function without zeros
For \(\zeta\in\C^{3}\) put
the product in the exponent being the bilinear one. Rests on Equation (A.918) and Definition 25.52.
The integral Equation (A.922) converges absolutely for every \(\zeta\in\C^{3}\) and defines a function holomorphic in each variable separately and jointly continuous, hence entire on \(\C^{3}\). Writing \(\zeta=\vect{x}_{0}-\ii s^{2}\vect{p}_{0}/\hbar\) with \(\vect{x}_{0},\vect{p}_{0}\in\R^{3}\),
and consequently
Rests on Definition A.605, Theorem 8.6 and Proposition 12.4.
Derives Lemma A.606. Write \(\zeta=\vect{\xi}+\ii\vect{\eta}\). Then \((\vect{x}-\zeta)\cdot(\vect{x}-\zeta) =\abs{\vect{x}-\vect{\xi}}^{2}-\abs{\vect{\eta}}^{2} -2\ii(\vect{x}-\vect{\xi})\cdot\vect{\eta}\), so the modulus of the kernel is \(\ee^{\abs{\vect{\eta}}^{2}/2s^{2}} \ee^{-\abs{\vect{x}-\vect{\xi}}^{2}/2s^{2}}\), a Gaussian in \(\vect{x}\); its product with \(\abs{\psi}\) is integrable by Cauchy–Schwarz (Proposition 12.4), and the bound is locally uniform in \(\zeta\). Differentiating under the integral sign in the real and imaginary parts of each \(\zeta_{k}\) — licensed by Theorem 7.77 on bounded regions together with the locally uniform domination just displayed, which controls the tails — shows that \(F\) has continuous first partial derivatives and that they satisfy the Cauchy–Riemann equations Equation (8.5) in each variable, because the integrand does. Hence \(F\) is holomorphic in each variable, and being continuous it is entire on \(\C^{3}\).
For Equation (A.923), substitute \(\zeta=\vect{x}_{0}-\ii s^{2}\vect{p}_{0}/\hbar\), so that \(\vect{x}-\zeta=(\vect{x}-\vect{x}_{0}) +\ii s^{2}\vect{p}_{0}/\hbar\) and
The first two terms are the exponent of \(\overline{\varphi_{\vect{x}_{0},\vect{p}_{0}}}\) up to the constant \((\pi s^{2})^{-3/4}\) and the phase \(\ee^{\ii\vect{x}_{0}\cdot\vect{p}_{0}/\hbar}\), and the third is \(\abs{\operatorname{Im}\zeta}^{2}/2s^{2}\) because \(\operatorname{Im}\zeta=-s^{2}\vect{p}_{0}/\hbar\). That is Equation (A.923). Finally \(\abs{\braket{\varphi}{\psi}}\leq\norm{\varphi}\norm{\psi}=1\) by Cauchy–Schwarz, which with \(\abs{\kappa}\) as displayed gives Equation (A.924); and every \(\zeta\in\C^{3}\) arises from exactly one pair \((\vect{x}_{0},\vect{p}_{0})\), namely \(\vect{x}_{0}=\operatorname{Re}\zeta\) and \(\vect{p}_{0}=-\hbar\operatorname{Im}\zeta/s^{2}\).
∎Let \(F\) be entire on \(\C^{n}\) with \(F(\zeta)\neq0\) for every \(\zeta\). Then there is an entire \(g\) on \(\C^{n}\) with \(F=\ee^{g}\). Rests on Lemma A.606, Theorem 8.6 and Theorem 7.77.
Derives Lemma A.607. Choose \(c_{0}\in\C\) with \(\ee^{c_{0}}=F(0)\), possible because \(F(0)\neq0\), and set
The integrand is continuous in \(t\) and holomorphic in \(\zeta\), since \(F\) is entire and non-vanishing, so \(g\) is entire by the same combination of Theorem 7.77 and Equation (8.5) used in Lemma A.606. Fix \(\zeta\) and substitute \(u=vt\) in Equation (A.926) written at the point \(t\zeta\); the result is \(g(t\zeta)=c_{0}+\int_{0}^{t}\sum_{j}\zeta_{j} (\pp_{j}F)(u\zeta)/F(u\zeta)\,\dd u\), so
the last step by the chain rule. Hence \(\frac{\dd}{\dd t}\bigl[F(t\zeta)\ee^{-g(t\zeta)}\bigr]=0\), so \(F(\zeta)\ee^{-g(\zeta)}=F(0)\ee^{-c_{0}}=1\).
∎Let \(u\) be entire on \(\C\) and suppose there are real constants \(A\) and \(B\) with \(\operatorname{Re}u(\lambda)\leq A+B\abs{\lambda}^{2}\) for all \(\lambda\). Then \(u\) is a polynomial of degree at most two. Rests on Theorem 8.20, Equation (8.13) and Theorem 8.18.
Derives Lemma A.608. By Theorem 8.20 the Taylor series \(u(\lambda)=\sum_{k\geq0}a_{k}\lambda^{k}\) converges on every disc. Fix \(r>0\). The series converges uniformly on the circle \(\abs{\lambda}=r\), so it may be integrated term by term against \(\ee^{-\ii k\vartheta}\), and for \(k\geq0\)
the first being Equation (8.13) written on the circle and the second holding because every term of the series then carries a strictly positive power of \(\ee^{\ii\vartheta}\). Conjugating the second identity and adding it to the first gives, for \(k\geq1\),
Write \(\mu(r)=A+Br^{2}\), so that \(\operatorname{Re}u(r\ee^{\ii\vartheta})\leq\mu(r)\) by hypothesis, and subtract the constant \(\mu(r)\) inside Equation (A.929), which changes nothing because \(\int_{0}^{2\pi}\ee^{-\ii k\vartheta}\dd\vartheta=0\) for \(k\geq1\). Bounding the modulus of the integral by the integral of the modulus, and using that \(\mu(r)-\operatorname{Re}u\) is non-negative so that its modulus is itself,
the mean of \(\operatorname{Re}u\) over the circle being \(\operatorname{Re}a_{0}\) by the case \(k=0\) of Equation (A.928). Hence for every \(k\geq1\)
and for \(k\geq3\) the right-hand side tends to \(0\) as \(r\to\infty\). So \(a_{k}=0\) for \(k\geq3\).
∎If \(W_{\psi}\geq0\) everywhere then there are a complex symmetric \(3\times3\) matrix \(\Gamma\), a vector \(\vect{\lambda}\in\C^{3}\) and a constant \(\nu\in\C\) with
Rests on Proposition A.604, Lemma A.607 and Lemma A.608.
Derives Proposition A.609. By Lemma A.606 and Proposition A.604 the entire function \(F\) has no zeros: \(F(\zeta)\) is a non-zero multiple of an overlap that never vanishes. So \(F=\ee^{g}\) with \(g\) entire, by Lemma A.607, and Equation (A.924) gives
since \(\abs{\operatorname{Im}\zeta}\leq\abs{\zeta}\).
Fix \(\vect{\beta}\in\C^{3}\) and consider \(u_{\vect{\beta}}(\lambda)=g(\lambda\vect{\beta})\), entire in the single variable \(\lambda\) because \(g\) is entire and \(\lambda\mapsto\lambda\vect{\beta}\) is holomorphic. By Equation (A.933), \(\operatorname{Re}u_{\vect{\beta}}(\lambda) \leq\frac{3}{4}\log(\pi s^{2}) +\abs{\vect{\beta}}^{2}\abs{\lambda}^{2}/2s^{2}\), so Lemma A.608 makes \(u_{\vect{\beta}}\) a polynomial of degree at most two:
exactly, with no remainder. The two derivatives are computed by the chain rule: \(\dd g(\lambda\vect{\beta})/\dd\lambda|_{0} =\sum_{j}\beta_{j}\,\pp_{j}g(0)\) and \(\dd^{2}g(\lambda\vect{\beta})/\dd\lambda^{2}|_{0} =\sum_{j,k}\beta_{j}\beta_{k}\,\pp_{j}\pp_{k}g(0)\). Setting \(\lambda=1\) and renaming \(\vect{\beta}\) as \(\zeta\), which is legitimate because \(\vect{\beta}\) was arbitrary,
which is Equation (A.932) with \(\Gamma_{jk}=-\tfrac{1}{2}\pp_{j}\pp_{k}g(0)\), symmetric because mixed holomorphic partial derivatives commute, \(\lambda_{j}=\pp_{j}g(0)\) and \(\nu=g(0)\). Note that no theorem about Taylor expansions in several variables was used: the expansion was performed along one complex line at a time, and Lemma A.608 guaranteed that on every such line it terminates.
∎What normalisability forces
The quadratic form \(\Gamma\) is not arbitrary, and the constraint it obeys is exactly the one that will make the recovered wavefunction square integrable. Write \(\Gamma_{\mathrm{r}}=\operatorname{Re}\Gamma\) and \(\Gamma_{\mathrm{i}}=\operatorname{Im}\Gamma\).
Let \(P_{11}\), \(P_{12}\) and \(P_{22}\) be real \(n\times n\) matrices with \(P_{11}\) and \(P_{22}\) symmetric. The real symmetric form \(Q(\vect{\xi},\vect{\eta}) =\vect{\xi}\cdot P_{11}\vect{\xi} +2\vect{\xi}\cdot P_{12}\vect{\eta} +\vect{\eta}\cdot P_{22}\vect{\eta}\) on \(\R^{2n}\) is positive definite if and only if \(P_{22}>0\) and \(P_{11}-P_{12}P_{22}^{-1}P_{12}\transpose>0\). Rests on Theorems 5.38 and 5.80.
Derives Lemma A.610. If \(P_{22}>0\), complete the square: with \(\vect{\eta}'=\vect{\eta}+P_{22}^{-1}P_{12}\transpose\vect{\xi}\),
as one checks by expanding. The map \((\vect{\xi},\vect{\eta})\mapsto(\vect{\xi},\vect{\eta}')\) is a linear bijection of \(\R^{2n}\), so \(Q>0\) if and only if the right-hand side is positive for every non-zero \((\vect{\xi},\vect{\eta}')\), which is the stated pair of conditions. Conversely \(Q>0\) forces \(P_{22}>0\), by setting \(\vect{\xi}=\vect{0}\), after which the displayed identity applies.
∎With \(\Gamma\) as in Equation (A.932) and \(W_{\psi}\geq0\), the real symmetric form
is positive definite on \(\R^{6}\). Equivalently, \(\Gamma_{\mathrm{r}}<\identity/2s^{2}\) and
Rests on Lemma A.603, Proposition A.609 and Lemma A.610.
Derives Proposition A.611. By Equation (A.923), \(\abs{\braket{\varphi_{\vect{x}_{0},\vect{p}_{0}}}{\psi}} =(\pi s^{2})^{-3/4}\abs{F(\zeta)} \ee^{-\abs{\operatorname{Im}\zeta}^{2}/2s^{2}}\) with \(\zeta=\vect{\xi}+\ii\vect{\eta}\), \(\vect{\xi}=\vect{x}_{0}\) and \(\vect{\eta}=-s^{2}\vect{p}_{0}/\hbar\). That substitution has Jacobian \(\dd^{3}x_{0}\,\dd^{3}p_{0} =(\hbar/s^{2})^{3}\dd^{3}\xi\,\dd^{3}\eta\), so Equation (A.920) says that
Now compute the real part of the exponent from Equation (A.932). Since \(\zeta\cdot\Gamma\zeta =\vect{\xi}\cdot\Gamma\vect{\xi} -\vect{\eta}\cdot\Gamma\vect{\eta} +2\ii\,\vect{\xi}\cdot\Gamma\vect{\eta}\),
with \(L\) real affine. The exponent of Equation (A.939) is therefore \(-2Q(\vect{\xi},\vect{\eta})+2L\), with \(Q\) as in Equation (A.937). An integral \(\int_{\R^{6}}\ee^{-2Q+2L}\) over the whole space is finite if and only if the quadratic part \(Q\) is positive definite: if some \((\vect{\xi}_{0},\vect{\eta}_{0})\neq0\) had \(Q(\vect{\xi}_{0},\vect{\eta}_{0})\leq0\) then along the ray \(t(\vect{\xi}_{0},\vect{\eta}_{0})\) the integrand would be at least \(\ee^{2tL_{0}}\) with \(L_{0}\) the value of the linear part, which is bounded below on the ray only if \(L_{0}\ge 0\) and in either case the integral over a solid cone about the ray diverges, \(Q\) being continuous and hence bounded above on a neighbourhood of the ray. So \(Q>0\), and Lemma A.610 with \(P_{22}=\identity/2s^{2}-\Gamma_{\mathrm{r}}\) and \(P_{12}=-\Gamma_{\mathrm{i}}\) turns that into Equation (A.938) together with \(P_{22}>0\).
∎Recovering the wavefunction
Let \(A\) be complex symmetric with \(A_{\mathrm{r}}>0\) and put \(N=A+\identity/2s^{2}\), so that \(N_{\mathrm{r}}>0\). Then the function \(\psi_{0}=\exp(-\vect{x}\cdot A\vect{x} +\vect{b}\cdot\vect{x}+c)\) has
for suitable \(\vect{\lambda}\) and \(\nu\). Conversely, a complex symmetric \(\Gamma\) arises this way from exactly one such \(A\) if and only if it satisfies the conditions of Proposition A.611. Rests on Lemma A.601, Lemma A.600 and Proposition A.611.
Derives Proposition A.612. Forward. Insert \(\psi_{0}\) into Equation (A.922). Expanding \((\vect{x}-\zeta)\cdot(\vect{x}-\zeta) =\abs{\vect{x}}^{2}_{\ast}-2\zeta\cdot\vect{x}+\zeta\cdot\zeta\), where \(\abs{\vect{x}}^{2}_{\ast}=\vect{x}\cdot\vect{x}\), the exponent is \(-\vect{x}\cdot N\vect{x} +(\vect{b}+\zeta/s^{2})\cdot\vect{x}+c-\zeta\cdot\zeta/2s^{2}\), and Equation (A.914) evaluates the integral as \(C_{N}\exp\bigl[(\vect{b}+\zeta/s^{2})\cdot N^{-1} (\vect{b}+\zeta/s^{2})/4+c-\zeta\cdot\zeta/2s^{2}\bigr]\). The part quadratic in \(\zeta\) is \(\zeta\cdot N^{-1}\zeta/4s^{4}-\zeta\cdot\zeta/2s^{2}\), which is \(-\zeta\cdot\Gamma\zeta\) with \(\Gamma\) as in Equation (A.941); the rest is affine in \(\zeta\) and supplies \(\vect{\lambda}\) and \(\nu\).
Converse. Solving Equation (A.941) for \(N\) gives \(N^{-1}=K\) with
The first condition of Proposition A.611, \(\Gamma_{\mathrm{r}}<\identity/2s^{2}\), is exactly \(K_{\mathrm{r}}>0\), which by Lemma A.600 makes \(K\) invertible; set \(N=K^{-1}\) and \(A=N-\identity/2s^{2}\), complex symmetric and uniquely determined. It remains to show that \(A_{\mathrm{r}}>0\) is equivalent to Equation (A.938). By Equation (A.913),
so, \(A_{\mathrm{r}}=N_{\mathrm{r}}-\identity/2s^{2}\) and both matrices being real symmetric with the inverse order-reversing on positive definite matrices,
Substituting Equation (A.942), the left-hand side is \(2s^{2}\identity-4s^{4}\Gamma_{\mathrm{r}} +4s^{4}\Gamma_{\mathrm{i}} \bigl(\identity/2s^{2}-\Gamma_{\mathrm{r}}\bigr)^{-1} \Gamma_{\mathrm{i}}\), the factors \(4s^{4}\) cancelling against the inverse of \(K_{\mathrm{r}}\); so the last inequality of Equation (A.944) is \(\Gamma_{\mathrm{i}}(\identity/2s^{2}-\Gamma_{\mathrm{r}})^{-1} \Gamma_{\mathrm{i}}<\Gamma_{\mathrm{r}}\), which is Equation (A.938).
∎Proof of Theorem A.598. Derives Theorem A.598. The “if” half is Proposition A.602, which also gives the last sentence of the theorem.
For the “only if” half, let \(W_{\psi}\geq0\) everywhere. By Proposition A.609 the entire function \(F\) of Equation (A.922) is \(\ee^{g}\) with \(g\) a quadratic polynomial, of matrix \(\Gamma\); by Proposition A.611 that \(\Gamma\) satisfies the positivity conditions; and by Proposition A.612 there is therefore exactly one complex symmetric \(A\) with \(A_{\mathrm{r}}>0\), together with a \(\vect{b}\) and a \(c\) matching the affine part, such that the Gaussian \(\psi_{0}=\exp(-\vect{x}\cdot A\vect{x}+\vect{b}\cdot\vect{x}+c)\) has \(F_{0}=F\). Because \(A_{\mathrm{r}}>0\), \(\psi_{0}\) lies in \(L^{2}(\R^{3})\).
Both \(\psi\) and \(\psi_{0}\) therefore have the same overlaps with every coherent state: by Equation (A.923) applied to each of them, and with the same non-vanishing factor \(\kappa\),
for every \(\vect{x}_{0}\) and \(\vect{p}_{0}\). Hence \(\psi-\psi_{0}\in L^{2}(\R^{3})\) is orthogonal to every \(\varphi_{\vect{x}_{0},\vect{p}_{0}}\), that is to every \(W(\zeta) \Omega_{s}\), and those vectors are total in \(L^{2}(\R^{3})\) by Proposition A.592. So \(\psi=\psi_{0}\) in \(L^{2}\), which is Equation (A.910) almost everywhere.
∎Scope, and what is quoted
Nothing. Remark A.596 explains why the Hadamard factorization theorem, named as the import in Remark 25.56, is not needed: Lemma A.608 replaces it and is proved from Theorem 8.20 and Equation (8.13). The other inputs are the spectral theorem for a real symmetric matrix (Theorem 5.80), differentiation under the integral sign (Theorem 7.77), the Cauchy–Riemann criterion (Theorem 8.6), the Gaussian transform pair (Equation (17.42)), the Wigner-function properties (Theorem 25.53, in particular the overlap identity Equation (25.56)), and the totality of the coherent states, proved in Proposition A.592 of The Stone–von Neumann Theorem. Every one of them is proved in this book.
The auxiliary width \(s\) is arbitrary, exactly as in Remark A.588: it enters the definitions of \(\varphi\), \(F\) and \(\Gamma\) and cancels out of the conclusion, which mentions only \(A\). A reader who repeats the argument with a different \(s\) obtains a different \(\Gamma\) and the same \(\psi\).
Two limits, both stated in Remark 25.56 and both worth repeating where the proof can be pointed at.
The theorem is about pure states, and the restriction is essential rather than technical: every step from Proposition A.604 onwards uses that \(W_{\psi}\) is built from a single wavefunction, through the identity Equation (A.919) in which the right-hand side is \(\abs{\braket{\varphi}{\psi}}^{2}\) and not merely a trace. Mixed states with everywhere non-negative Wigner function that are not mixtures of Gaussians do exist, so “non-negative Wigner function” and “classical” are not synonyms in general, and Corollary 25.54 is an exact witness only inside the pure family.
And the attribution stands without a source the reader of this book can check: Hudson's 1974 paper, When is the Wigner quasi-probability density non-negative? (Reports on Mathematical Physics 6, 249–252), has no entry in this bibliography, and the several-variable form used here is due to Soto and Claverie (1983). Both are cited in words only, on the footing of Darboux's memoir in Remark A.81. That is precisely why the proof is written out in full: nothing in The Poisson Algebra and the Canonical Bridge to Quantum Mechanics rests on the attribution.
Hudson's Theorem: the Pure States of Non-Negative Wigner Function discharges the derivation owed at Theorem 25.55 of The Poisson Algebra and the Canonical Bridge to Quantum Mechanics. Its use there is as the converse half of Corollary 25.54: that corollary exhibits states whose Wigner function reaches the most negative value Equation (25.57) permits, and the theorem proved here says that the ones which avoid negativity altogether are exactly the Gaussians — so that within the pure states, negativity of \(W\) and non-Gaussianity are the same property, and an experiment that reconstructs a Wigner function and finds it negative has measured a departure from the Gaussian family and nothing weaker. The two sections of this appendix that serve The Poisson Algebra and the Canonical Bridge to Quantum Mechanics are linked in one direction: The Stone–von Neumann Theorem supplies the totality of the coherent states used at the end of the proof above, and both are built on the same Gaussian, Equation (A.902) there and Equation (A.918) here.
The Jacobi Identity for the Dirac Bracket
This appendix proves the one clause of Theorem 26.21 of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism that the chapter defers: the Dirac bracket Equation (26.25) obeys the Jacobi identity Equation (22.46). Bilinearity, antisymmetry, the Leibniz rule Equation (24.24), the strong vanishing Equation (26.26) and the two statements Equations (26.27) and (26.28) are one-line computations and are done in the chapter; the identity below is what earns the object the name bracket, because without it the correspondence rule Equation (25.31) would send it to something that is not a commutator.
Two independent routes are written out, as the chapter's prose promises. Route A: the direct expansion is the direct expansion: it uses nothing but the Poisson Jacobi identity Equation (22.46), the derivation property Proposition 25.4 and the antisymmetry of \(C^{-1}\), and it is the route that shows where the identity comes from — every one of the four orders in \(C^{-1}\) cancels for its own reason, and three of the four reasons are the Poisson Jacobi identity applied to a different triple. Route B: the bracket of the induced symplectic form is the structural route: the Dirac bracket is the Poisson bracket of the symplectic manifold that Proposition 26.25 constructs on the second-class surface, and the identity is then read off the closure of a two-form. The second is three paragraphs long and is the one to remember; it is written second because it consumes Proposition 26.25, which is proved in the chapter after the Dirac bracket is introduced.
Nothing outside this treatise is used anywhere in this section.
Notation, and the four identities consumed
Throughout, \(\chi_{\alpha}\), \(\alpha=1,\ldots,S\), is a complete set of second-class constraints in the sense of Definition 26.12, \(C_{\alpha\beta}=\pb{\chi_{\alpha}} {\chi_{\beta}}\) is the matrix Equation (26.24), which is invertible on a neighbourhood of the constraint surface by Proposition 26.19 and continuity, and
The antisymmetry of \(D\) is proved in the chapter, in the antisymmetry step of Theorem 26.21: \(C\transpose=-C\) gives \(\left(C^{-1}\right)\transpose=\left(C\transpose\right)^{-1}=-C^{-1}\). That \(S\) is even, so that such a \(C\) can exist at all, is Proposition 26.19. With this notation Equation (26.25) reads
the sign having flipped because \(\pb{\chi_{\beta}}{B}=-B_{\beta}\). Two further abbreviations carry the whole computation:
Note that \(A_{\alpha\beta}\) is not symmetric — its antisymmetric part is Equation (A.949) below — and that the index of \(\widehat{A}\) is raised from the left, so that \(A_{\alpha}D^{\alpha\beta}B_{\beta}=\widehat{A}^{\beta}B_{\beta}\) while \(D^{\alpha\beta}B_{\beta}=-\widehat{B}^{\alpha}\).
Dimensionally there is nothing to watch: \(D^{\alpha\beta}\) carries the inverse of the dimension of \(C_{\alpha\beta}\), so the correction term of Equation (A.947) has exactly the dimension of \(\pb{A}{B}\), whatever SI dimensions the individual \(\chi_{\alpha}\) happen to have. Nothing below depends on the constraints being dimensionally homogeneous among themselves.
For all phase-space functions \(A\), \(B\), \(E\), wherever \(C^{-1}\) is defined,
Rests on Equation (22.46), Proposition 25.4 and Equation (26.24).
Derives Lemma A.616. Three of the four are the Poisson Jacobi identity Equation (22.46) applied to a different triple, and the fourth is the derivative of an inverse.
Equation (A.949). Take the triple \(\left(A,\chi_{\alpha},\chi_{\beta}\right)\) in Equation (22.46):
The second term is \(\pb{\chi_{\alpha}}{-A_{\beta}}=+A_{\beta\alpha}\) and the third is \(\pb{\chi_{\beta}}{A_{\alpha}}=-A_{\alpha\beta}\), which gives Equation (A.949).
Equation (A.950). Take the triple \(\left(\chi_{\delta},B,E\right)\): \(\pb{\chi_{\delta}}{\pb{B}{E}}+\pb{B}{\pb{E}{\chi_{\delta}}} +\pb{E}{\pb{\chi_{\delta}}{B}}=0\), that is \(-\pb{\pb{B}{E}}{\chi_{\delta}}+\pb{B}{E_{\delta}} -\pb{E}{B_{\delta}}=0\).
Equation (A.951). The map \(X\longmapsto\pb{A}{X}\) is a derivation (Proposition 25.4), so applying it to \(D^{\alpha\mu}C_{\mu\beta}=\delta^{\alpha}_{\beta}\) gives \(\pb{A}{D^{\alpha\mu}}C_{\mu\beta} +D^{\alpha\mu}\pb{A}{C_{\mu\beta}}=0\); contracting with \(D^{\beta\nu}\) isolates \(\pb{A}{D^{\alpha\nu}}\). This is the only fact about the inverse matrix that the computation needs, and it is why \(D\) may be kept as an unexpanded symbol throughout.
Equation (A.952). Take the triple \(\left(\chi_{\gamma},\chi_{\alpha},\chi_{\beta}\right)\): each of the three terms of Equation (22.46) is minus one of the \(T\)'s, in the cyclic order stated.
∎Route A: the direct expansion
Wherever \(C_{\alpha\beta}\) is invertible, and for all phase-space functions \(A\), \(B\), \(E\),
strongly — as an identity of functions, not merely on the constraint surface. Rests on Definition 26.20, Equation (22.46) and Lemma A.616.
Derives Theorem A.617. Write \(\mathfrak{S}\) for the sum over the three cyclic images \(\left(A,B,E\right)\to\left(B,E,A\right)\to\left(E,A,B\right)\), so that the left-hand side of Equation (A.954) is \(\mathfrak{S}\,\pb{A}{\pb{B}{E}_{\text{D}}}_{\text{D}}\).
Step 1: expand once and sort by the number of \(D\) factors. Put \(F:=\pb{B}{E}_{\text{D}}=\pb{B}{E}+B_{\alpha}D^{\alpha\beta}E_{\beta}\). By Equation (A.947) applied a second time, \(\pb{A}{F}_{\text{D}}=\pb{A}{F}+A_{\gamma}D^{\gamma\delta}F_{\delta}\) with \(F_{\delta}=\pb{F}{\chi_{\delta}}\). Expanding both pieces by the Leibniz rule (Proposition 25.4) and using Equation (A.950) on \(\pb{\pb{B}{E}}{\chi_{\delta}}\),
Assembling, and using Equation (A.951) on \(\pb{A}{D^{\alpha\beta}}\) and on \(\pb{D^{\alpha\beta}}{\chi_{\delta}}\), the eight surviving terms sort by the number of factors of \(D\) they carry:
The three order-two terms and the order-three term are obtained by substituting Equation (A.951) and then absorbing each \(A_{\gamma}D^{\gamma\delta}\) and \(B_{\alpha}D^{\alpha\mu}\) into the hatted abbreviation of Equation (A.948), remembering from it that \(D^{\nu\beta}E_{\beta}=-\widehat{E}^{\nu}\), which is where the sign of the first order-two term and of the order-three term comes from.
Step 2: order zero. \(\mathfrak{S}\,\pb{A}{\pb{B}{E}}=0\) is the Poisson Jacobi identity Equation (22.46) itself.
Step 3: order one. Introduce the trilinear object
The first order-one term is \(G(A,B,E)\). For the second, relabel \(\alpha\leftrightarrow\beta\) and use \(D^{\beta\alpha}=-D^{\alpha\beta}\): \(B_{\alpha}D^{\alpha\beta}\pb{A}{E_{\beta}}=-G(A,E,B)\). For the third, the same manoeuvre on each of the two pieces gives \(A_{\gamma}D^{\gamma\delta}\pb{B}{E_{\delta}}=-G(B,E,A)\) and \(-A_{\gamma}D^{\gamma\delta}\pb{E}{B_{\delta}}=+G(E,B,A)\). The order-one part of Equation (A.957) is therefore
Now apply \(\mathfrak{S}\). A cyclic sum of \(G\) is a sum of three terms that depends only on which of the two orientations of the triple its arguments carry, so \(\mathfrak{S}G(B,E,A)=\mathfrak{S}G(A,B,E)\) and \(\mathfrak{S}G(E,B,A)=\mathfrak{S}G(A,E,B)\): in each case the three summands are the same three terms in a different order. Hence the first term of Equation (A.959) cancels the third and the second cancels the fourth, and the cyclic sum vanishes. Concretely, the summand \(G(A,B,E)\) produced by the first term of the \(\left(A,B,E\right)\) copy is cancelled by the summand \(-G(A,B,E)\) produced by the third term of the \(\left(E,A,B\right)\) copy; the pairing of the second and fourth terms is the same one cycle over. This is the step that a reader cannot reconstruct without being told which term meets which, and it is the reason Equation (A.950) had to be used in Step 1: without it the third order-one term is a bracket of a bracket and does not have the shape Equation (A.958) at all.
Step 4: order two. Replace \(\pb{A}{C_{\mu\nu}}\) by \(A_{\mu\nu}-A_{\nu\mu}\), which is Equation (A.949), and relabel so that every term carries the free index pair in the order \(\left(p,q\right)\):
the second term coming from \(-\widehat{B}^{p}\widehat{E}^{q}A_{qp}\) by the interchange \(p\leftrightarrow q\). Now collect the cyclic sum by which of \(A_{pq}\), \(B_{pq}\), \(E_{pq}\) each term carries. The coefficient of \(A_{pq}\) receives \(\widehat{B}^{p}\widehat{E}^{q}-\widehat{B}^{q}\widehat{E}^{p}\) from \(X(A,B,E)\), \(+\widehat{B}^{q}\widehat{E}^{p}\) from the fourth term of \(X(B,E,A)\) and \(-\widehat{E}^{q}\widehat{B}^{p}\) from the third term of \(X(E,A,B)\); the four contributions cancel in pairs. The coefficients of \(B_{pq}\) and of \(E_{pq}\) are the same computation read one and two cycles on, so \(\mathfrak{S}X=0\).
Step 5: order three. By Equation (A.952) the totally contracted coefficient is a cyclic sum of \(T\). Writing the order-three term of Equation (A.957) as \(\widehat{A}^{p}\widehat{B}^{q}\widehat{E}^{r}T_{qrp}\) and applying \(\mathfrak{S}\), which permutes \(\widehat{A},\widehat{B},\widehat{E}\) cyclically, the three summands are \(\widehat{A}^{p}\widehat{B}^{q}\widehat{E}^{r}\) times, in turn, \(T_{qrp}\), \(T_{rpq}\) and \(T_{pqr}\). Their sum vanishes by Equation (A.952).
All four orders vanish separately, so Equation (A.954) holds. No constraint was set to zero anywhere in the argument, so the identity is strong, as claimed.
∎It is worth recording the audit, because the four cancellations look alike and are not. Order zero is the Poisson Jacobi identity on the triple \(\left(A,B,E\right)\). Order one is the Poisson Jacobi identity on \(\left(\chi_{\delta},B,E\right)\), which turns a bracket of a bracket into two terms of the shape Equation (A.958), plus the antisymmetry of \(D\). Order two is the Poisson Jacobi identity on \(\left(A,\chi_{\alpha},\chi_{\beta}\right)\), which is what makes the derivative of the inverse matrix collapse onto the second brackets \(A_{\alpha\beta}\) already present in the sum, plus the antisymmetry of \(D\). Order three is the Poisson Jacobi identity on three constraints \(\left(\chi_{\gamma},\chi_{\alpha},\chi_{\beta}\right)\). The antisymmetry of \(D\) enters at orders one, two and three, and never on its own: at every order it does no more than allow the terms to be written with their indices in a common order, and the vanishing is always supplied by Equation (22.46). So the Dirac bracket inherits its Jacobi identity from the Poisson bracket four times over, once for each way of feeding a constraint into the triple.
Route B: the bracket of the induced symplectic form
The second proof is shorter, gives the reason rather than the verification, and is the one worth remembering. Its input is Proposition 26.25, proved in the chapter: with \(C_{\alpha\beta}\) invertible at \(x\), the tangent space splits as \(T_{x}M=W\oplus T_{x}\Sigma_{\chi}\) along the projector Equation (26.42), the two summands are \(\omega\)-orthogonal, and the restriction \(\omega_{\Sigma}\) of the symplectic form to \(T_{x}\Sigma_{\chi}\) is non-degenerate, so that \(\left(\Sigma_{\chi},\omega_{\Sigma}\right)\) is a symplectic manifold.
Let \(X_{A}\) be the Hamiltonian vector field of \(A\) in the ambient space, Equation (24.5), and let \(P\) be the projector Equation (26.42). Then, for all \(A\) and \(B\),
On \(\Sigma_{\chi}\) the vector field \(P(X_{A})\) is the Hamiltonian vector field of the restriction \(A|_{\Sigma_{\chi}}\) with respect to \(\omega_{\Sigma}\), and consequently
Rests on Proposition 26.25, Equation (24.7) and Equation (26.42).
Derives Proposition A.619. By Equation (24.7), \(\pb{u}{v}=\dd u\left(X_{v}\right)\) for every pair of functions. Hence \(\pb{B}{A}=\dd B\left(X_{A}\right)\), \(\pb{B}{\chi_{\alpha}}=\dd B\left(X_{\chi_{\alpha}}\right)\) and \(\pb{\chi_{\beta}}{A}=\dd\chi_{\beta}\left(X_{A}\right)\), so that Equation (26.25) reads
and the argument of \(\dd B\) is exactly \(P\left(X_{A}\right)\) of Equation (26.42). That is the whole of Equation (A.961), and it is the whole point: the Dirac bracket differs from the Poisson bracket only in that the Hamiltonian vector field is projected onto the constraint surface before it is used.
For the second statement, let \(v\in T_{x}\Sigma_{\chi}\). The \(\omega\)-orthogonality of the splitting gives \(\omega\!\left(X_{\chi_{\alpha}},v\right)=\dd\chi_{\alpha}(v)=0\), so
Since \(P\left(X_{A}\right)\) is tangent to \(\Sigma_{\chi}\) and \(\dd A(v)\) for tangent \(v\) depends on \(A\) only through \(A|_{\Sigma_{\chi}}\), this says precisely that \(P\left(X_{A}\right)\) is the \(\omega_{\Sigma}\)-Hamiltonian vector field of \(A|_{\Sigma_{\chi}}\). Feeding that back into Equation (A.961) with the roles of \(A\) and \(B\) exchanged, and using antisymmetry of both brackets, gives Equation (A.962).
∎Second derivation of the Jacobi identity. Derives Theorem A.617. This is a second, independent proof of Theorem A.617. Fix a point \(x\) at which \(C_{\alpha\beta}\) is invertible, and let \(c_{\alpha}:=\chi_{\alpha}(x)\). The shifted functions \(\chi_{\alpha}-c_{\alpha}\) have the same brackets as the \(\chi_{\alpha}\), hence the same invertible matrix \(C_{\alpha\beta}\), and they define the level surface \(\Sigma_{c}\) through \(x\). Everything in Proposition 26.25 and in Proposition A.619 therefore applies verbatim to \(\Sigma_{c}\), and the Dirac bracket built from \(\chi_{\alpha}-c_{\alpha}\) is, term by term in Equation (26.25), the Dirac bracket built from the \(\chi_{\alpha}\): constants drop out of every bracket. So the level surfaces of the constraints foliate a neighbourhood of \(\Sigma_{\chi}\) by symplectic manifolds \(\left(\Sigma_{c},\omega_{c}\right)\), and by Equation (A.962) the Dirac bracket of two functions restricted to \(\Sigma_{c}\) is the Poisson bracket of \(\left(\Sigma_{c},\omega_{c}\right)\) applied to their restrictions.
Now \(\omega_{c}\) is the pullback \(i_{c}^{*}\omega\) of the ambient symplectic form along the inclusion, so \(\dd\omega_{c}=i_{c}^{*}\dd\omega=0\): it is closed because \(\omega\) is, and non-degenerate by Proposition 26.25. A symplectic manifold's Poisson bracket satisfies the Jacobi identity — that is Equation (22.46) on \(\left(\Sigma_{c},\omega_{c}\right)\), which by Theorem 24.12 may be computed in canonical coordinates, where it is the same computation as in the ambient space. Hence the cyclic sum Equation (A.954), restricted to \(\Sigma_{c}\), vanishes.
Finally, the value of the cyclic sum at \(x\) is determined by its restriction to the leaf through \(x\): by Equation (A.962) each of its three terms is, at \(x\), the corresponding double bracket of \(\left(\Sigma_{c},\omega_{c}\right)\) applied to restrictions. Since \(x\) was an arbitrary point of the neighbourhood on which \(C^{-1}\) exists, Equation (A.954) holds there identically.
∎The last paragraph is the step at which a shorter-looking argument fails, and the failure is worth naming. It is tempting to say that Equation (26.26) lets one add any multiple of a constraint to either argument without changing the Dirac bracket, so that an identity holding on \(\Sigma_{\chi}\) holds everywhere. It does not: by the Leibniz rule, \(\pb{A}{c^{\alpha}\chi_{\alpha}}_{\text{D}} =\chi_{\alpha}\pb{A}{c^{\alpha}}_{\text{D}}\), which vanishes on \(\Sigma_{\chi}\) and not off it. What Equation (26.26) does say — and this is the right reading — is that every \(\chi_{\alpha}\) is a Casimir of the Dirac bracket, so its level sets are unions of symplectic leaves, and the foliation used above is not an artifice but the canonical structure the bracket itself defines. The Dirac bracket makes a neighbourhood of \(\Sigma_{\chi}\) a Poisson manifold in the sense of Definition 24.36, whose symplectic leaves are the \(\Sigma_{c}\); Theorem A.617 is the assertion that it is a Poisson bracket at all.
Nothing. Both routes are carried out from material proved in this treatise: the Poisson Jacobi identity Equation (22.46), the derivation property Proposition 25.4, the relation between the bracket and the symplectic form Equation (24.7), Darboux's theorem Theorem 24.12 in the form proved in Darboux's Theorem: Local Canonical Coordinates, and the splitting of Proposition 26.25. The one hypothesis that is not proved but assumed is the regularity of the constraint set (Remark 26.7), which is what makes \(\Sigma_{\chi}\) a submanifold in the first place; it is an assumption about the system under study and is stated as such in the chapter.
The Jacobi Identity for the Dirac Bracket discharges the derivation owed at Theorem 26.21 of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism, whose proof establishes every other property of the Dirac bracket and defers this one. With it the object of Definition 26.20 is a bracket in the full sense, so that Postulate 26.47 may send it to a commutator, and the three facts the chapter draws together are seen to be one fact in three costumes: the Jacobi identity proved here, the degree-of-freedom count of Theorem 26.17, and the reduced Liouville measure of Proposition 26.25 are the tangential, the dimensional and the volumetric readings of the single statement that \(\left(\Sigma_{\chi},\omega_{\Sigma}\right)\) is a symplectic manifold. Remark 26.26 makes the same point from the measure side; Example 26.22 is the worked instance, where the leaves are the tangent bundles of a surface in \(\R^{3}\) and the reduced bracket is the one Lagrangian Mechanics obtains by eliminating a coordinate.
The Gauss–Codazzi Form of the Einstein–Hilbert Lagrangian
This appendix proves Proposition 26.41 of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism: that the \(3+1\) decomposition Equation (26.78) gives \(\sqrt{\abs{g}}=N\sqrt{h}\), that the four-dimensional curvature scalar splits into the intrinsic curvature of the leaf, a quadratic in the extrinsic curvature and two divergences, and — the conclusion the chapter actually consumes — that no time derivative of the lapse or of the shift survives anywhere in the Einstein–Hilbert Lagrangian.
The mathematics is the Gauss–Codazzi decomposition of the Riemann tensor of an ambient manifold along a hypersurface. The difficulty is not the decomposition, which is four projections of one identity; it is the signature bookkeeping, and this treatise sits in exactly the place where that bookkeeping is worst. The metric of Minkowski Space and Its Symmetries is mostly minus, so the unit normal to a spacelike leaf satisfies \(n\cdot n=+1\); the curvature conventions of Theorem 13.152 and Definition 13.153 are the ones for which the round two-sphere has \(R>0\); and Definition 26.40 takes \(h_{ij}\) to be the positive-definite Riemannian metric of the leaf, so that the metric \(g_{ij}\) actually induced on the leaf by Equation (26.78) is \(-h_{ij}\). Three sign conventions, and every formula in the literature is written with a different combination of them. Nothing below is inherited from a mostly-plus source. The discipline used instead is this. Write
carry \(\epsilon\) as a symbol through every step, do the whole computation with the induced metric \(\gamma_{\mu\nu}=g_{\mu\nu}-\epsilon n_{\mu}n_{\nu}\) — whatever its sign — and only at the very end substitute \(\epsilon=+1\) and \(\gamma_{ij}=-h_{ij}\). A reader checking against a mostly-plus text sets \(\epsilon=-1\) and \(\gamma_{ij}=+h_{ij}\) and compares line by line; Remark A.633 records what happens when the substitution is made, which is the one place where the answer differs from the printed statement of Proposition 26.41 and the reason the chapter's prose insists that the sign be settled here rather than assumed.
Throughout, the ambient connection \(\nabla\) is the Levi-Civita connection of \(g_{\mu\nu}\) (Theorem 13.150 with vanishing torsion), \(R^{\lambda}{}_{\rho\mu\nu}\) is Equation (13.306), and \(R_{\mu\nu}\), \(R\) are the contractions Equation (13.307). Indices \(\mu,\nu,\ldots\) run over \(0,1,2,3\) and \(i,j,k,\ldots\) over \(1,2,3\); \(x^{0}=ct\), as in Notation 26.28. Since \(h_{ij}\) and \(N\), \(N^{i}\) are dimensionless and the \(x^{i}\) carry metres, every term of Equation (26.80) has the SI dimension \(/\mathrm{m}^{2}\), and the identity is dimensionally homogeneous as it stands; no factor of \(c\) or \(G\) enters until the Lagrangian is multiplied by \(1/2\kappa\).
The kit: normal, inverse metric, projector, determinant
Write Equation (26.78) in terms of the coframe
Then, with \(N>0\):
-
the dual frame is \(e_{0}=\pp_{0}-N^{i}\pp_{i}\), \(e_{i}=\pp_{i}\), and the inverse metric is
\begin{equation}\tag{A.967} g^{00}=\frac{1}{N^{2}}\ec\qquad g^{0i}=-\frac{N^{i}}{N^{2}}\ec\qquad g^{ij}=\frac{N^{i}N^{j}}{N^{2}}-h^{ij}\ec \end{equation}with \(h^{ij}\) the inverse of \(h_{ij}\);
-
the future-directed unit normal to the leaves \(x^{0}=\text{const}\) is
\begin{equation}\tag{A.968} n_{\mu}=N\delta^{0}_{\mu}\ec\qquad n^{\mu}=\left(\frac{1}{N},\,-\frac{N^{i}}{N}\right) =\frac{1}{N}e_{0}\ec\qquad \epsilon=n^{\mu}n_{\mu}=+1\ec \end{equation}so the normal is timelike and the leaves are spacelike;
-
\(\det g=-N^{2}h\) with \(h=\det h_{ij}>0\), hence
\begin{equation}\tag{A.969} \sqrt{\abs{g}}=N\sqrt{h}\ep \end{equation}
Rests on Definition 26.40, Equation (26.78) and Definition 13.117.
Derives Lemma A.623. (i). That Equation (A.966) reproduces Equation (26.78) is the definition of \(\theta^{i}\). For the dual frame, \(\theta^{0}(e_{0})=1\), \(\theta^{i}(e_{0})=-N^{i}+N^{i}=0\), \(\theta^{0}(e_{i})=0\) and \(\theta^{j}(e_{i})=\delta^{j}_{i}\), so \(\set{e_{0},e_{i}}\) is indeed dual to \(\set{\theta^{0},\theta^{i}}\). In that basis the metric is the block array \(\diag\left(N^{2},-h_{ij}\right)\), whose inverse is the block array \(\diag\left(N^{-2},-h^{ij}\right)\); hence
and expanding \(e_{0}=\pp_{0}-N^{k}\pp_{k}\) in coordinates gives Equation (A.967).
(ii). A covector annihilating every vector tangent to a leaf — that is, every \(\pp_{i}\) — is proportional to \(\delta^{0}_{\mu}\). Normalizing, \(g^{\mu\nu}\left(N\delta^{0}_{\mu}\right) \left(N\delta^{0}_{\nu}\right)=N^{2}g^{00}=1\) by Equation (A.967), so Equation (A.968) is a unit covector and \(\epsilon=+1\). Raising with Equation (A.967), \(n^{\mu}=Ng^{\mu0}\) gives the components stated, and \(n^{0}=1/N>0\) makes it future directed.
(iii). This is the step Proposition 26.41 asserts and does not prove. The change of coframe Equation (A.966) has the matrix \(\theta^{a}=\Lambda^{a}{}_{\mu}\dd x^{\mu}\) with \(\Lambda^{0}{}_{0}=1\), \(\Lambda^{0}{}_{i}=0\), \(\Lambda^{i}{}_{0}=N^{i}\), \(\Lambda^{i}{}_{j}=\delta^{i}_{j}\): it is lower triangular with unit diagonal, so \(\det\Lambda=1\). Since \(g_{\mu\nu}=\Lambda^{a}{}_{\mu}\hat{g}_{ab}\Lambda^{b}{}_{\nu}\) with \(\hat{g}=\diag\left(N^{2},-h_{ij}\right)\),
the factor \(\left(-1\right)^{3}\) being the three minus signs the signature puts on the spatial block. Completing the square in \(\dd x^{i}+N^{i}\dd x^{0}\) is thus exactly the statement that the shift can be removed from the determinant by a unimodular change of coframe.
∎The projector onto a leaf and the induced metric are
A tensor is tangential if every index is annihilated by contraction with \(n\), equivalently if projecting every index returns it. For a tangential tensor \(T\) the intrinsic covariant derivative is
Rests on Equation (A.968) and Definition 13.145.
Three facts about Equation (A.972) are used without further comment. First, \(\gamma^{\mu}{}_{\nu}n^{\nu}=n^{\mu}-\epsilon^{2}n^{\mu}=0\) and \(\gamma^{\mu}{}_{\nu}v^{\nu}=v^{\mu}\) for \(v\) tangent to a leaf, since such a \(v\) has \(n_{\nu}v^{\nu}=0\); so \(\gamma\) is the projection along \(n\), and \(\gamma^{\mu}{}_{\alpha}\gamma^{\alpha}{}_{\nu} =\gamma^{\mu}{}_{\nu}\). Second, \(D\) is the Levi-Civita connection of \(\gamma\) on the leaf: it is torsion free, and \(D_{\mu}\gamma_{\nu\rho}=\gamma\gamma\gamma\nabla \left(g-\epsilon nn\right)=0\), the first term because \(\nabla g=0\) and the second because each projector annihilates the \(n\) it meets. Third — and this is the sign that will have to be paid for at the end — in the coordinates of Lemma A.623 one has \(n_{i}=0\), so
where \(\gamma^{ij}\) means the inverse of \(\gamma_{ij}\) as a three-by-three matrix, which is also \(\gamma^{\mu\nu}\) restricted to spatial indices. The induced metric of Equation (26.78) is negative definite; the positive-definite \(h_{ij}\) of Definition 26.40 is its negative.
The extrinsic curvature, and the dictionary
Put
Rests on Equation (A.972) and Definition 13.145.
\(\mathcal{K}_{\mu\nu}\) and \(a_{\mu}\) are tangential, \(\mathcal{K}\) is symmetric, and
Rests on Definition A.625, Lemma A.623 and Equation (13.295).
Derives Lemma A.626. Tangentiality of \(\mathcal{K}\) is built into Equation (A.975). For \(a\), differentiate \(n^{\nu}n_{\nu}=\epsilon\): \(n^{\nu}\nabla_{\mu}n_{\nu}=0\), and contracting with \(n^{\mu}\) gives \(n^{\nu}a_{\nu}=0\).
Equation (A.976). Insert \(\delta^{\alpha}_{\mu}=\gamma^{\alpha}{}_{\mu} +\epsilon n^{\alpha}n_{\mu}\) in both slots of \(\nabla_{\mu}n_{\nu}\). The \(\gamma\gamma\) term is \(\mathcal{K}_{\mu\nu}\). Any term carrying \(n^{\beta}\) in the second slot vanishes, because \(n^{\beta}\nabla_{\alpha}n_{\beta}=0\). The remaining term is \(\epsilon n^{\alpha}n_{\mu}\gamma^{\beta}{}_{\nu}\nabla_{\alpha}n_{\beta} =\epsilon n_{\mu}\gamma^{\beta}{}_{\nu}a_{\beta} =\epsilon n_{\mu}a_{\nu}\), since \(a\) is tangential.
Symmetry of \(\mathcal{K}\) follows: the leaves are level sets of the coordinate \(x^{0}\), so \(n_{\mu}/N=\pp_{\mu}x^{0}\) is a gradient and \(\nabla_{[\mu}\left(n_{\nu]}/N\right)=0\), whence
Projecting both slots kills the right-hand side, so \(\mathcal{K}_{\mu\nu}=\mathcal{K}_{\nu\mu}\).
Equations (A.977) and (A.978). Trace Equation (A.976) with \(g^{\mu\nu}\): the second term gives \(\epsilon n^{\nu}a_{\nu}=0\) and the first gives \(g^{\mu\nu}\mathcal{K}_{\mu\nu}=\gamma^{\mu\nu}\mathcal{K}_{\mu\nu} =\mathcal{K}\), because \(\mathcal{K}\) is tangential and \(g\) and \(\gamma\) agree on tangential tensors. Squaring Equation (A.976),
each cross term carrying an \(n\) contracted with a tangential index and the last term carrying \(a^{\nu}n_{\nu}=0\).
Equation (A.979). Compare Equation (A.980) with the antisymmetric part of Equation (A.976), which is \(\epsilon\left(n_{\mu}a_{\nu}-n_{\nu}a_{\mu}\right)\), and contract with \(n^{\mu}\). The left-hand side gives \(\epsilon\left(\epsilon a_{\nu}-0\right)=a_{\nu}\) and the right-hand side gives \(n_{\nu}\left(n^{\mu}\pp_{\mu}\ln N\right) -\epsilon\pp_{\nu}\ln N\). Projecting with \(\gamma\), which fixes the tangential \(a_{\nu}\) and kills \(n_{\nu}\), leaves Equation (A.979).
∎In the coordinates of Lemma A.623, with \(K_{ij}\) the extrinsic curvature Equation (26.79) of Definition 26.40, \(K=h^{ij}K_{ij}\), \(K^{ij}=h^{ik}h^{jl}K_{kl}\) and \({}^{(3)}\!R\) the Ricci scalar of \(h_{ij}\):
\(R[\gamma]\) denoting the curvature scalar Equation (13.307) of the induced metric \(\gamma_{ij}\) of Equation (A.974). The Levi-Civita connections of \(\gamma_{ij}\) and of \(h_{ij}\) coincide, so \(D_{i}\) of Equation (A.973) is the \(D_{i}\) of Definition 26.40. Rests on Equation (26.79), Equation (A.974) and Definition 13.153.
Derives Lemma A.627. Connections and Riemann tensors. \(\gamma_{ij}=-h_{ij}\) differs from \(h_{ij}\) by a constant factor, and the Christoffel symbols Equation (13.299) are homogeneous of degree zero in the metric — one inverse metric against one derivative of the metric — so they are the same for \(\gamma\) and for \(h\). By Equation (13.306) the Riemann tensors \(R^{i}{}_{jkl}\) and hence the Ricci tensors \(R_{jl}\) agree. The scalars do not: \(R[\gamma]=\gamma^{jl}R_{jl}=-h^{jl}R_{jl}=-{}^{(3)}\!R\), which is the fourth relation.
Extrinsic curvature. Since \(n_{i}=0\), \(\gamma^{\alpha}{}_{i} =\delta^{\alpha}_{i}\), so \(\mathcal{K}_{ij}=\nabla_{i}n_{j} =\pp_{i}\left(N\delta^{0}_{j}\right) -\Gamma^{\lambda}{}_{ij}n_{\lambda}=-N\Gamma^{0}{}_{ij}\). Compute \(\Gamma^{0}{}_{ij}\) from Equation (13.299) with \(g_{0i}=-N_{i}\), \(g_{ij}=-h_{ij}\) and Equation (A.967):
every term having picked up one overall minus from \(g_{ij}=-h_{ij}\), \(g_{0i}=-N_{i}\) and, in the second group, a further minus from \(g^{0k}=-N^{k}/N^{2}\). Now the Christoffel symbols of \(h_{ij}\) give \(D_{i}N_{j}+D_{j}N_{i}=\pp_{i}N_{j}+\pp_{j}N_{i} -2\Gamma^{(3)k}{}_{ij}N_{k}\) and \(2\Gamma^{(3)k}{}_{ij}N_{k} =N^{l}\left(\pp_{i}h_{lj}+\pp_{j}h_{li}-\pp_{l}h_{ij}\right)\), so the bracket of Equation (A.983) is \(\pp_{0}h_{ij}-D_{i}N_{j}-D_{j}N_{i}\) and
by Equation (26.79). Hence \(\mathcal{K}_{ij}=-K_{ij}\), the first relation. This is not a choice: it is forced, because \(\mathcal{K}\) is one half the Lie derivative of \(\gamma\) along \(n\) while \(K\) is one half the Lie derivative of \(h\), and \(\gamma=-h\).
Trace and square. \(\mathcal{K}=\gamma^{ij}\mathcal{K}_{ij} =\left(-h^{ij}\right)\left(-K_{ij}\right)=K\): the two minus signs cancel, so the traces agree even though the tensors differ in sign. Likewise \(\mathcal{K}_{\mu\nu}\mathcal{K}^{\mu\nu} =\gamma^{ik}\gamma^{jl}\mathcal{K}_{ij}\mathcal{K}_{kl} =h^{ik}h^{jl}K_{ij}K_{kl}=K_{ij}K^{ij}\), four minus signs cancelling in pairs. Only the curvature scalar, which carries one inverse metric rather than an even number, survives the substitution with its sign changed — and it is exactly that asymmetry which Remark A.633 is about.
∎The Gauss equation
For a hypersurface with unit normal \(n\), \(\epsilon=n\cdot n\),
\(R[\gamma]\) being the Riemann tensor of the induced metric. Rests on Equation (13.305), Definition A.625 and Equation (A.973).
Derives Lemma A.628. Let \(\omega_{\mu}\) be a tangential covector field. By Equation (A.973), \(\left(D\omega\right)_{\alpha\beta} =\gamma^{\sigma}{}_{\alpha}\gamma^{\tau}{}_{\beta} \nabla_{\sigma}\omega_{\tau}\), and applying \(D\) once more,
Expand the derivative by the Leibniz rule. Because \(\nabla_{\lambda}\gamma^{\sigma}{}_{\alpha} =-\epsilon\left(n^{\sigma}\nabla_{\lambda}n_{\alpha} +n_{\alpha}\nabla_{\lambda}n^{\sigma}\right)\), and because every \(n_{\alpha}\) or \(n_{\beta}\) meets a projector and dies, the two terms in which \(\nabla\) falls on a projector are
The first is symmetric in \(\rho\leftrightarrow\mu\) and therefore drops out of the commutator \(D_{\rho}D_{\mu}-D_{\mu}D_{\rho}\). In the second, differentiate \(\omega_{\tau}n^{\tau}=0\) to get \(n^{\tau}\nabla_{\sigma}\omega_{\tau} =-\omega_{\tau}\nabla_{\sigma}n^{\tau}\), and then \(\gamma^{\sigma}{}_{\mu}\omega_{\tau}\nabla_{\sigma}n^{\tau} =\omega_{\tau}\mathcal{K}_{\mu}{}^{\tau}\) because \(\omega\) is tangential; so the second term equals \(+\epsilon\,\mathcal{K}_{\rho\nu}\mathcal{K}_{\mu}{}^{\tau} \omega_{\tau}\).
The remaining term of Equation (A.986) is \(\gamma^{\lambda}{}_{\rho}\gamma^{\sigma}{}_{\mu}\gamma^{\tau}{}_{\nu} \nabla_{\lambda}\nabla_{\sigma}\omega_{\tau}\). Antisymmetrizing in \(\rho\leftrightarrow\mu\) and using the Ricci identity Equation (13.305) in its covector form, \(\comm{\nabla_{\lambda}}{\nabla_{\sigma}}\omega_{\tau} =-R^{\kappa}{}_{\tau\lambda\sigma}\omega_{\kappa}\) (which follows from Equation (13.305) by applying the commutator to the scalar \(\omega_{\lambda}V^{\lambda}\)), and the same identity on the leaf for \(D\),
Both sides are tangential in the free index \(\kappa\) and \(\omega\) was an arbitrary tangential covector, so the coefficients agree after projecting \(\kappa\), which is Equation (A.985).
∎The sign of the quadratic term in Equation (A.985) can be checked on the one case every reader already knows, and in the signature that matters here. Take the sphere of radius \(r\) in Euclidean \(\R^{3}\): the ambient is flat, \(\epsilon=+1\) because the normal is a unit vector of a positive-definite metric, and \(\mathcal{K}_{ij}=\nabla_{i}n_{j}=\gamma_{ij}/r\) for the outward normal. Then Equation (A.985), with indices lowered, gives \(R[\gamma]_{\kappa\nu\rho\mu} =r^{-2}\left(\gamma_{\mu\nu}\gamma_{\rho\kappa} -\gamma_{\rho\nu}\gamma_{\mu\kappa}\right)\), whose double contraction is \(R[\gamma]=2/r^{2}\). That is \(2\) times the Gaussian curvature, which is what Definition 13.153 says a two-dimensional curvature scalar must be, and it is Gauss's theorema egregium: the sphere's intrinsic curvature is fixed by its second fundamental form alone. A sign error in Equation (A.985) would make the sphere intrinsically hyperbolic.
The Codazzi equation
With the same notation,
and contracting \(\nu\) with \(\rho\),
No factor of \(\epsilon\) appears in either. Rests on Lemma A.626, Equation (13.305) and Definition A.625.
Derives Lemma A.630. By Equation (A.976), \(\mathcal{K}_{\beta\lambda}=\nabla_{\beta}n_{\lambda} -\epsilon n_{\beta}a_{\lambda}\), so
the second term because \(\gamma^{\beta}{}_{\nu}\nabla_{\alpha}\left(n_{\beta}a_{\lambda}\right) =a_{\lambda}\gamma^{\beta}{}_{\nu}\nabla_{\alpha}n_{\beta}\), the other piece dying on \(\gamma^{\beta}{}_{\nu}n_{\beta}=0\). Antisymmetrizing in \(\mu\leftrightarrow\nu\), the term \(a_{\rho}\mathcal{K}_{\mu\nu}\) is symmetric and drops, and the Ricci identity in covector form leaves Equation (A.989).
For the contraction, apply \(\gamma^{\nu\rho}\). On the left it gives \(D_{\mu}\mathcal{K}-D_{\nu}\mathcal{K}^{\nu}{}_{\mu}\), which is minus the left-hand side of Equation (A.990). On the right, \(\gamma^{\beta\lambda}=g^{\beta\lambda}-\epsilon n^{\beta}n^{\lambda}\); the \(g^{\beta\lambda}\) piece contracts the second and fourth slots of \(R_{\kappa\lambda\alpha\beta}\), which by the antisymmetry within each index pair (Proposition 13.154) is the same as contracting the first and third, that is again the Ricci tensor, giving \(n^{\kappa}R_{\kappa\alpha}\), while the \(n^{\beta}n^{\lambda}\) piece vanishes because \(R_{\kappa\lambda\alpha\beta}n^{\kappa}n^{\lambda}=0\) by antisymmetry in the first pair. Rearranging gives Equation (A.990).
∎Equation (A.990) is worth reading before it is used. Under the dictionary Equation (A.982) the mixed-index extrinsic curvature carries two inverse metrics' worth of sign and is unchanged, \(\mathcal{K}^{j}{}_{i}=\gamma^{jk}\mathcal{K}_{ki} =\left(-h^{jk}\right)\left(-K_{ki}\right)=K^{j}{}_{i}\), so the left-hand side is \(D_{j}K^{j}{}_{i}-D_{i}K\). By Equation (26.82) that is \(2\kappa h^{-1/2}D_{j}\pi^{j}{}_{i} =-\kappa h^{-1/2}\Ham_{i}\) with \(\Ham_{i}\) the momentum constraint Equation (26.85); and the right-hand side is a projection of the Ricci tensor onto one normal and one tangential index. So the momentum constraint of Theorem 26.42 is the contracted Codazzi equation, which is the precise content of the first paragraph of Remark 26.44: the constraints are the \(G^{0}{}_{\mu}\) components of the field equations, and they are conditions on data laid down on one leaf because Codazzi's equation contains no derivative off the leaf.
The twice-normal contraction
The double contraction of Equation (A.985) will leave behind the term \(R_{\mu\nu}n^{\mu}n^{\nu}\), which is neither intrinsic nor a projection of anything on the leaf. Converting it is the step that produces the total derivatives, and it is the one usually passed over. It is not passed over here.
For any unit normal field,
Rests on Equation (13.305), Lemma A.626 and Definition 13.153.
Derives Lemma A.631. Apply the Ricci identity Equation (13.305) to \(n^{\lambda}\), set \(\lambda=\mu\) and contract: since \(R^{\mu}{}_{\rho\mu\nu}=R_{\rho\nu}\) by Equation (13.307),
Contract with \(n^{\nu}\) and rewrite each side as a divergence minus the term the Leibniz rule left over:
Subtracting, and inserting Equations (A.977) and (A.978) for the two quadratic terms and \(a^{\mu}=n^{\nu}\nabla_{\nu}n^{\mu}\) for the first divergence, gives Equation (A.992).
∎Equation (A.992) is the trace of what is called the Ricci, or Mainardi, equation — the twice-normal projection of the Riemann tensor. Only the trace is needed below, and only the trace is proved: the untraced equation would be an identity for \(\gamma^{\alpha}{}_{\mu}\gamma^{\beta}{}_{\nu}n^{\lambda}n^{\sigma} R_{\alpha\lambda\beta\sigma}\) in terms of the Lie derivative of \(\mathcal{K}\) along \(n\), and nothing in this appendix or in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism consumes it.
Assembling the Lagrangian
For a foliation by hypersurfaces with unit normal \(n\) and \(\epsilon=n\cdot n\),
Specialized to Equation (26.78) by Lemmas A.623 and A.627, that is \(\epsilon=+1\), \(\gamma_{ij}=-h_{ij}\),
with \(\sqrt{\abs{g}}=N\sqrt{h}\). In particular \({}^{(3)}\!R\) contains \(h_{ij}\) and its spatial derivatives only, and \(K_{ij}\) contains \(\pp_{0}h_{ij}\) linearly and no other time derivative, so neither \(\pp_{0}N\) nor \(\pp_{0}N^{i}\) occurs anywhere on the right. Rests on Lemmas A.627, A.628 and A.631.
Derives Theorem A.632. Step 1: contract the Gauss equation once. Put \(\rho=\kappa\) in Equation (A.985) and sum. On the right, \(\gamma^{\kappa}{}_{\alpha}\gamma^{\lambda}{}_{\kappa} =\gamma^{\lambda}{}_{\alpha}\), so, lowering the free index,
Writing \(\gamma^{\lambda\alpha}=g^{\lambda\alpha} -\epsilon n^{\lambda}n^{\alpha}\) and using \(g^{\lambda\alpha}R_{\alpha\tau\lambda\sigma}=R_{\tau\sigma}\) from Equation (13.307),
Step 2: contract again. Apply \(\gamma^{\nu\mu}\). The first bracket gives \(\gamma^{\tau\sigma}R_{\tau\sigma}=R-\epsilon R_{\tau\sigma} n^{\tau}n^{\sigma}\). The second gives \(\gamma^{\tau\sigma}n^{\alpha}n^{\lambda} R_{\alpha\tau\lambda\sigma}=R_{\alpha\lambda}n^{\alpha}n^{\lambda}\): the \(g^{\tau\sigma}\) part is again a Ricci contraction on the second and fourth slots, and the \(n^{\tau}n^{\sigma}\) part vanishes because \(R_{\alpha\tau\lambda\sigma}n^{\alpha}n^{\tau}=0\). The quadratic terms give \(\mathcal{K}_{\mu\nu}\mathcal{K}^{\mu\nu}-\mathcal{K}^{2}\). Hence
Step 3: eliminate the normal–normal Ricci term. Substitute Equation (A.992) into Equation (A.1000) and solve for \(R\):
which is Equation (A.996): the quadratic term appears twice with opposite weight and what survives is \(-\epsilon\left(\mathcal{K}\mathcal{K}-\mathcal{K}^{2}\right)\). It is worth noticing that the sign of the quadratic term in the final answer is opposite to the sign it carries in the Gauss equation, and that the flip is entirely the work of Lemma A.631; this is the step at which a derivation that quotes the Ricci equation instead of proving it can go wrong without leaving a trace.
Step 4: substitute. By Lemma A.623, \(\epsilon=+1\); by Lemma A.627, \(R[\gamma]=-{}^{(3)}\!R\), \(\mathcal{K}_{\mu\nu}\mathcal{K}^{\mu\nu}=K_{ij}K^{ij}\) and \(\mathcal{K}=K\). So
Multiplying by \(\sqrt{\abs{g}}=N\sqrt{h}\) and using the standard identity \(\sqrt{\abs{g}}\,\nabla_{\mu}V^{\mu} =\pp_{\mu}\left(\sqrt{\abs{g}}V^{\mu}\right)\) for the divergence of a vector field — which follows from \(\Gamma^{\mu}{}_{\mu\nu}=\pp_{\nu}\ln\sqrt{\abs{g}}\), itself a contraction of Equation (13.299) — gives Equation (A.997).
Step 5: the velocities. \({}^{(3)}\!R\) is built from \(h_{ij}\) and its spatial derivatives alone. \(K_{ij}\) is given by Equation (26.79), in which \(\pp_{0}\) acts on \(h_{ij}\) and on nothing else: \(N\) and \(N^{i}\) enter algebraically and through spatial derivatives \(D_{i}N_{j}\) only. The divergence term contains \(\pp_{0}\) of \(\sqrt{\abs{g}}\left(a^{0}-Kn^{0}\right)\), hence of \(N\), \(N^{i}\) and \(h_{ij}\); but it is a total derivative and is removed by the boundary term of The boundary term, so no time derivative of the lapse or the shift survives in the bulk Lagrangian.
∎Equation (A.997) carries an overall minus sign that Equation (26.80) does not display, and it is worth saying plainly where it comes from and what it costs, because this is exactly the bookkeeping the chapter's prose defers to this appendix.
The source is isolated in Lemma A.627 and is a single asymmetry: of the four quantities in the dictionary, three are built with an even number of inverse metrics and are blind to \(\gamma_{ij}=-h_{ij}\), while the curvature scalar carries exactly one and changes sign. So with the curvature conventions of Theorem 13.152 and Definition 13.153 — the conventions for which the round two-sphere has \(R=2/r^{2}\), checked in Remark A.629 — and the mostly-minus signature of Minkowski Space and Its Symmetries, the four-scalar that equals \(N\sqrt{h}\left({}^{(3)}\!R+K_{ij}K^{ij}-K^{2}\right)\) up to divergences is \(-\sqrt{\abs{g}}\,R\), and not \(+\sqrt{\abs{g}}\,R\).
Nothing physical turns on this, and three things should be said about it. First, the gravitational Lagrangian density in this signature is therefore \(-\sqrt{\abs{g}}\,R/2\kappa\); that is the sign carried, for precisely this reason, by every mostly-minus treatment of general relativity, and it is the combination whose ADM form is \(+N\sqrt{h}\left({}^{(3)}\!R+K_{ij}K^{ij}-K^{2}\right)/2\kappa\). Second, that is exactly the Lagrangian \(L\) used in the proof of Theorem 26.42 — read the first display of that proof — so the momentum Equation (26.82), the constraint densities Equations (26.84) and (26.85) and the degree-of-freedom count all stand unchanged. Third, and this is why Proposition 26.41 can be used as it is written, neither of the two facts the chapter draws from it is touched: \(\sqrt{\abs{g}} =N\sqrt{h}\) is Equation (A.969), and the absence of \(\pp_{0}N\) and \(\pp_{0}N^{i}\) is Step 5 above, and both are independent of the overall sign. A reader working in a mostly-plus convention sets \(\epsilon=-1\) and \(\gamma_{ij}=+h_{ij}\) in Equation (A.996) and recovers the familiar \(R={}^{(3)}\!R+K_{ij}K^{ij}-K^{2} -2\nabla_{\mu}\left(a^{\mu}-Kn^{\mu}\right)\) with no minus sign in front, which is the form printed in the standard references [Arnowitt:1962] [Misner:1973] [Wald:1984] and the reason the sign is so easy to import unexamined.
The boundary term
The divergence in Equation (A.997) is a genuine coordinate divergence, so over a region \(\mathcal{R}\) bounded by two leaves it contributes only a surface integral. Evaluate it. The acceleration is tangential, so \(n_{\mu}a^{\mu}=Na^{0}=0\) and \(a^{0}=0\): the acceleration has no flux through a leaf. And \(\sqrt{\abs{g}}\,n^{0}=N\sqrt{h}\cdot N^{-1}=\sqrt{h}\) by Equations (A.968) and (A.969). Hence the \(\mu=0\) component of the divergence's argument is
and, integrating Equation (A.997) over \(\mathcal{R}\),
the sign of the last term following the outward orientation of \(\pp\mathcal{R}\) once that orientation is fixed. Dividing by \(2\kappa\), the surface term is \(\kappa^{-1}\oint\sqrt{h}\,K\), which is the Gibbons–Hawking–York boundary action [Gibbons:1977] [York:1972].
Two statements are made about it here and only the first is proved. The first is Equation (A.1004) itself: the bulk ADM Lagrangian and the gravitational action differ by exactly that surface integral, which is why Proposition 26.41 may write “plus total derivatives” and Theorem 26.42 may drop them. The second is quoted: adding the Gibbons–Hawking–York term to the action makes the variational problem well posed with \(h_{ij}\) — and only \(h_{ij}\) — held fixed on \(\pp\mathcal{R}\), because the unwanted normal derivatives of the metric variation that \(\sqrt{\abs{g}}\,R\) produces on the boundary are precisely what it cancels. That is proved in [Gibbons:1977] and not here; the constrained-dynamics consequence of the same term, that the numerical value of the Hamiltonian of an isolated system is carried entirely by a surface integral, is Remark 26.45.
Two imports, and one debt of Differentiable Manifolds, Tensors, and Curvature.
Quoted. That the Gibbons–Hawking–York term makes the variational problem well posed, as stated in the paragraph above; it is used nowhere in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism except as the justification for discarding total derivatives, which Equation (A.1004) establishes independently.
Owed by Differentiable Manifolds, Tensors, and Curvature. A general-signature Gauss–Codazzi theory of hypersurfaces. That chapter carries the classical version Equation (13.92), but it is the theory of a surface in flat Euclidean \(\R^{3}\), built from a position vector \(\vect{x}\) and its derivatives, and it is not usable for a spacelike leaf of a Lorentzian spacetime: there is no position vector, the ambient is curved, and the normal's square is \(\epsilon\) rather than \(+1\). The chapter's later treatment of a quadric in a flat ambient of signature \((p,q)\) carries the right sign structure but is again an embedding into a flat space given by an explicit position vector. Lemmas A.628, A.630 and A.631 are therefore proved here, from Theorem 13.152 and the projector alone, and the general statement belongs in the manifolds chapter. This is recorded as a Part II debt, in the same terms as Remarks 26.4 and 26.18 record theirs.
Everything else — the inverse metric, the determinant, the extrinsic curvature in coordinates, the dictionary, the two contractions of the Gauss equation and the elimination of the normal–normal Ricci term — is carried out above from the Levi-Civita connection of Theorem 13.150 and the Riemann tensor of Theorem 13.152.
The Gauss–Codazzi Form of the Einstein–Hilbert Lagrangian discharges the derivation owed at Proposition 26.41 of Section 26.6.5 in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism. What the chapter takes from it is narrow and is now established: \(\sqrt{\abs{g}}=N\sqrt{h}\) (Equation (A.969)), and that the Einstein–Hilbert Lagrangian contains no time derivative of the lapse or the shift (Step 5 of Theorem A.632). Those two facts are what make \(\pi_{N}\) and \(\pi_{i}\) vanish identically in Equation (26.81), hence what make \(N\) and \(N^{i}\) Lagrange multipliers rather than dynamical fields, hence what make \(\Ham_{\perp}\) and \(\Ham_{i}\) secondary constraints by outcome 2 of Proposition 26.10 — so the whole of Theorem 26.42, and with it the count of two degrees of freedom per point, rests on this section. The overall sign of Equation (26.80) is settled in Remark A.633 and changes none of that. The brackets of the constraints so obtained are computed in The Hypersurface-Deformation Algebra.
The Hypersurface-Deformation Algebra
This appendix proves Proposition 26.43 of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism: the three brackets Equations (26.89), (26.90) and (26.91) of the smeared constraints of Theorem 26.42. With them the eight constraints of general relativity are shown to be first class in the sense of Definition 26.12, so that Equation (26.21) applies and the count of two propagating degrees of freedom per point of space is established; and the last of the three exhibits, on the right-hand side, the inverse spatial metric \(h^{ij}\) — a function on phase space where a Lie algebra would carry a constant.
Everything is computed with smeared constraints. That is not a matter of taste. The unsmeared densities have brackets proportional to \(\delta^{3}(\vect{x}-\vect{y})\) and to its first and second derivatives, and Equation (26.91) in unsmeared form is an identity between distributions in which the two sides must be compared after integrating twice by parts against test functions anyway. Smearing does that integration once and for all, at the start; every step below is then an ordinary integration by parts on smooth compactly supported data.
Two things are assumed rather than proved and are named here. The first is the constraint densities themselves, Equations (26.84) and (26.85), which Theorem 26.42 derives; nothing below depends on the overall sign discussed in Remark A.633, because the brackets are computed from the densities as the chapter displays them. The second is that all fields fall off fast enough for the surface terms of the integrations by parts to vanish — on a region with a boundary they do not, and what they contribute is exactly the ADM charges of Remark 26.45.
Conventions: weights, smearing functions, and the bracket
The canonical pair is \(\left(h_{ij},\pi^{ij}\right)\) of Equation (26.82), with the field bracket Equation (26.47) read as
functional derivatives with respect to \(h_{ij}\) and \(\pi^{ij}\) being taken with both regarded as symmetric. The weights are: \(h_{ij}\) is a tensor of weight \(0\); \(\pi^{ij}\) is a tensor density of weight \(1\), as Equation (26.82) shows through its factor \(\sqrt{h}\); \(\Ham_{\perp}\) of Equation (26.84) is a scalar density of weight \(1\), its kinetic part carrying \(\pi\pi/\sqrt{h}\) and its curvature part \(\sqrt{h}\,{}^{(3)}\!R\); and \(\Ham_{i}\) of Equation (26.85) is a covector density of weight \(1\). The smearing functions \(f\), \(g\) are scalars and \(\xi^{i}\), \(\zeta^{i}\) vectors, all of weight \(0\), all smooth, all of compact support, and — this is used constantly and is easy to forget — all independent of the canonical variables, so that they pass through every functional derivative untouched. Finally
both of which are ordinary numbers because a weight-one density integrates without a metric factor. Rests on Notation 26.28, Equation (26.82) and Theorem 26.42.
Getting a weight wrong is the commonest way this computation fails, so the one fact used about them is stated once and checked: for a weight-one vector density \(V^{i}\),
for \(V\) of compact support. This is immediate from the covariant derivative of a density, \(D_{i}V^{i}=\pp_{i}V^{i} +\Gamma^{i}{}_{ik}V^{k}-\Gamma^{k}{}_{ki}V^{i}\), the last term being the weight term; the two cancel. Every integration by parts below is an application of Equation (A.1007). For a weight-zero vector the corresponding statement carries a \(\sqrt{h}\), \(\sqrt{h}D_{i}V^{i}=\pp_{i}\left(\sqrt{h}V^{i}\right)\), and that form is used where the integrand carries an explicit \(\sqrt{h}\).
The sign convention for the commutator of vector fields is
which is what makes the right-hand side of Equation (26.89) carry a plus sign rather than a minus.
The momentum constraint generates the spatial drag
For every smearing vector \(\xi\),
and consequently
the second Lie derivative being that of a weight-one contravariant density, \(\mathcal{L}_{\xi}\pi^{ij}=\xi^{k}\pp_{k}\pi^{ij} -\pi^{kj}\pp_{k}\xi^{i}-\pi^{ik}\pp_{k}\xi^{j} +\pi^{ij}\pp_{k}\xi^{k}\). Rests on Equation (26.85), Equation (A.1005) and Notation A.636.
Derives Lemma A.637. Equation (A.1009). By Equation (26.85), \(H_{\parallel}[\xi]=-2\int\dd^{3}x\;\xi^{i}D_{j}\pi^{j}{}_{i} =-2\int\dd^{3}x\;\xi_{k}D_{j}\pi^{jk}\), the index having been moved through \(D\) because \(D h=0\). The quantity \(\xi_{k}\pi^{jk}\) is a weight-one vector density, so Equation (A.1007) gives \(\int D_{j}\left(\xi_{k}\pi^{jk}\right)=0\), whence
the symmetrization being free because \(\pi^{jk}\) is symmetric. That \(D_{j}\xi_{k}+D_{k}\xi_{j}=\mathcal{L}_{\xi}h_{jk}\) is Equation (13.275).
First of Equation (A.1010). Take \(F=h_{ij}(\vect{x})\) in Equation (A.1005): only the first term survives, and \(\pb{h_{ij}(\vect{x})}{G} =\delta G/\delta\pi^{ij}(\vect{x})\). Applied to Equation (A.1009), in which \(\pi\) appears linearly and undifferentiated, this is \(\mathcal{L}_{\xi}h_{ij}\).
Second. Likewise \(\pb{\pi^{ij}(\vect{x})}{G} =-\delta G/\delta h_{ij}(\vect{x})\). Write out Equation (A.1009) in partial derivatives, \(\mathcal{L}_{\xi}h_{ab}=\xi^{k}\pp_{k}h_{ab}+h_{kb}\pp_{a}\xi^{k} +h_{ak}\pp_{b}\xi^{k}\), and vary. The first term contributes, after one integration by parts, \(-\pp_{k}\left(\xi^{k}\pi^{ij}\right)\); the second contributes \(\pi^{kj}\pp_{k}\xi^{i}\) and the third \(\pi^{ik}\pp_{k}\xi^{j}\). Hence
and the minus of the bracket cancels it. Note where the weight shows itself: the term \(\pi^{ij}\pp_{k}\xi^{k}\), which distinguishes the Lie derivative of a density from that of a tensor, comes from the integration by parts in the first term and from nowhere else. Had \(\pi^{ij}\) been treated as a weight-zero tensor, the density term would be spurious and Equation (26.89) would fail.
∎The first two brackets
Both follow from Lemma A.637 and one observation, which is worth isolating because it does the work twice.
Let \(\Phi\) be a functional of the canonical variables and of one smearing field \(\sigma\) (a scalar \(f\) or a vector \(\xi\)),
with \(\mathcal{F}\) a scalar density of weight one built covariantly from its arguments. Then
Rests on Lemma A.637 and Equation (A.1005).
Derives Lemma A.638. By Equation (A.1005) and Equation (A.1010),
the change of \(\Phi\) when the canonical fields alone are dragged by \(\zeta\). But \(\Phi\) is invariant when everything is dragged: if all of \(\sigma\), \(h\), \(\pi\) are transported, \(\mathcal{F}\) changes by \(\mathcal{L}_{\zeta}\mathcal{F}=\pp_{k}\left(\zeta^{k}\mathcal{F}\right)\), because \(\mathcal{F}\) is a weight-one scalar density, and that integrates to zero. Hence dragging the fields alone gives minus the result of dragging \(\sigma\) alone, which is Equation (A.1014).
∎Equations (26.89) and (26.90) hold:
Rests on Lemma A.638, Equation (26.84) and Equation (A.1008).
Derives Proposition A.639. Both integrands satisfy the hypothesis of Lemma A.638: \(\xi^{i}\Ham_{i}\) and \(f\Ham_{\perp}\) are weight-one scalar densities built covariantly from the smearing field and the canonical pair, by Notation A.636.
For the first, apply Equation (A.1014) with \(\Phi=H_{\parallel}\), \(\sigma=\xi\): \(\pb{H_{\parallel}[\xi]}{H_{\parallel}[\zeta]} =-H_{\parallel}\!\left[\mathcal{L}_{\zeta}\xi\right] =-H_{\parallel}\!\left[\comm{\zeta}{\xi}\right] =H_{\parallel}\!\left[\comm{\xi}{\zeta}\right]\) by Equation (A.1008). Nothing about the detailed form of \(\Ham_{i}\) was used beyond Lemma A.637; the identity is \(\comm{\mathcal{L}_{\xi}}{\mathcal{L}_{\zeta}} =\mathcal{L}_{\comm{\xi}{\zeta}}\) in disguise.
For the second, apply it with \(\Phi=H_{\perp}\), \(\sigma=f\) a scalar, so \(\mathcal{L}_{\xi}f=\xi^{i}\pp_{i}f\): \(\pb{H_{\perp}[f]}{H_{\parallel}[\xi]} =-H_{\perp}\!\left[\xi^{i}\pp_{i}f\right]\), and the antisymmetry of the bracket gives the stated form. The content is that \(\Ham_{\perp}\) is a scalar density of weight one — that dragging \(f\Ham_{\perp}\) produces \(\mathcal{L}_{\xi}\left(f\Ham_{\perp}\right)\), whose total-divergence part integrates away and whose remainder is \(\left(\xi^{i}\pp_{i}f\right)\Ham_{\perp}\).
∎The variation of the intrinsic curvature
The third bracket needs the piece a reader cannot supply from the chapter: how \(\int\sqrt{h}\,{}^{(3)}\!R\) responds to a variation of \(h_{ij}\).
For any smooth \(f\) of compact support, independent of \(h_{ij}\),
with \({}^{(3)}\!G^{ij}={}^{(3)}\!R^{ij} -\tfrac{1}{2}h^{ij}\,{}^{(3)}\!R\) the three-dimensional Einstein tensor and \(D^{2}=h^{ij}D_{i}D_{j}\). Rests on Definition 13.153, Equation (13.299) and Theorem 13.152.
Derives Lemma A.640. Three pieces. First, \(\delta\sqrt{h} =\tfrac{1}{2}\sqrt{h}\,h^{ij}\delta h_{ij}\), from \(\delta\ln\det h=\tr\left(h^{-1}\delta h\right)\). Second, \(\delta h^{ij}=-h^{ik}h^{jl}\delta h_{kl}\), from varying \(h^{ik}h_{kj}=\delta^{i}_{j}\); so \(\delta\left({}^{(3)}\!R\right) ={}^{(3)}\!R_{ij}\delta h^{ij}+h^{ij}\delta\,{}^{(3)}\!R_{ij} =-{}^{(3)}\!R^{ij}\delta h_{ij}+h^{ij}\delta\,{}^{(3)}\!R_{ij}\).
Third, the term \(h^{ij}\delta\,{}^{(3)}\!R_{ij}\). The variation of a Christoffel symbol is a tensor, and from Equation (13.299) with vanishing torsion,
which one checks by evaluating both sides in normal coordinates at a point, where \(\Gamma=0\) and \(D=\pp\). Then Equation (13.306) gives \(\delta\,{}^{(3)}\!R_{ij} =D_{k}\delta\Gamma^{k}{}_{ij}-D_{i}\delta\Gamma^{k}{}_{kj}\), so \(h^{ij}\delta\,{}^{(3)}\!R_{ij}=D_{k}v^{k}\) with \(v^{k}=h^{ij}\delta\Gamma^{k}{}_{ij}-h^{kj}\delta\Gamma^{i}{}_{ij}\). Evaluate the two pieces from Equation (A.1018). In the second, the first and third terms of the bracket cancel on contraction, leaving \(\delta\Gamma^{i}{}_{ij} =\tfrac{1}{2}D_{j}\left(h^{il}\delta h_{il}\right)\). In the first, the two symmetric terms combine, leaving \(h^{ij}\delta\Gamma^{k}{}_{ij}=h^{kl}D^{j}\delta h_{lj} -\tfrac{1}{2}D^{k}\left(h^{ij}\delta h_{ij}\right)\). Hence
Collecting,
the first term assembling \(\tfrac{1}{2}h^{ij}\,{}^{(3)}\!R -{}^{(3)}\!R^{ij}\). Multiply by \(f\), integrate, and move the two derivatives of the second term onto \(f\) using \(\sqrt{h}\,D_{k}V^{k}=\pp_{k}\left(\sqrt{h}V^{k}\right)\) twice. The result is Equation (A.1017); note that \(D^{i}D^{j}f\) is automatically symmetric, \(f\) being a scalar.
∎The third bracket
Equation (26.91) holds:
Rests on Lemma A.640, Equation (26.84) and Equation (26.85).
Derives Theorem A.641. Step 1: the two functional derivatives. From Equation (26.84), in which \(\pi\) occurs only algebraically,
since \(\delta\left(\pi_{kl}\pi^{kl}\right)/\delta\pi^{ij}=2\pi_{ij}\) and \(\delta\left(\pi^{2}\right)/\delta\pi^{ij}=2\pi h_{ij}\). From Lemma A.640 together with the variation of the kinetic part — which involves \(h_{ij}\) algebraically only, through \(\pi_{ij}\), \(\pi\) and \(\sqrt{h}\), and which therefore produces no derivative of \(f\) —
where \(A^{ij}\) is a local expression in \(h\) and \(\pi\) whose exact form is never needed: it collects the kinetic variation and the term \(+\sqrt{h}\,{}^{(3)}\!G^{ij}/2\kappa\), and the only property used is that it multiplies \(f\) undifferentiated.
Step 2: what cancels, and why. Insert Equations (A.1022) and (A.1023) into Equation (A.1005):
because the terms \(fgA^{ij}B_{ij}\) are symmetric under \(f\leftrightarrow g\) and cancel between the two halves of the bracket. This is the load-bearing observation of the whole computation and it deserves to be stated in words: everything ultralocal in the smearing functions drops out. The Einstein-tensor part of Equation (A.1017) goes with it, and so does the entire kinetic variation; what survives is only the group \(\Delta^{ij}\), which is precisely the part carrying two spatial derivatives. That is why the answer will carry a metric where a structure constant would sit: two derivatives must have their indices raised, and \(h^{ij}\) is what raises them.
Step 3: contract. With \(\sqrt{h}B_{ij}/2\kappa=2\left(\pi_{ij} -\tfrac{1}{2}\pi h_{ij}\right)\),
the three trace terms cancelling exactly, the last carrying \(h_{ij}h^{ij}=3\). Hence
Step 4: one integration by parts. \(\pi^{ij}\) is a weight-one density, so Equation (A.1007) applies directly to \(\pi^{ij}fD_{j}g\):
and likewise with \(f\) and \(g\) exchanged. In the difference, the term \(\pi^{ij}\left(D_{i}f\,D_{j}g-D_{i}g\,D_{j}f\right)\) is antisymmetric in \(i,j\) contracted with the symmetric \(\pi^{ij}\), and vanishes. What remains is
Step 5: recognize the momentum constraint. Lower the index on \(\pi\) and raise it on the bracket, which is legitimate because \(Dh=0\): \(\left(D_{j}\pi^{ji}\right)\left(fD_{i}g-gD_{i}f\right) =\left(D_{j}\pi^{j}{}_{k}\right)h^{ki} \left(f\pp_{i}g-g\pp_{i}f\right)\), the covariant derivatives of the scalars \(f\), \(g\) being ordinary ones. By Equation (26.85), \(-2D_{j}\pi^{j}{}_{k}=\Ham_{k}\), so
which is Equation (A.1021).
∎Every right-hand side in Equations (A.1016) and (A.1021) is a smeared constraint, and the primary constraints Equation (26.81) have vanishing bracket with everything, since no \(\Ham\) contains \(N\) or \(N^{i}\). So all eight constraints of Theorem 26.42 are first class in the sense of Definition 26.12, \(F=8\) and \(S=0\), and Equation (26.21) gives \(2\times10-2\times8=4\) per point of space. Rests on Theorem A.641, Proposition A.639 and Theorem 26.17.
Derives Corollary A.642. By Proposition 26.13 it suffices that the brackets of the generators close on the constraint set, which is what the three displayed identities say — each right-hand side is \(H_{\perp}\) or \(H_{\parallel}\) of some smearing field, hence weakly zero. The smearing fields on the right of Equation (A.1021) depend on \(h_{ij}\), but that is irrelevant to the weak vanishing: a phase-space-dependent coefficient times a constraint still vanishes on \(\Sigma\). The count is Equation (26.21) with \(2n=20\) per point.
∎The geometric reading
The three brackets are the composition law for deformations of a spacelike surface embedded in a Lorentzian spacetime, and read that way they explain themselves.
\(H_{\parallel}[\xi]\) deforms the surface within itself: it slides the points along \(\xi\) and leaves the surface, as a subset of spacetime, where it was. Two such slides compose as diffeomorphisms of a three-manifold compose, which is Equation (A.1016) left, and their commutator is again a slide. That part of the algebra is a Lie algebra — the Lie algebra of vector fields on the leaf — and it is the only part that is.
\(H_{\perp}[f]\) pushes the surface off itself, by a proper time \(f/c\) per point along the normal. The mixed bracket says only that \(f\) is a scalar: sliding and then pushing differs from pushing and then sliding by pushing with the dragged function \(\xi^{i}\pp_{i}f\).
The third bracket is the one with content. Push by \(f\), then by \(g\); push by \(g\), then by \(f\). Each pair of normal pushes lands on the same surface — to first order the two orders of deformation reach the same place — but not on the same points of it: the two routes differ by a slide. The slide is easy to see. The normal at a point of the first intermediate surface is not the normal at the corresponding point of the second, because the two intermediate surfaces are tilted relative to one another by an amount proportional to the gradient of the pushing function; and tilting a normal is exactly what turns a normal deformation into a tangential one. The resulting tangent vector is \(f\,\nabla g-g\,\nabla f\) — the gradient of one function weighted by the other, antisymmetrized — and a gradient is a covector, so to be a deformation vector it must have its index raised. There is only one object available to raise it: the metric of the surface itself. Hence \(h^{ij}\) in Equation (A.1021).
And hence the sentence that closes Proposition 26.43. The coefficient on the right of the third bracket is a function on phase space, not a number, so Equation (26.16) for general relativity has structure functions and the first-class algebra is not a Lie algebra. This is not a technicality: it is the reason the constraint algebra of gravity cannot be treated as the algebra of a symmetry group with a fixed structure, why the BRST charge of Existence and Uniqueness of a Nilpotent BRST Charge needs terms beyond the two displayed in Equation (26.97), and why the comparison with Yang–Mills of Section 26.6.4, where the same brackets do close on constants, is the sharpest single contrast the chapter draws.
Nothing outside the treatise. The constraint densities Equations (26.84) and (26.85) are taken from Theorem 26.42, which proves them; the Levi-Civita apparatus is Theorems 13.150 and 13.152, and the component form of the Lie derivative is Proposition 13.127. Two assumptions are made and are not hidden. The smearing fields are independent of the canonical variables — for field-dependent smearings the brackets acquire extra terms and Equation (A.1016) is false as it stands. And the fields fall off fast enough that every integration by parts loses its surface term; on a region with a boundary they do not, the constraints are not even functionally differentiable without supplementary surface integrals, and those integrals are the ADM energy and momentum of Remark 26.45.
The Hypersurface-Deformation Algebra discharges the derivation owed at Proposition 26.43 of Section 26.6.5 in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism. It completes the last step of Theorem 26.42, which asserts that the eight constraints are first class and points forward to Equation (26.89) for the proof; with Corollary A.642 the count of two degrees of freedom per point of space — the same count Proposition 26.32 gives for light, and the two polarizations detected in Experiment: Gravitational Waves — rests on nothing unproved. The Lagrangian from which those constraints were obtained is established in The Gauss–Codazzi Form of the Einstein–Hilbert Lagrangian.
Existence and Uniqueness of a Nilpotent BRST Charge
This appendix proves Theorem 26.51 of Section 26.7.3 in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism: for every regular, irreducible set of first-class constraints the series Equation (26.97) can be continued to a Grassmann-odd function \(\Omega\) of ghost number \(+1\) satisfying Equation (26.98) exactly; two such functions with the same leading term differ by a canonical transformation of the extended phase space; and when the structure functions of Equation (26.16) are constants the series stops at the two terms Equation (26.97) displays.
This is the statement that makes BRST a method rather than a device. Without it the definition Definition 26.50 would be a recipe that happens to work for Yang–Mills and is silent about general relativity, whose structure functions are genuinely functions (Proposition 26.43, proved in The Hypersurface-Deformation Algebra); with it, the cohomological characterization Equation (26.99) of the physical state space is available for every first-class system.
Scope. BRST is in this treatise because it is the machinery behind the perturbative calculations of Quantum Chromodynamics and the path-integral development of Path-Integral Quantization, whose predictions are measured; it is not here as a formal programme. Accordingly this section stops where the chapter stops — at the Hamiltonian BRST charge of a finite-dimensional constrained system — and does not go on to the antifield or Batalin–Vilkovisky formulation, which is a different construction with different inputs and is used nowhere in this book.
Attribution. The existence of a nilpotent charge for a general first-class system is due to Fradkin and Vilkovisky and, in the form with structure functions, to Batalin, Fradkin and Vilkovisky; the homological proof organized around the Koszul–Tate differential, which is the one given below, is due to Henneaux and Teitelboim. None of these works has an entry in this treatise's bibliography, so — exactly as Remark 26.2 records for the rest of the chapter — the attribution is made in words and nothing below rests on it. The original BRST papers [Becchi:1976] [Tyutin:1975] are cited in the chapter.
The extended phase space and its two gradings
Let \(z^{c}\), \(c=1,\ldots,2n\), be the original phase-space variables and \(\gamma_{A}\), \(A=1,\ldots,F\), the first-class constraints, with Equation (26.16), \(\pb{\gamma_{A}}{\gamma_{B}}=f_{AB}{}^{C}\gamma_{C}\). Adjoin the ghost pairs \(\left(\eta^{A},\mathcal{P}_{A}\right)\) of Definition 26.50. Functions of \(\left(z,\eta,\mathcal{P}\right)\) carry a Grassmann parity \(\varepsilon\in\set{0,1}\) and the graded bracket obeys
together with the graded Jacobi identity. It splits as
where \(\pb{\cdot}{\cdot}_{M}\) differentiates only the \(z^{c}\) — so that the ghosts pass through it as constants — and \(\pb{\cdot}{\cdot}_{\text{gh}}\) only the ghost pair, with \(\pb{\eta^{A}}{\mathcal{P}_{B}}=-\delta^{A}_{B}\).
Three gradings are used. The pure ghost number counts the \(\eta\)'s; the antighost number counts the \(\mathcal{P}\)'s,
and the ghost number of Definition 26.50 is their difference, \(\operatorname{gh}=\operatorname{pgh}-\operatorname{antigh}\), so that \(\operatorname{gh}\eta^{A}=+1\) and \(\operatorname{gh}\mathcal{P}_{A}=-1\). Parity is pgh \(+\) antigh modulo two. Rests on Definition 26.50, Equation (26.16) and Equation (26.97).
Two grading facts are used constantly and are recorded once. \(\pb{\cdot}{\cdot}_{M}\) preserves the antighost number, while \(\pb{\cdot}{\cdot}_{\text{gh}}\) removes one \(\eta\) and one \(\mathcal{P}\) and so lowers it by one; both add ghost numbers. And an \(\Omega\) that is odd with \(\operatorname{gh}\Omega=+1\) decomposes as
in which \(\Omega_{(k)}\) has pure ghost number \(k+1\): it is a polynomial carrying exactly \(k+1\) ghosts and \(k\) ghost momenta, hence \(2k+1\) Grassmann-odd factors, hence odd, as it must be. The first two terms are the ones Equation (26.97) displays,
The Koszul–Tate differential
\(\delta\) is the odd derivation of the algebra of functions on the extended phase space fixed by
It lowers the antighost number by one and raises the ghost number by one. Rests on Notation A.646 and Equation (26.16).
For every \(X\), \(\delta X=\pb{X}{\Omega_{(0)}}_{\text{gh}}\), and \(\delta^{2}=0\). Rests on Definition A.647, Equation (A.1030) and Equation (A.1034).
Derives Lemma A.648. \(\Omega_{(0)}\) contains no \(\mathcal{P}\), so \(\pb{\cdot}{\Omega_{(0)}}_{\text{gh}}\) annihilates \(z^{c}\) and \(\eta^{A}\), as \(\delta\) does. On \(\mathcal{P}_{A}\), the graded Leibniz rule Equation (A.1030) gives \(\pb{\mathcal{P}_{A}}{\eta^{B}\gamma_{B}}_{\text{gh}} =\pb{\mathcal{P}_{A}}{\eta^{B}}\gamma_{B}\), and graded antisymmetry between two odd objects gives \(\pb{\mathcal{P}_{A}}{\eta^{B}}=\pb{\eta^{B}}{\mathcal{P}_{A}} =-\delta^{B}_{A}\), so the value is \(-\gamma_{A}=\delta\mathcal{P}_{A}\). Both maps are odd derivations agreeing on generators, hence equal.
For nilpotency, \(\delta^{2}\) is the square of an odd derivation and is therefore an even derivation; it vanishes on \(z\) and \(\eta\) trivially and on \(\mathcal{P}_{A}\) because \(\delta^{2}\mathcal{P}_{A}=-\delta\gamma_{A}=0\), the constraints being functions of \(z\) alone. A derivation vanishing on generators is zero.
∎Remark 26.7 assumes two things about the constraint set, and Theorem 26.51 adds a third word, irreducible. Spelled out, what is assumed is: (i) the differentials \(\dd\gamma_{A}\) are linearly independent at every point of the surface \(\Sigma\) they define; (ii) every smooth function vanishing on \(\Sigma\) is \(c^{A}\gamma_{A}\) with smooth \(c^{A}\); and (iii) every \(\lambda^{A}\) with \(\lambda^{A}\gamma_{A}=0\) identically is of the form \(\lambda^{A}=\mu^{AB}\gamma_{B}\) with \(\mu^{AB}=-\mu^{BA}\) — there are no relations among the constraints beyond the trivial ones. Item (iii) is what irreducible means; a reducible set, such as the one obtained by writing a two-form constraint in components with a Bianchi identity among them, fails it, and the construction below then needs ghosts for ghosts and is not carried out in this treatise.
Lemma A.650 shows that (i) already implies (ii) and (iii) locally, so the three are not independent assumptions; the content is entirely in (i), which is the same regularity that Proposition 26.5 and Theorem 26.17 rest on.
About each point of \(\Sigma\) there are local coordinates \(\left(y^{A},w^{\alpha}\right)\) on phase space in which \(\gamma_{A}=y^{A}\). In such coordinates, (ii) and (iii) of Remark A.649 hold. Rests on Remark 26.4, Theorem A.286 and Remark 26.7.
Derives Lemma A.650. The map \(z\longmapsto\left(\gamma_{1}(z),\ldots,\gamma_{F}(z)\right)\) has surjective differential at each point of \(\Sigma\) by (i), hence constant maximal rank on a neighbourhood; the constant-rank theorem stated in Remark 26.4, itself a consequence of Theorem A.286, supplies coordinates in which it is the projection onto the first \(F\) of them. Those coordinates are the \(\left(y^{A},w^{\alpha}\right)\).
For (ii), let \(F(y,w)\) vanish at \(y=0\). By the fundamental theorem of calculus applied along the segment \(t\longmapsto\left(ty,w\right)\),
and \(\lambda_{A}\) is smooth by differentiation under the integral sign. That is (ii). Item (iii) is the case \(k=1\) of Lemma A.651 below, whose proof uses only Equation (A.1036) and the coordinates just constructed, so nothing is circular.
∎On the coordinate patch of Lemma A.650: \(H_{k}(\delta)=0\) for every \(k\ge1\) — every \(\delta\)-closed function of antighost number \(k\ge1\) is \(\delta\) of a function of antighost number \(k+1\) — and \(H_{0}(\delta)\) is the algebra of smooth functions on \(\Sigma\) (tensored with the ghosts). Rests on Lemma A.650, Definition A.647 and Remark A.649.
Derives Lemma A.651. Work in the coordinates of Lemma A.650, so that \(\delta\mathcal{P}_{A}=-y^{A}\). The \(\eta\)'s play no part: \(\delta\) annihilates them and they simply multiply everything, so the complex is a free module over the Grassmann algebra they generate and it suffices to treat \(\eta\)-independent coefficients.
A homotopy. Let \(\sigma\) be the odd derivation with
so that on functions \(\sigma f=\mathcal{P}_{A}\,\pp f/\pp y^{A}\). The anticommutator of two odd derivations is an even derivation, so
is a derivation, and it is determined by its values on generators: \(\mathcal{N}y^{A}=-\delta\mathcal{P}_{A}=y^{A}\), \(\mathcal{N}\mathcal{P}_{A}=-\sigma\left(-y^{A}\right) =\mathcal{P}_{A}\), and \(\mathcal{N}w^{\alpha} =\mathcal{N}\eta^{A}=0\). So \(\mathcal{N}\) is the Euler operator counting the joint degree in \(y\) and \(\mathcal{P}\). Because \(\delta^{2}=0\), \(\mathcal{N}\) commutes with \(\delta\).
Inverting \(\mathcal{N}\) in positive antighost degree. A general element of antighost number \(k\) is \(a=\frac{1}{k!}a^{A_{1}\ldots A_{k}}(y,w)\, \mathcal{P}_{A_{1}}\cdots\mathcal{P}_{A_{k}}\), on which \(\mathcal{N}=k+\mathcal{N}_{y}\) with \(\mathcal{N}_{y}=y^{A}\pp/\pp y^{A}\). For \(k\ge1\) define
acting on the coefficient functions and leaving the \(\mathcal{P}\)'s alone; it is smooth by differentiation under the integral sign. It is the inverse: from \(\dv{}{t}\left[t^{k}a(ty,w)\right] =t^{k-1}\left[k\,a+\mathcal{N}_{y}a\right](ty,w)\), integrating from \(0\) to \(1\) — where the boundary term at \(t=0\) vanishes because \(k\ge1\) — gives \(a=\mathcal{N}^{-1}\left(k+\mathcal{N}_{y}\right)a =\mathcal{N}^{-1}\mathcal{N}a\). Since \(\mathcal{N}\) commutes with \(\delta\) and is invertible in every antighost degree \(\ge1\), so does \(\mathcal{N}^{-1}\).
Acyclicity. Let \(\delta a=0\) with \(\operatorname{antigh}a=k\ge1\). Then by Equation (A.1038), \(\mathcal{N}a=-\delta\sigma a-\sigma\delta a=-\delta\sigma a\), and applying \(\mathcal{N}^{-1}\), which commutes with \(\delta\),
so \(a\) is \(\delta\)-exact, with a primitive of antighost number \(k+1\).
Degree zero. At antighost number \(0\) every element is closed, and the image of \(\delta\) from degree one is \(\set{\lambda^{A}\gamma_{A}}\), which by Equation (A.1036) is exactly the ideal of functions vanishing on \(\Sigma\). Hence \(H_{0}(\delta)=C^{\infty}(\Sigma)\).
This is the mathematical heart of the section, and it is worth naming what was consumed: the coordinates supplied by regularity, and nothing else. Irreducibility — item (iii) of Remark A.649 — is the case \(k=1\) of what has just been proved, since \(\lambda^{A}\mathcal{P}_{A}\) is \(\delta\)-closed precisely when \(\lambda^{A}\gamma_{A}=0\), and a primitive of antighost number \(2\) is an antisymmetric \(\mu^{AB}\) with \(\lambda^{A}=\mu^{AB}\gamma_{B}\).
∎The recursion
With \(\Omega\) expanded as in Equation (A.1033), the antighost-\(k\) component of \(\pb{\Omega}{\Omega}\) is
where
depends on \(\Omega_{(0)},\ldots,\Omega_{(k)}\) only. So Equation (26.98) is the sequence of equations \(\delta\Omega_{(k+1)}=D_{(k)}\), \(k\ge0\). Explicitly
and \(\Omega_{(1)}\) of Equation (A.1034) solves \(\delta\Omega_{(1)}=D_{(0)}\). Rests on Lemma A.648, Notation A.646 and Equation (26.98).
Derives Lemma A.652. Expand \(\pb{\Omega}{\Omega}\) bilinearly and sort by antighost number, using that \(\pb{\cdot}{\cdot}_{M}\) preserves it and \(\pb{\cdot}{\cdot}_{\text{gh}}\) lowers it by one. The terms of the ghost sum with \(i=0\) or \(j=0\) are \(\pb{\Omega_{(0)}}{\Omega_{(k+1)}}_{\text{gh}} +\pb{\Omega_{(k+1)}}{\Omega_{(0)}}_{\text{gh}}\), and the bracket of two odd functions is graded-symmetric by Equation (A.1030), so these are equal and their sum is \(2\pb{\Omega_{(k+1)}}{\Omega_{(0)}}_{\text{gh}} =2\delta\Omega_{(k+1)}\) by Lemma A.648. Everything else is Equation (A.1042). That \(D_{(k)}\) involves no \(\Omega_{(j)}\) with \(j>k\) is immediate from the ranges of the two sums.
For Equation (A.1043): since \(\pb{\cdot}{\cdot}_{M}\) differentiates only the \(z^{c}\), the ghosts pass through it, and moving one odd \(\eta\) past the even object \(\pp_{c}\gamma_{A}\) costs no sign, so
by Equation (26.16). Halving and negating gives Equation (A.1043). Equation (A.1044) is the case \(k=1\) of Equation (A.1042), the ghost sum having only the term \(i=j=1\) and the \(M\) sum only the two equal terms \(\pb{\Omega_{(0)}}{\Omega_{(1)}}_{M}\).
Finally, \(\delta\) passes through the even coefficient \(-\tfrac{1}{2}\eta^{B}\eta^{A}f_{AB}{}^{C}\) without a sign and \(\delta\mathcal{P}_{C}=-\gamma_{C}\), so
the middle step by anticommutativity of the ghosts. So the term Equation (26.97) displays is not a guess: it is the unique solution of the first equation of the recursion, up to the \(\delta\)-closed ambiguity that Uniqueness disposes of.
∎Notice what the first step already required. \(D_{(0)}\) has antighost number \(0\), where \(\delta\) is not acyclic, so \(\delta\Omega_{(1)}=D_{(0)}\) is solvable only if \(D_{(0)}\) lies in the image of \(\delta\) — that is, by Lemma A.651, only if it vanishes on \(\Sigma\). Equation (A.1043) shows it is proportional to the constraints, and that is precisely the statement that the constraints are first class. The whole construction is powered by Equation (26.16), and it fails at the first step for a set that is not first class.
Existence
Let \(\gamma_{A}\) be regular and irreducible in the sense of Remark A.649. Then on a neighbourhood of \(\Sigma\) there is an odd \(\Omega\) of ghost number \(+1\), of the form Equation (A.1033) with the first two terms Equation (A.1034), satisfying \(\pb{\Omega}{\Omega}=0\). Rests on Lemma A.651, Lemma A.652 and Equation (26.16).
Derives Theorem A.653. Induction on the antighost number. By Equation (A.1046), \(\Omega_{(0)}\) and \(\Omega_{(1)}\) are constructed and \(\pb{\Omega}{\Omega}\) has no antighost-\(0\) component. Suppose \(\Omega_{(0)},\ldots,\Omega_{(m)}\) have been found, \(m\ge1\), such that \(R:=\pb{\Omega^{[m]}}{\Omega^{[m]}}\), with \(\Omega^{[m]}:=\sum_{j\le m}\Omega_{(j)}\), has vanishing components of antighost number \(\le m-1\).
The obstruction is \(\delta\)-closed. For an odd \(\Omega\) the graded Jacobi identity applied to the triple \(\left(\Omega^{[m]},\Omega^{[m]},\Omega^{[m]}\right)\) collapses to the identity
which holds whatever \(\Omega^{[m]}\) is, nilpotent or not: the three cyclic terms of the graded Jacobi identity are equal, and their common sign is such that three times one of them must vanish. Take the antighost-\(\left(m-1\right)\) component of Equation (A.1047). The \(M\)-part contributes \(\sum_{i+j=m-1}\pb{\Omega_{(i)}}{R_{(j)}}_{M}\), in which every \(j\) is at most \(m-1\) and every \(R_{(j)}\) therefore vanishes by hypothesis. The ghost part contributes \(\sum_{i+j=m}\pb{\Omega_{(i)}}{R_{(j)}} _{\text{gh}}\), in which the only surviving term is \(i=0\), \(j=m\). Hence \(\pb{\Omega_{(0)}}{R_{(m)}}_{\text{gh}}=0\), and since \(R\) is even and \(\Omega_{(0)}\) odd this is \(-\delta R_{(m)}\). So \(\delta R_{(m)}=0\).
Solve. \(R_{(m)}\) has antighost number \(m\ge1\), so by Lemma A.651 there is a \(c\) with \(\delta c=R_{(m)}\) and \(\operatorname{antigh}c=m+1\). Counting gradings: \(R\) is even with ghost number \(2\), so \(R_{(m)}\) has pure ghost number \(m+2\); \(c\) therefore has pure ghost number \(m+2\) and antighost number \(m+1\), hence ghost number \(+1\) and odd parity — exactly the type of an \(\Omega_{(m+1)}\). Put \(\Omega_{(m+1)}:=-\tfrac{1}{2}c\). Adding it to \(\Omega^{[m]}\) changes nothing of antighost number \(<m\), because the new term enters \(\pb{\Omega}{\Omega}\) only at antighost number \(\ge m\), and it changes the antighost-\(m\) component by \(2\delta\Omega_{(m+1)}=-\delta c=-R_{(m)}\), cancelling it. This is the statement \(\delta\Omega_{(m+1)}=D_{(m)}\) of Lemma A.652.
The induction produces \(\Omega_{(k)}\) for every \(k\). It terminates for a reason of degree: \(\Omega_{(k)}\) carries \(k\) factors \(\mathcal{P}_{A}\), which anticommute, so \(\Omega_{(k)}=0\) identically for \(k>F\). The series Equation (A.1033) is therefore a finite sum, and \(\Omega\) is a genuine function on the extended phase space, not a formal series.
∎Termination at the displayed terms
If the \(f_{AB}{}^{C}\) of Equation (26.16) are constants, then \(D_{(1)}=0\) and \(\Omega=\Omega_{(0)}+\Omega_{(1)}\) is exactly nilpotent: the series Equation (26.97) stops at the two terms displayed. Rests on Lemma A.652, Equation (26.16) and Equation (22.46).
Derives Proposition A.654. Both terms of Equation (A.1044) vanish.
The second does so for a trivial reason: with constant \(f\), the function \(\Omega_{(1)}\) contains no \(z^{c}\) at all, and \(\pb{\cdot}{\cdot}_{M}\) differentiates only the \(z^{c}\), so \(\pb{\Omega_{(0)}}{\Omega_{(1)}}_{M}=0\). This is exactly the place where structure functions would obstruct: for non-constant \(f\) this bracket is a nonzero function of antighost number \(1\), and it is what forces \(\Omega_{(2)}\) to exist.
The first is the Jacobi identity. The ghost bracket pairs the \(\pp/\pp\eta\) of one \(\Omega_{(1)}\) with the \(\pp/\pp\mathcal{P}\) of the other, producing a term with three \(\eta\)'s, one \(\mathcal{P}\) and two factors \(f\); because the three ghosts anticommute, the coefficient is the total antisymmetrization \(f_{[AB}{}^{D}f_{C]D}{}^{E}\). That vanishes: applying the Poisson Jacobi identity Equation (22.46) to the triple \(\left(\gamma_{A},\gamma_{B},\gamma_{C}\right)\) and using Equation (26.16) twice with constant coefficients gives \(f_{[AB}{}^{D}f_{C]D}{}^{E}\gamma_{E}=0\), and the \(\gamma_{E}\) are independent by irreducibility, so the coefficient itself vanishes.
With \(D_{(1)}=0\) one may take \(\Omega_{(2)}=0\), and then every later \(D_{(k)}\) is built from \(\Omega_{(0)}\), \(\Omega_{(1)}\) and vanishing terms and is itself zero by the same two arguments, so \(\Omega_{(k)}=0\) for all \(k\ge2\) solves the recursion. That is the case of Section 26.6.4; general relativity, whose structure functions are genuinely functions by Proposition 26.43, is not it, and the omitted terms of Equation (26.97) are exactly what Theorem A.653 supplies there.
∎Uniqueness
Let \(\Omega\) and \(\Omega'\) both satisfy the hypotheses and conclusion of Theorem A.653 with the same \(\Omega_{(0)}=\Omega'_{(0)}=\eta^{A}\gamma_{A}\). Then there is a canonical transformation of the extended phase space, generated by an even function of ghost number \(0\), carrying \(\Omega\) to \(\Omega'\). Rests on Theorem A.653, Lemma A.651 and Lemma A.652.
Derives Proposition A.655. Induction on the lowest order at which the two differ. Suppose \(\Omega_{(j)}=\Omega'_{(j)}\) for \(j<k\) and put \(\Delta_{(k)}:=\Omega'_{(k)}-\Omega_{(k)}\), with \(k\ge1\). By Lemma A.652 both satisfy \(\delta\Omega_{(k)}=D_{(k-1)}\) with the same right-hand side, since \(D_{(k-1)}\) is built from the orders below \(k\), which agree. Hence \(\delta\Delta_{(k)}=0\).
By Lemma A.651 there is a \(K\) with \(\delta K=\Delta_{(k)}\) and \(\operatorname{antigh}K=k+1\). Its type is forced: \(\Delta_{(k)}\) is odd with ghost number \(+1\) and pure ghost number \(k+1\), so \(K\) has pure ghost number \(k+1\) and antighost number \(k+1\), hence ghost number \(0\) and even parity — a legitimate generator of a canonical transformation preserving both gradings.
Let \(T_{K}\) be the canonical transformation generated by \(K\),
which is a finite sum here because \(K\) has positive antighost number and each bracket lowers it by at most one while raising the number of \(\eta\)'s, so the series terminates. Since \(K\) has antighost number \(k+1\) and the ghost bracket lowers antighost number by one, \(\pb{\Omega}{K}\) has antighost number \(\ge k\), and its antighost-\(k\) part is \(\pb{\Omega_{(0)}}{K}_{\text{gh}}=-\delta K=-\Delta_{(k)}\). So \(T_{K}\Omega'\) agrees with \(\Omega\) through antighost number \(k\) and, being the image of a nilpotent charge under a canonical transformation, is itself nilpotent — a canonical transformation preserves the graded bracket, so it carries \(\pb{\Omega'}{\Omega'}=0\) to \(\pb{T_{K}\Omega'}{T_{K}\Omega'}=0\).
Repeat at the next order at which the two still differ. The generators so produced have strictly increasing antighost number, and all vanish beyond \(F\) by the degree argument of Theorem A.653, so their composition is a single canonical transformation carrying \(\Omega'\) to \(\Omega\).
∎Two things, and one limitation.
Quoted. The constant-rank theorem, in the form Remark 26.4 states it, is used once, in Lemma A.650, to flatten the constraints. That theorem is owed by Real Analysis and Differentiable Manifolds, Tensors, and Curvature and is a standing debt of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism, not a new one; the analytic input from which it is proved, Theorem A.286, is in this appendix.
Not quoted, though it usually is. The acyclicity of the Koszul complex in positive degree (Lemma A.651) is normally imported from homological algebra, where it is the statement that the Koszul complex of a regular sequence is a resolution. This treatise carries no homological algebra, and opening a chapter of it to prove one lemma would be out of proportion; the case at hand is therefore proved here outright, by an explicit contracting homotopy Equations (A.1037) and (A.1039), and the general theorem is recorded as a debt of Part II rather than used.
The limitation. Lemma A.650 is local: the coordinates in which \(\gamma_{A}=y^{A}\) exist on a neighbourhood of a point of \(\Sigma\), so Lemma A.651 and with it Theorem A.653 are local statements. Passing to a global \(\Omega\) requires patching the local solutions with a partition of unity, which works whenever the \(\gamma_{A}\) are globally defined and the surface admits a tubular neighbourhood — the case in every application in this book, where the constraints are given by global formulas — but it is a further step and it is not carried out here. This is the same locality that Remark 26.27 records for gauge fixing, and for the same reason: what is easy near a point of the constraint surface can be obstructed in the large.
Existence and Uniqueness of a Nilpotent BRST Charge discharges the derivation owed at Theorem 26.51 of Section 26.7.3 in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism. Three consequences are worth carrying back. The construction shows that the two terms Equation (26.97) displays are not an ansatz but the first two solutions of a recursion, the second of them forced; Proposition A.654 shows that the series stops there exactly when Equation (26.16) closes on constants, which is the Yang–Mills case of Section 26.6.4 and is why the chapter can display two terms and say “\(+\cdots\)”; and Proposition A.655 is what makes Equation (26.99) well posed, since a cohomology defined by a charge that was only unique up to nothing in particular would not be a property of the theory. The first-class algebra that the whole construction consumes is Proposition 26.13; for general relativity it is the hypersurface-deformation algebra of The Hypersurface-Deformation Algebra, whose structure functions are the reason a general existence theorem was needed at all.
Bertrand's Theorem
This appendix proves Theorem 27.41 of Central Forces and Statics: among all central forces there are exactly two for which every bounded orbit closes, namely the inverse-square attraction \(f(r)=-k/r^{2}\) and the linear attraction \(f(r)=-\kappa r\) of the isotropic oscillator. Here \(k\) carries the unit \(\mathrm{N}\,\mathrm{m}^{2}\) and \(\kappa\) the unit \(\mathrm{N}/\mathrm{m}\), so that \(f\) is a force in newtons in both cases.
Nothing is quoted here. The whole argument is carried out from the Binet equation Equation (27.30) of Central Forces and Statics, Taylor's theorem with remainder (Theorem 7.38) and the elementary fact that a continuous rational-valued function on an interval is constant. There is no imported theorem and therefore no What is quoted here remark.
What the section does contain, and what the reader should watch for, is the division of labour between its two stages. The chapter's roadmap after Theorem 27.41 states it, and it is worth stating again because the point is easy to get wrong: the first-order stage does not isolate the two force laws. It narrows the candidates from all central forces to a one-parameter family of power laws, and no more. The third-order stage is where the two laws are picked out of that family, and it is the real content of the theorem.
The orbit equation and its circular solutions
Throughout, the motion is that of a particle of mass \(m\) (in \(\mathrm{kg}\); for a two-body problem read the reduced mass) in a central field \(f(r)\), with angular momentum \(L=mr^{2}\dot\phi\neq0\) of Equation (27.20), in \(\mathrm{kg}\,\mathrm{m}^{2}/\mathrm{s}\). Write \(u:=1/r\), in \(/\mathrm{m}\).
For a central force \(f\) and a fixed angular momentum \(L\neq0\), the orbit function is
defined for those \(u>0\) at which \(f\) is defined, and carrying the unit \(/\mathrm{m}\). Combining \(ma_{r}=f(r)\) from Equation (19.4) with the Binet equation Equation (27.30) puts the shape of every orbit of that angular momentum in the form
The force is attractive where \(f<0\), that is where \(J>0\). Rests on Equation (27.30), Proposition 27.23 and Equation (19.4).
The time has been eliminated: Equation (A.1050) is an autonomous second-order equation for the shape alone, and closure of an orbit is the statement that its solution \(u(\phi)\) is periodic in \(\phi\) with a period commensurable with \(2\pi\).
A circular orbit of the field at angular momentum \(L\) is a constant solution \(u\equiv u_{0}>0\) of Equation (A.1050), that is a root of
Its index is the dimensionless number
Rests on Definition A.658 and Equation (A.1050).
Let \(r_{0}=1/u_{0}\) be the radius of a circular orbit. Then
which depends on the force law alone and not on \(L\) or \(m\) except through the radius at which it is evaluated. Rests on Definition A.659 and Equation (A.1049).
Derives Lemma A.660. Differentiate Equation (A.1049), writing \(f\) and \(f'\) for the values at \(r=1/u\) and using \(\dd(1/u)/\dd u=-u^{-2}\):
At a circular orbit Equation (A.1051) reads \(u_{0}=-\left(m/L^{2}\right)u_{0}^{-2}f(r_{0})\), so
which incidentally requires \(f(r_{0})<0\): a circular orbit exists only where the force attracts. Substituting Equation (A.1055) into Equation (A.1054) at \(u=u_{0}\),
and Equation (A.1052) gives \(\beta^{2}=1-J'(u_{0})=3+r_{0}f'(r_{0})/f(r_{0})\), which is Equation (A.1053).
∎First stage: closure near a circle forces a power law
Suppose every bounded orbit of the field \(f\) closes, and suppose \(f\) admits circular orbits at every radius of some interval \(0<r_{-}<r<r_{+}\) — which is the case as \(L\) ranges over an interval, by Equation (A.1055). Then \(\beta\) is one fixed rational number on that interval and
where \(C>0\) carries whatever unit makes the right-hand side a force in \(\mathrm{N}\), namely \(\mathrm{N}\) times a length raised to the power \(3-\beta^{2}\). Equivalently, in the orbit variable,
This is as far as the first-order calculation reaches: it leaves a one-parameter family, indexed by \(\beta\), and does not select any member of it. Rests on Lemma A.660 and Equation (27.39).
Derives Proposition A.661. Perturb the circular orbit: put \(u=u_{0}+\delta\) in Equation (A.1050) and expand \(J\) about \(u_{0}\) by Theorem 7.38. Using Equation (A.1051) to cancel the constant terms and Equation (A.1052) to name the coefficient of \(\delta\),
To first order in \(\delta\) the right-hand side is absent and the solution is \(\delta=a\cos\left(\beta\phi\right)\) with the origin of \(\phi\) at an apsis. If \(\beta^{2}\le0\) the solution grows without bound and there are no bounded orbits in the neighbourhood at all, so \(\beta^{2}>0\) and \(\delta\) oscillates. Successive turning points of \(r\) are successive extrema of \(\delta\), separated by \(\Delta\phi=\pi/\beta\); that is, the apsidal angle Equation (27.39) of these orbits satisfies
An orbit closes exactly when the total angle swept between one apsis and the next repetition of the whole configuration is a whole multiple of \(2\pi\), that is when \(\Phi/\pi\) is rational. So \(\beta\) is rational for every circular radius in the interval. By Equation (A.1053), \(\beta\) is a continuous function of \(r_{0}\) wherever \(f\) is continuously differentiable and non-vanishing; a continuous function on an interval taking only rational values has connected image contained in \(\Q\), hence is constant. So \(\beta\) is one fixed rational number.
With \(\beta\) constant, Equation (A.1053) is the separable first-order equation \(rf'/f=\beta^{2}-3\) on the interval, whose solution is \(\log\abs{f}=\left(\beta^{2}-3\right)\log r+\text{const}\), that is Equation (A.1057), with \(C>0\) because \(f<0\) at a circular orbit. Substituting into Equation (A.1049),
which is Equation (A.1058). Every \(\beta>0\) produces a power law satisfying the hypotheses used so far, so nothing has yet been excluded.
∎The condition just used is closure of the orbits infinitesimally near a circle, which is strictly weaker than closure of all bounded orbits. Equation (A.1060) is a limit as the amplitude \(a\) tends to zero, and it fixes the apsidal angle only in that limit; at finite amplitude the apsidal angle of a power law generally depends on \(a\), and an orbit whose apsidal angle drifts with amplitude cannot close at every amplitude even though it closes in the limit. The second stage is precisely the computation of that drift.
For \(J\) of the form Equation (A.1058) with circular orbit at \(u_{0}\),
Derives Lemma A.663. Equation (A.1051) applied to Equation (A.1058) reads \(u_{0}=Au_{0}^{1-\beta^{2}}\), that is
Differentiating Equation (A.1058) three times,
and evaluating at \(u_{0}\) with Equation (A.1065) gives Equations (A.1062), (A.1063) and (A.1064) in turn. Equation (A.1062) is a consistency check on Equation (A.1052): it returns \(\beta^{2}=1-J'(u_{0})\), as it must.
∎Second stage: the third-order solvability condition
The apsidal angle at finite amplitude is computed by the standard device for a nonlinear oscillator: allow the frequency itself to depend on the amplitude, and fix that dependence by demanding that the solution stay periodic. The alternative bookkeeping — keeping the frequency at \(\beta\) and requiring that no term proportional to \(\phi\sin\left(\beta\phi\right)\) appear — gives the same condition, because such a term is exactly the first-order Taylor expansion of the frequency shift.
Let \(J\) be the power law Equation (A.1058) with \(\beta^{2}>0\), and let the radial oscillation about the circular orbit \(u_{0}\) have amplitude \(a\). Then its angular frequency is
with \(J''\), \(J'''\) evaluated at \(u_{0}\), and the apsidal angle is \(\Phi=\pi/\Omega+O(a^{4})\). If every bounded orbit closes then \(\omega_{2}=0\), and
so that \(\beta^{2}=1\) or \(\beta^{2}=4\). Rests on Lemma A.663, Proposition A.661 and Equation (A.1059).
Derives Theorem A.664. Introduce the stretched angle \(\theta=\Omega\phi\), so that \(\dd/\dd\phi=\Omega\,\dd/\dd\theta\) and \(\delta\) is sought as a \(2\pi\)-periodic function of \(\theta\). Writing a dot for \(\dd/\dd\theta\), Equation (A.1059) becomes
Expand in the amplitude,
the frequency carrying only even powers of \(a\) because Equation (A.1068) is unchanged by \(a\mapsto-a\) together with \(\theta\mapsto\theta+\pi\).
Order \(a\). \(\beta^{2}\left(\ddot\delta_{1}+\delta_{1}\right)=0\), so \(\delta_{1}=\cos\theta\), the amplitude being carried by \(a\) and the phase fixed by putting an apsis at \(\theta=0\).
Order \(a^{2}\). Using \(\cos^{2}\theta=\tfrac{1}{2}\left(1+\cos2\theta\right)\),
which has no resonant forcing — neither \(1\) nor \(\cos2\theta\) is a solution of the homogeneous equation — and is solved by
as is checked by applying \(\ddot{\ }+1\): the constant reproduces itself, and \(-\tfrac{1}{3}\cos2\theta\) returns \(-\tfrac{1}{3}\left(-4+1\right)\cos2\theta=\cos2\theta\). The homogeneous solution that could be added here is absorbed into the definition of \(a\).
Order \(a^{3}\). Collecting the terms of that order in Equation (A.1068), and using \(\ddot\delta_{1}=-\cos\theta\),
Reduce the two nonlinear terms with \(\cos\theta\cos2\theta =\tfrac{1}{2}\left(\cos3\theta+\cos\theta\right)\) and \(\cos^{3}\theta =\tfrac{3}{4}\cos\theta+\tfrac{1}{4}\cos3\theta\):
The operator \(\ddot{\ }+1\) annihilates \(\cos\theta\), so Equation (A.1072) has a \(2\pi\)-periodic solution only if the total coefficient of \(\cos\theta\) on its right-hand side vanishes. By Equations (A.1073) and (A.1074) that coefficient is
which is the second member of Equation (A.1066). (The \(\cos3\theta\) terms are non-resonant and merely fix \(\delta_{3}\); they play no part in what follows.) The apsidal angle is half the period of \(\delta\) in \(\phi\), that is \(\pi/\Omega\).
Now impose closure. For each amplitude \(a\) the orbit closes only if \(\Phi/\pi=1/\Omega(a)\) is rational; \(\Omega\) is a continuous function of \(a\) by Equation (A.1066), and a continuous rational-valued function on an interval is constant — the same step used in Proposition A.661, now applied in the amplitude rather than in the radius. Hence \(\Omega\) is independent of \(a\) and \(\omega_{2}=0\). Multiplying Equation (A.1075) by \(-24\beta^{2}\) gives the first member of Equation (A.1067).
It remains to evaluate it. Substituting Equations (A.1063) and (A.1064),
since \(5-5\beta^{2}+3+3\beta^{2}=8-2\beta^{2}\). With \(\beta^{2}>0\) the factor \(\beta^{4}\) cannot vanish, so \(\beta^{2}=1\) or \(\beta^{2}=4\).
∎Each root switches off the condition for a different reason, and it is worth seeing both. At \(\beta^{2}=1\) the power law is \(J(u)=A\), a constant: every derivative of \(J\) vanishes, the right-hand side of Equation (A.1059) is empty to all orders, and the orbit equation is exactly linear — which is Equation (27.47), and is why the Kepler problem has an apsidal angle independent of amplitude rather than merely to third order. At \(\beta^{2}=4\) the derivatives do not vanish; by Equations (A.1063) and (A.1064), \(u_{0}J''=12\) and \(u_{0}^{2}J'''=-60\), and the two terms of Equation (A.1067) cancel against each other: \(u_{0}^{2}\left[5\left(J''\right)^{2}+3\beta^{2}J'''\right] =5\cdot12^{2}+3\cdot4\cdot\left(-60\right)=720-720=0\). The oscillator therefore closes by a cancellation between the second-order feedback of Equation (A.1071) and the direct cubic term, not because its orbit equation is linear.
The two laws, and that they really do close
The argument so far is one of necessity: no force law other than these two can have every bounded orbit closed. Sufficiency has to be checked separately, and must be checked for all bounded orbits and not only for those near a circle, since that is the property claimed.
For \(f(r)=-k/r^{2}\) every bounded orbit is an ellipse with the force centre at a focus, with apsidal angle \(\Phi=\pi\); for \(f(r)=-\kappa r\) every orbit is an ellipse with the force centre at its centre, with apsidal angle \(\Phi=\pi/2\). In both cases \(\Phi\) is independent of the energy and of the angular momentum. Rests on Equations (27.47) and (A.1053).
Derives Proposition A.666. The inverse square. With \(f=-k/r^{2}\), \(rf'/f=r\left(2k/r^{3}\right)/\left(-k/r^{2}\right)=-2\), so Equation (A.1053) gives \(\beta^{2}=1\), consistent with Theorem A.664. The orbit is Equation (27.47), \(1/r=\left(mk/L^{2}\right) \left(1+e\cos\phi\right)\), obtained there without any expansion in amplitude. It is \(2\pi\)-periodic in \(\phi\) for every \(e\), and bounded precisely for \(e<1\); \(r\) is least at \(\phi=0\) and greatest at \(\phi=\pi\), so \(\Phi=\pi\) exactly, at every eccentricity. The force centre sits at \(r=0\), which is a focus of the conic.
The linear law. With \(f=-\kappa r\), \(rf'/f=r\left(-\kappa\right)/\left(-\kappa r\right)=1\), so \(\beta^{2}=4\) and \(\beta=2\). Here the direct route is better than the orbit equation. The force is \(\vect{F}=-\kappa\vect{x}\), so each Cartesian component of Equation (19.4) is an independent harmonic oscillator of the same angular frequency \(\omega_{0}=\sqrt{\kappa/m}\), in \(\mathrm{rad}/\mathrm{s}\). In the plane of the motion,
which is a closed ellipse centred on the origin, traversed once per period \(2\pi/\omega_{0}\), for every \(a_{1}\), \(a_{2}\) and phase \(\gamma\). Every orbit is bounded and every orbit closes. Its radius \(r=\sqrt{x^{2}+y^{2}}\) is a function of \(\cos^{2}\) and \(\sin^{2}\) of \(\omega_{0}t\) and therefore has period \(\pi/\omega_{0}\), half the orbital period: the particle passes two pericentres and two apocentres per revolution, one pair at each end of each principal axis, and the apsidal angle is a quarter of a revolution,
in agreement with Equation (A.1060) — and here the agreement holds at every amplitude, not merely in the limit, which is the content of \(\omega_{2}=0\).
The contrast between the two cases is not a detail of geometry. For the inverse square the centre of force is a focus of the ellipse and the two apsidal distances differ; for the oscillator it is the centre, the two axes are traversed symmetrically, and the orbit closes after half as much angle. That is the same fact as \(\beta=1\) against \(\beta=2\).
∎Three distinct properties are in play and only the strongest is Bertrand's. (i) Some bounded orbit closes: true for every attractive power law, since the circular orbits themselves close. (ii) Every orbit infinitesimally near a circle closes: true for every power law with \(\beta\) rational — a countable but infinite family, including \(f\propto r^{-5/4}\) with \(\beta=\tfrac{1}{2}\) and \(f\propto r^{6}\) with \(\beta=3\). (iii) Every bounded orbit closes: true only for \(\beta=1\) and \(\beta=2\). The whole distance between (ii) and (iii) is Equation (A.1075), which measures how the apsidal angle drifts as the orbit is made less circular. For any other power law, \(\Phi\) varies continuously with the amplitude, takes irrational multiples of \(\pi\) on a dense set of amplitudes, and those orbits fill the annulus \(r_{-}\le r\le r_{+}\) without ever repeating, exactly as Theorem 27.41 asserts.
The theorem is Bertrand's, announced in the Comptes rendus of
the Paris Academy in 1873. That memoir is not among the sources
verified for this treatise and has no key in
references.bib, so the attribution is made here in words, on
the same footing as Darboux's memoir in
Remark A.81. Nothing above rests on it: the
proof given is self-contained, and the third-order stage in particular
is carried out from Equation (A.1059) alone.
Bertrand's Theorem discharges the derivation owed at Theorem 27.41 of Central Forces and Statics. The chapter uses the theorem twice. It uses it to say what the closure of planetary orbits is evidence for — the inverse-square law and nothing weaker — and it uses it, through Phenomenon 27.42, to turn Mercury's residual perihelion advance into a measurement: by Equation (A.1053) an apsidal angle differing from \(\pi\) is a force law differing from \(r^{-2}\), and the size of the difference is the size of the departure. The reader returning there should carry back the division of labour set out at the head of this section, since the chapter's roadmap states it and this appendix is where it is earned: Proposition A.661 narrows all central forces to one parameter, and Theorem A.664 spends that parameter.
Isotropic and Cubic Elastic Tensors
This appendix proves Lemma 30.26 of Continuum Mechanics and Elasticity and the statement the chapter takes from it: that the twenty-one independent constants of Proposition 30.25 collapse to two for an isotropic solid, the Lamé constants of Equation (30.21). It also supplies the intermediate result the chapter never states, and which is the more useful of the two in a laboratory: a solid invariant only under the rotation group of the cube has three independent elastic constants, and the single number by which the general cubic tensor fails to be isotropic is directly measurable.
What is proved here, and what is proved elsewhere
When Continuum Mechanics and Elasticity was written, the classification of the isotropic Cartesian tensors of low rank was carried by no chapter of Part II — Mathematical Methods, and the chapter recorded the gap as a debt owed by the Cartesian-tensor algebra of Differentiable Manifolds, Tensors, and Curvature. That debt is discharged, and not where it was expected: the classification for ranks one to four is Theorem 5.128 of Linear Algebra and Representation Theory — the representation chapter, not the manifolds one, because the statement is invariant theory for \(\SO(3)\) rather than tensor algebra — proved there in full from the definition Equation (5.164) of an isotropic Cartesian tensor. Section 13.2.2 carries the pointer and the constitutive-law reading. Remark 30.27 has been rewritten accordingly and now records what the chapter rests on rather than what it is owed; the Helmholtz decomposition named there is likewise in place, as Theorem 7.105.
This section therefore does not reprove the classification. It quotes Theorem 5.128, adds the two things that theorem does not contain — the independence of the three delta products, and the strictly weaker classification under the finite rotation group of the cube — and then carries out the specialisation to elasticity, which is physics and belongs here.
For reference, the result quoted is this. A Cartesian tensor \(T\) of rank \(N\) is isotropic (Definition 5.126) when its components are unchanged by every \(R\in\SO(3)\),
and Theorem 5.128 states that the isotropic tensors of rank one, two, three and four are respectively \(0\), the multiples of \(\delta_{ij}\), the multiples of \(\varepsilon_{ijk}\), and the linear combinations
which is Equation (30.20) and hence Lemma 30.26. Nothing below reopens that proof.
The three arrays \(\delta_{ik}\delta_{lm}\), \(\delta_{il}\delta_{km}\) and \(\delta_{im}\delta_{kl}\) are linearly independent, so the space of isotropic rank-four Cartesian tensors has dimension exactly three and the constants \(a\), \(b\), \(c\) of Equation (A.1080) are uniquely determined by \(T\). Rests on Theorem 5.128 and Equation (5.168).
Derives Lemma A.671. Evaluate Equation (A.1080) on the three index quadruples \((1,1,2,2)\), \((1,2,1,2)\) and \((1,2,2,1)\). A product of two Kronecker deltas is \(1\) when both pairs it joins carry equal values and \(0\) otherwise, so
Each probe isolates one coefficient and annihilates the other two; hence a vanishing combination has \(a=b=c=0\), and the coefficients of a given \(T\) are read off by Equation (A.1081).
∎The rotation group of the cube leaves four constants
A crystal is not isotropic. Its elasticity tensor is invariant only under the finite point group of its lattice, and for the cubic classes — which include the two commonest structural metals, body-centred iron and face-centred aluminium — that group is the rotation group of the cube. The classification under this smaller group is a strictly weaker statement than Theorem 5.128, and the difference between the two is one number.
The octahedral group \(\mathcal{O}\subset\SO(3)\) is the group of the \(24\) rotations carrying a cube centred at the origin, with faces normal to the coordinate axes, to itself. In the standard frame its elements are exactly the \(3\times3\) matrices carrying a single entry \(\pm1\) in each row and each column and having determinant \(+1\); such a matrix acts as \(R_{ij}=s_{i}\,\delta_{i\,\pi(j)}\) for a permutation \(\pi\) of \(\set{1,2,3}\) and signs \(s_{i}=\pm1\) whose product is \(\sgn\pi\). A tensor obeying Equation (A.1079) for every \(R\in\mathcal{O}\), but not necessarily for every \(R\in\SO(3)\), is called cubic. Rests on Definitions 5.126 and 14.26.
The rank-four cubic Cartesian tensors form a four-dimensional space, spanned by the three products of Equation (A.1080) together with the cubic array
whose components are taken in the crystal frame of Definition A.672. Every cubic \(T\) is therefore
and \(T\) is isotropic if and only if \(f=0\). Rests on Definition A.672 and Lemma A.671.
Derives Proposition A.673. Step 1: the half-turn filter. For each \(p\in\set{1,2,3}\) the diagonal matrix with \(+1\) in the \(p\)-th place and \(-1\) in the other two is a half-turn about the \(p\)-axis; it has determinant \(+1\) and is a signed permutation, so it lies in \(\mathcal{O}\). The argument of Lemma 5.127 uses these three matrices and nothing else, so its conclusion holds verbatim for a cubic tensor: a component \(T_{iklm}\) vanishes unless every value \(p\) occurs among its four indices an even number of times. At rank four the multiplicities \((m_{1},m_{2},m_{3})\) sum to four, so the surviving components have multiplicities \((4,0,0)\) or \((2,2,0)\) in some order. These are the three components \(T_{pppp}\) and, for each of the six ordered pairs \(p\neq q\), the three arrangements \(T_{ppqq}\), \(T_{pqpq}\), \(T_{pqqp}\): twenty-one components in all, and \(81-21=60\) vanishing ones.
Step 2: the three-fold axis. The cyclic matrix \(C\) with \(C_{21}=C_{32}=C_{13}=1\) and every other entry zero is a signed permutation of determinant \(+1\) — a three-cycle is even — so it too lies in \(\mathcal{O}\), and it is the rotation by \(2\pi/3\) about the body diagonal \((1,1,1)/\sqrt{3}\). Substituted into Equation (A.1079) it says that the components are unchanged when every index value is relabelled by the cycle \(\sigma:1\mapsto3\mapsto2\mapsto1\). Hence \(T_{1111}=T_{2222}=T_{3333}\), a common value \(d\), and the ordered pairs \((p,q)\) fall into the two orbits \(\set{(1,2),(3,1),(2,3)}\) and \(\set{(2,1),(1,3),(3,2)}\), within each of which the three arrangements have common values.
Step 3: the four-fold axis. The quarter turn \(Q\) about the third axis, with \(Q_{21}=1\), \(Q_{12}=-1\), \(Q_{33}=1\), is a signed permutation of determinant \(+1\) and lies in \(\mathcal{O}\). It exchanges the values \(1\) and \(2\) and multiplies a component by \((-1)^{m_{1}}\). The components \(T_{1122}\), \(T_{1212}\), \(T_{1221}\) each have \(m_{1}=2\) and so carry the sign \(+1\), giving \(T_{1122}=T_{2211}\), \(T_{1212}=T_{2121}\), \(T_{1221}=T_{2112}\): the two orbits of Step 2 are joined to one another. The three numbers
are therefore independent of which ordered pair is used, and together with \(d=T_{pppp}\) they determine every component: the invariant space has dimension at most four.
Step 4: it has dimension exactly four, and Equation (A.1083) is its general element. The right-hand side of Equation (A.1083) is cubic: \(\Delta\) is unchanged by any relabelling of values and, having all four indices equal on each nonvanishing component, by any sign change; and the three delta products are isotropic by Theorem 5.128, hence a fortiori cubic. Reading off its components with Equation (A.1081) gives \(T_{ppqq}=a\), \(T_{pqpq}=b\), \(T_{pqqp}=c\) and \(T_{pppp}=a+b+c+f\), so the four constants \((a,b,c,f)\) of the ansatz reproduce the four constants \((a,b,c,d)\) of Steps 2 and 3 through
The map \((a,b,c,f)\mapsto(a,b,c,d)\) is a bijection, so the four constants of the ansatz are in one-to-one correspondence with those found in Steps 2 and 3. The four arrays are independent: the three delta products are independent of one another by Lemma A.671, and none of them is a multiple of \(\Delta\), since \(\Delta\) alone changes \(T_{1111}\) without changing \(T_{1122}\), \(T_{1212}\) or \(T_{1221}\). Hence the dimension is exactly four.
Step 5: isotropy is \(f=0\). If \(f=0\) then \(T\) is a combination of the three isotropic products, hence isotropic. Conversely, an isotropic \(T\) is of the form Equation (A.1080) by Theorem 5.128, and its coefficients are those read off by Equation (A.1081), so \(d=T_{1111}=a+b+c\) and \(f=0\) by Equation (A.1085).
∎The whole difference between the cubic and the isotropic classifications is the single relation \(d=a+b+c\), and it is worth being clear about which rotations supply it. Steps 1 to 3 used only the finite group: three half-turns, one three-fold axis, one four-fold axis. Those fix every component in terms of four numbers and can do no more, because the cubic array Equation (A.1082) is genuinely invariant under all \(24\) of them. The relation is produced instead by a rotation through a general angle — in the proof of Theorem 5.128 it is the rotation by \(\theta\) about the third axis, which forces \(2\cos^{2}\theta\sin^{2}\theta\,(a+b+c-d)=0\) and hence, at \(\theta=\pi/4\), \(d=a+b+c\). So the elementary argument stopped one step early does not fail: it delivers cubic anisotropy, which is a physically meaningful and experimentally accessible intermediate result rather than a defective version of isotropy.
From twenty-one constants to three, and then to two
We now impose on Equation (A.1083) the symmetries that an elasticity tensor has by Proposition 30.25: the minor symmetries \(C_{iklm}=C_{kilm}=C_{ikml}\), which follow from the symmetry of the strain and of the stress, and the major symmetry \(C_{iklm}=C_{lmik}\), which follows from the existence of the energy density of Definition 30.24.
Let \(C_{iklm}\) be an elasticity tensor in the sense of Definition 30.24, so that it obeys the minor and major symmetries of Proposition 30.25.
-
If \(C\) is cubic, then in the crystal frame
\begin{equation}\tag{A.1086} C_{iklm}=\lambda\,\delta_{ik}\delta_{lm} +\mu\left(\delta_{il}\delta_{km}+\delta_{im}\delta_{kl}\right) +f\,\Delta_{iklm}\ec \end{equation}three independent constants, all of SI dimension \(\mathrm{Pa}\). The minor symmetries are what force \(b=c=\mu\); the major symmetry is then automatic and imposes nothing further.
-
If \(C\) is isotropic, then \(f=0\) and \(\sigma_{ik}=C_{iklm}u_{lm}\) is Equation (30.21), with two constants.
Rests on Proposition A.673, Proposition 30.25 and Definition 30.24.
Derives Theorem A.675. The minor symmetries. Exchange \(i\) and \(k\) in Equation (A.1083). The first term is unchanged, since \(\delta_{ik}\) is symmetric; the cubic array is unchanged, since it is symmetric under every permutation of its indices; and the middle two terms are exchanged,
Equality of Equation (A.1087) with the original, together with the independence of the two arrays \(\delta_{il}\delta_{km}\) and \(\delta_{im}\delta_{kl}\) established in Lemma A.671, forces \(b=c\). Writing \(a=\lambda\) and \(b=c=\mu\) gives Equation (A.1086). Exchanging \(l\) and \(m\) instead produces the same equation and hence no new condition.
The major symmetry. Send \((ik)\mapsto(lm)\) in Equation (A.1086). The first term becomes \(\lambda\delta_{lm}\delta_{ik}\), which is itself; the second becomes \(\mu(\delta_{li}\delta_{mk}+\delta_{lk}\delta_{mi})\), which is itself because each Kronecker delta is symmetric; and \(\Delta\) is symmetric under every index permutation. So the major symmetry is satisfied identically, and a cubic crystal has exactly the three constants \((\lambda,\mu,f)\).
Isotropy. If \(C\) is isotropic then \(f=0\) by Proposition A.673, and contracting Equation (A.1086) with the symmetric strain \(u_{lm}\) gives
which is Equation (30.21), the second equality using \(u_{ik}=u_{ki}\). Counting: the general tensor has \(21\) constants by Proposition 30.25, a cubic crystal three, an isotropic solid two.
∎In the Voigt notation, in which an index pair \((ik)\) is coded as one of \(1,\ldots,6\) by \(11\mapsto1\), \(22\mapsto2\), \(33\mapsto3\), \(23\mapsto4\), \(13\mapsto5\), \(12\mapsto6\), the three constants of Equation (A.1086) are
so that
and \(A=1\) exactly when \(f=0\), that is exactly when the crystal is elastically isotropic. \(A\) is the Zener anisotropy ratio, and the measured values [Simmons:1971] show that the isotropic idealisation is a poor description of a single crystal and an accidentally good one for a few materials: for body-centred iron \(C_{11}=230\,\mathrm{GPa}\), \(C_{12}=135\,\mathrm{GPa}\), \(C_{44}=117\,\mathrm{GPa}\) give \(A=2.5\); for face-centred aluminium \(C_{11}=108\,\mathrm{GPa}\), \(C_{12}=61\,\mathrm{GPa}\), \(C_{44}=28.5\,\mathrm{GPa}\) give \(A=1.2\); and for tungsten \(C_{11}=523\,\mathrm{GPa}\), \(C_{12}=203\,\mathrm{GPa}\), \(C_{44}=160\,\mathrm{GPa}\) give \(A=1.00\), so that tungsten is isotropic to the precision of the measurement — a coincidence of the numbers, not a consequence of any symmetry. A polycrystalline specimen of iron, on the other hand, is isotropic on scales large against the grain, because the orientations average; that is why Equation (30.21) describes structural steel and would not describe a single iron whisker. Rests on Theorem A.675 and Definition 30.31.
The two lower-rank cases of Theorem 5.128 carry physics of their own and are used elsewhere in Part III — Classical Mechanics. Rank two: an isotropic tensor is \(a\,\delta_{ik}\), so every rank-two material property of an isotropic body — thermal expansion, thermal conductivity, electrical conductivity — is a single scalar times the identity, and the response is parallel to the drive. Rank three: an isotropic tensor is \(a\,\varepsilon_{ikl}\), and \(\varepsilon\) is a pseudotensor: it is invariant under \(\SO(3)\) but changes sign under the inversion \(R=-\identity\), by Equation (5.59). A material property that is a true rank-three tensor and is invariant under the full orthogonal group, inversion included, therefore vanishes identically. The piezoelectric coupling \(P_{i}=d_{ikl}\sigma_{kl}\) is such a property, so no centrosymmetric medium is piezoelectric — quartz is not centrosymmetric and is, a fact used in every quartz oscillator, and no isotropic solid can be. The same parity argument is why an odd-rank isotropic tensor of rank one is zero. Rests on Theorem 5.128 and Remark 5.129.
The dimension three found in Lemma A.671 can be checked against the character theory of Linear Algebra and Representation Theory without recomputing a single component. Write \(V=\R^{3}\) for the defining representation of \(\SO(3)\). An invariant of \(V^{\otimes4}\) is, after using the invariant inner product to identify \(V\) with its dual, an \(\SO(3)\)-equivariant endomorphism of \(V\otimes V\); and the Clebsch–Gordan series Theorem 14.47 decomposes \(V\otimes V\) into the trace, the antisymmetric part and the symmetric traceless part, of dimensions \(1\), \(3\) and \(5\), each occurring once. Schur's lemmas (Theorems 5.153 and 5.155) then say that an equivariant map is a scalar on each summand and zero between distinct ones, so the space of such maps has dimension \(1+1+1=3\).
This is offered as a check and not as the proof, for a reason worth stating. Over the complex numbers Schur's second lemma delivers a scalar; over the reals it delivers only that the endomorphism algebra of an irreducible is a division algebra, which may be \(\R\), \(\C\) or the quaternions, and the multiplicity count above is correct only because each real integer-spin irreducible of \(\SO(3)\) has endomorphism algebra \(\R\). That last step is not carried by Linear Algebra and Representation Theory, whereas the elementary argument of Theorem 5.128 is complete as it stands. Where the two routes agree, as here, the agreement is evidence that neither has dropped a constant.
Isotropic and Cubic Elastic Tensors discharges the derivation owed at Lemma 30.26 of Continuum Mechanics and Elasticity. The classification itself is Theorem 5.128 of Linear Algebra and Representation Theory and is not reproved here (Remark A.670); what this section adds is the uniqueness of the three coefficients (Lemma A.671), the cubic classification (Proposition A.673) that shows exactly which rotations the isotropy of an elastic solid actually consumes, and the reduction \(21\rightarrow3\rightarrow2\) of Theorem A.675, which is the step the derivation of Phenomenon 30.28 takes for granted when it writes \(b=c\) and \(a=\lambda\), \(b=\mu\). Everything downstream of Equation (30.21) in that chapter — the moduli of Definition 30.31, the elastic waves, the plate of Reduction of Three-Dimensional Elasticity to the Kirchhoff Plate Equation and the contact of Hertz's Solution for the Elastic Half-Space Under an Axisymmetric Pressure — rests on the two-constant form proved here.
Reduction of Three-Dimensional Elasticity to the Kirchhoff Plate Equation
This appendix proves Theorem 30.64 of Continuum Mechanics and Elasticity: that the transverse deflection \(w(x,y,t)\) of a thin isotropic plate obeys \(\rho h\,\pp_{t}^{2}w+D\nabla^{4}w=q\) with the flexural rigidity Equation (30.57), and that a free edge carries exactly two boundary conditions rather than the three a naive force balance demands. The reduction has three stages and each is a separate piece of work: a kinematic hypothesis about how the material moves, a constitutive statement that the plate is in plane stress and not plane strain, and a variational calculation whose boundary terms are the interesting part of the answer.
Everything used is in the book. The constitutive law is Equation (30.21), proved in Isotropic and Cubic Elastic Tensors; the conversions among the moduli are Proposition 30.33; the flexural rigidity is Definition 30.63; the variational machinery is that of Calculus of Variations, and in particular the fundamental lemma Lemma 16.18 and the natural boundary conditions Theorem 16.31. Nothing is quoted from outside the treatise. One point of bookkeeping is worth making in advance: Theorem 16.38 states the Euler–Poisson equation for a functional of one independent variable, while the plate functional depends on two, so the corresponding equation is derived here by carrying out the two integrations by parts explicitly — which is necessary anyway, because the boundary terms those integrations produce are the content of The free edge: three conditions are one too many.
Throughout, Greek indices \(\alpha,\beta,\gamma\) run over the in-plane values \(x,y\) and are summed when repeated; Latin indices run over \(x,y,z\). The mid-surface is \(z=0\), the thickness is \(h\) and the material occupies \(\abs{z}\le h/2\). Overdots are \(\pp_{t}\).
Kinematics: the Kirchhoff hypothesis
The plate is said to deform in the Kirchhoff manner when the mid-surface points move only transversely and material lines initially normal to the mid-surface remain straight, normal and unstretched. Writing \(w(x,y,t)\) for the transverse displacement of the mid-surface, this is
so that a cross-section rotates through the angle \(\pp_{\alpha}w\), as in the rod of Proposition 30.51. All three displacements carry \(\mathrm{m}\) and \(w\) is assumed small enough that the strain Definition 30.6 may be linearised. Rests on Definition 30.6 and Proposition 30.51.
Under Equation (A.1091) the in-plane strains are proportional to the curvatures of the deflected mid-surface and the transverse shear strains vanish identically,
while \(u_{zz}\) is not determined by the hypothesis and is fixed instead by the constitutive statement of Plane stress, not plane strain. Rests on Definitions 30.6 and A.680.
Derives Lemma A.681. From Definition 30.6, \(u_{\alpha\beta}=\tfrac{1}{2} \left(\pp_{\alpha}u_{\beta}+\pp_{\beta}u_{\alpha}\right) =-\tfrac{1}{2}z\left(\pp_{\alpha}\pp_{\beta}w +\pp_{\beta}\pp_{\alpha}w\right) =-z\,\pp_{\alpha}\pp_{\beta}w\), the two mixed partials being equal for a twice continuously differentiable \(w\). For the transverse shear,
That the shear vanishes is exactly the statement that normals stay normal, and it is the hypothesis rather than a consequence: a real plate carries a transverse shear stress, and must, since the transverse load has to reach the supports somehow. What Equation (A.1093) asserts is that the shear strain it produces is negligible against the bending strain, which is true to relative order \(\left(h/L\right)^{2}\) for a plate of lateral dimension \(L\).
The remaining component, \(u_{zz}=\pp_{z}u_{z}\), would be zero if Equation (A.1091) were read as an exact statement about every material point. It must not be so read. The ansatz is a statement about the in-plane displacement field and about the rotation of normals; the thickness of the plate does change, by an amount of the same order as the in-plane strains, and \(u_{zz}\) is left free here so that the correct constitutive condition can determine it.
∎Plane stress, not plane strain
This is the step at which a plate differs from a two-dimensional block of material, and it is where the factor \(\left(1-\nu^{2}\right)^{-1}\) of Equation (30.57) is born.
Let the faces \(z=\pm h/2\) be free of traction. Then throughout a thin plate \(\sigma_{zz}\) is negligible, and eliminating \(u_{zz}\) from Equation (30.21) by \(\sigma_{zz}=0\) gives
Both \(\lambda,\mu\) and \(E,\nu\) are the constants of Definition 30.31, related by Proposition 30.33, and every quantity in Equation (A.1094) except \(\nu\) carries \(\mathrm{Pa}\). Rests on Equation (30.21), Proposition 30.33 and Definition 30.63.
Derives Proposition A.682. Why \(\sigma_{zz}\) and not \(u_{zz}\) vanishes. The traction on the faces is \(\sigma_{iz}n_{z}\) with \(n_{z}=\pm1\), so \(\sigma_{zz}=0\) exactly on \(z=\pm h/2\). Between the faces \(\sigma_{zz}\) is of the order of the applied load \(q\), whereas the in-plane stresses are of order \(qL^{2}/h^{2}\), larger by the square of the aspect ratio; setting \(\sigma_{zz}=0\) throughout is therefore an error of relative order \(\left(h/L\right)^{2}\), the same order already discarded in Equation (A.1093). Imposing \(u_{zz}=0\) instead — plane strain — would be a different and wrong theory: it describes a slab confined between rigid platens, and it would put \(\lambda+2\mu\) where Equation (30.57) has \(E/(1-\nu^{2})\), overstating the flexural rigidity by \(\left(\lambda+2\mu\right)\left(1-\nu^{2}\right)/E =\left(1-\nu\right)^{2}/\left(1-2\nu\right)\), a factor \(1.22\) at \(\nu=0.3\).
The elimination. The \(zz\) component of Equation (30.21) reads \(\sigma_{zz}=\lambda\left(u_{\gamma\gamma}+u_{zz}\right)+2\mu u_{zz}\), where \(u_{\gamma\gamma}=u_{xx}+u_{yy}\) is the in-plane trace. Setting this to zero gives the first of Equation (A.1094). Substituting it back into the in-plane components,
since \(\lambda\left[1-\lambda/(\lambda+2\mu)\right]=2\lambda\mu/(\lambda+2\mu)\).
In engineering constants. By Equation (30.28), \(\lambda=2\mu\nu/(1-2\nu)\) and \(2\mu=E/(1+\nu)\), so
Substituting Equation (A.1096) into Equation (A.1095) and pulling out \(2\mu=E/(1+\nu)\) gives the second of Equation (A.1094).
∎Contracting the second of Equation (A.1094) for a state of pure uniaxial in-plane strain, \(u_{xx}=\epsilon\) and \(u_{yy}=0\), gives \(\sigma_{xx}=\left[E/(1+\nu)\right]\left[1+\nu/(1-\nu)\right]\epsilon =E\epsilon/\left(1-\nu^{2}\right)\): the plate is stiffer than a bar of the same material by \(\left(1-\nu^{2}\right)^{-1}\), because its own width prevents the transverse contraction that a bar is free to make. That is the sentence Definition 30.63 states in words, and it is the whole difference between Equation (30.57) and the \(EI\) of Equation (30.49). Rests on Proposition A.682 and Definition 30.63.
Integrating the energy across the thickness
Under Definition A.680 and Proposition A.682, the elastic energy of a plate occupying the region \(\Omega\) of the mid-plane is
with \(D=Eh^{3}/\left[12\left(1-\nu^{2}\right)\right]\) the flexural rigidity Equation (30.57), of SI dimension \(\mathrm{N}\,\mathrm{m}\); \(U\) is then in \(\mathrm{J}\). Rests on Proposition A.682, Lemma A.681 and Definition 30.63.
Derives Proposition A.684. The density. By Equation (30.22) and Equation (30.19) the energy density is \(F=\tfrac{1}{2}\sigma_{ik}u_{ik}\). Of the nine terms, \(\sigma_{\alpha z}u_{\alpha z}\) vanishes because the strain does (Equation (A.1093)) and \(\sigma_{zz}u_{zz}\) vanishes because the stress does (Proposition A.682), so only the in-plane block survives:
Note that the plate stores no energy in the eliminated component: the work done by \(\sigma_{zz}\) through \(u_{zz}\) is zero because \(\sigma_{zz}\) is zero, which is the reason the elimination costs nothing.
Through the thickness. Insert Equation (A.1092). Every strain carries one factor of \(z\), so every term of Equation (A.1098) carries \(z^{2}\), and
Writing \(w_{,\alpha\beta}=\pp_{\alpha}\pp_{\beta}w\) and abbreviating
the integrated energy is
The rearrangement. Expanding \(Q\) and comparing with \(P\),
because \(Q=\left(\pp_{x}^{2}w\right)^{2} +2\,\pp_{x}^{2}w\,\pp_{y}^{2}w+\left(\pp_{y}^{2}w\right)^{2}\). Write \(2G\) for the right-hand side of Equation (A.1102), so that \(P=Q+2G\) with \(G=\left(\pp_{x}\pp_{y}w\right)^{2}-\pp_{x}^{2}w\,\pp_{y}^{2}w\). Then the bracket of Equation (A.1101) is
and the prefactor of Equation (A.1101) combines with the \(\left(1-\nu\right)^{-1}\) to give
which is Equation (A.1097).
∎The integrand \(G=\left(\pp_{x}\pp_{y}w\right)^{2}-\pp_{x}^{2}w\,\pp_{y}^{2}w\) of Equation (A.1097) is an exact divergence,
so that \(\int_{\Omega}G\,\dd A\) depends only on \(w\) and its first derivatives on \(\pp\Omega\). It therefore contributes nothing to the field equation, and nothing at all when the edge is clamped. Rests on Proposition A.684.
Derives Lemma A.685. Differentiate the two products: \(\pp_{x}\left(\pp_{y}w\,\pp_{x}\pp_{y}w\right) =\left(\pp_{x}\pp_{y}w\right)^{2} +\pp_{y}w\,\pp_{x}^{2}\pp_{y}w\) and \(\pp_{y}\left(\pp_{y}w\,\pp_{x}^{2}w\right) =\pp_{y}^{2}w\,\pp_{x}^{2}w +\pp_{y}w\,\pp_{y}\pp_{x}^{2}w\). The two third-derivative terms are equal and cancel in the difference, leaving \(G\). By Theorem 7.101 in the plane, \(\int_{\Omega}G\,\dd A=\oint_{\pp\Omega} \left(\pp_{y}w\,\pp_{x}\pp_{y}w\,n_{x} -\pp_{y}w\,\pp_{x}^{2}w\,n_{y}\right)\dd s\), an expression in the boundary values of \(w\) and \(\nabla w\) alone. A variation \(\delta w\) vanishing together with \(\nabla\delta w\) on \(\pp\Omega\) therefore does not change it, so it cannot enter the Euler equation; and for a clamped edge, where \(w\) and \(\pp_{n}w\) are prescribed, the whole term is a constant that may be discarded from the functional.
∎Lemma A.685 is what makes the plate equation as simple as it is, and a reader who does not see it will suspect an error: the energy Equation (A.1097) is manifestly not proportional to \(\int(\nabla^{2}w)^{2}\), it depends on Poisson's ratio, and yet the field equation contains only the biharmonic operator and no \(\nu\) at all. The resolution is that the \(\nu\)-dependent part is a boundary functional. Geometrically, \(-G\) is the Gaussian curvature of the deflected mid-surface to leading order in \(\nabla w\) — the determinant of the Hessian — so Equation (A.1105) is the small-deflection shadow of the Gauss–Bonnet theorem: the total Gaussian curvature of a surface is determined by its boundary. Poisson's ratio does not disappear from the problem; it survives exactly where the boundary functional lives, in the edge conditions of The free edge: three conditions are one too many, which is why two plates of the same \(D\) and different \(\nu\) have the same equation and different free-edge behaviour. Rests on Lemma A.685 and Definition 30.63.
The field equation
To leading order in \(h/L\) the kinetic energy of the plate is
with \(\rho\) the density in \(\mathrm{kg}/\mathrm{m}^{3}\), so that \(\rho h\) is a mass per unit area in \(\mathrm{kg}/\mathrm{m}^{2}\). The discarded term is the rotatory inertia of the cross-sections, \(\left(\rho h^{3}/24\right)\int_{\Omega} \abs{\nabla\dot{w}}^{2}\dd A\). Rests on Definition A.680.
Derives Lemma A.687. Differentiating Equation (A.1091) in time gives \(\dot{u}_{z}=\dot{w}\) and \(\dot{u}_{\alpha}=-z\,\pp_{\alpha}\dot{w}\), so
using Equation (A.1099). For a disturbance of lateral scale \(L\) one has \(\abs{\nabla\dot{w}}^{2}\sim\dot{w}^{2}/L^{2}\), so the second term is smaller than the first by \(\left(h/L\right)^{2}/12\), the order already discarded twice above, and is dropped. Retaining it, together with the transverse shear strain discarded in Equation (A.1093), gives the Mindlin–Reissner theory, whose flexural waves have a finite limiting speed instead of the unbounded \(\omega\propto k^{2}\) that Equation (30.58) predicts; the two agree for \(kh\ll1\), which is the regime in which the plate theory is being used.
∎Let \(w\) extremise the action
built from Equation (A.1097) and Equation (A.1106), with \(q\) the transverse load per unit area in \(\mathrm{Pa}\). Then in the interior of \(\Omega\)
which is Equation (30.58), and the variation leaves on \(\pp\Omega\) the boundary integral Equation (A.1115) below. Rests on Proposition A.684, Lemma A.687 and Lemma 16.18.
Derives Theorem A.688. Moments. It shortens every formula to introduce the bending moments per unit length of edge,
of SI dimension \(\mathrm{N}\,\mathrm{m}/\mathrm{m}=\mathrm{N}\). With the notation Equation (A.1100), the energy Equation (A.1097) is
the first equality being Equation (A.1103) read backwards, \(Q+2(1-\nu)G=(1-\nu)P+\nu Q\).
The variation. \(U\) is a quadratic form in the second derivatives of \(w\) and \(M_{\alpha\beta}\) is linear in them, so \(\delta U=-\int_{\Omega}M_{\alpha\beta}\, \pp_{\alpha}\pp_{\beta}\delta w\,\dd A\). Integrating by parts twice with Theorem 7.101,
with \(\vect{n}\) the outward unit normal of the edge. Taking the double divergence of Equation (A.1110),
the \(\nu\) cancelling exactly as Remark A.686 promised. The kinetic term contributes \(\delta\int T\dd t=-\int\dd t\int_{\Omega} \rho h\,\pp_{t}^{2}w\,\delta w\,\dd A\) after one integration by parts in \(t\), the endpoint terms vanishing because \(\delta w\) is held zero at \(t_{1}\) and \(t_{2}\). Collecting,
where \(\mathcal{B}\) is the edge integral
For variations supported strictly inside \(\Omega\), \(\mathcal{B}\) vanishes and Lemma 16.18 applied to the area integral gives Equation (A.1109). Since \(D\) carries \(\mathrm{N}\,\mathrm{m}\) and \(\nabla^{4}w\) carries \(/\mathrm{m}^{3}\), each term of Equation (A.1109) is a pressure, as \(q\) is.
∎The free edge: three conditions are one too many
Let \(\vect{n}\) and \(\vect{s}\) be the outward unit normal and the unit tangent of \(\pp\Omega\), forming a right-handed pair, and let \(\pp_{n}=n_{\alpha}\pp_{\alpha}\) and \(\pp_{s}=s_{\alpha}\pp_{\alpha}\). Define
Explicitly, from Equation (A.1110), \(M_{nn}=-D\left[\pp_{n}^{2}w+\nu\,\pp_{s}^{2}w\right]\) for a straight edge, and \(M_{ns}=-D\left(1-\nu\right)\pp_{n}\pp_{s}w\), the trace term dropping out of the second because \(\delta_{\alpha\beta}n_{\alpha} s_{\beta}=\vect{n}\cdot\vect{s}=0\): the twisting moment measures the twist \(\pp_{n}\pp_{s}w\) of the mid-surface at the edge and nothing else. Rests on Equation (A.1110) and Proposition A.682.
Read as a force balance, a free edge — one at which nothing is attached — carries no bending moment, no twisting moment and no transverse shear, which is three conditions: \(M_{nn}=0\), \(M_{ns}=0\), \(Q_{n}=0\). This is Poisson's count, and Remark 30.67 records that it was believed for two decades. It cannot be right. Equation (A.1109) is of fourth order in the two space variables, so along each edge exactly two conditions may be imposed; three over-determine the problem and in general leave it with no solution at all. The resolution is not to discard one of the three arbitrarily but to observe that they are not independent, which is what the variational form Equation (A.1115) makes visible and a force balance cannot. Rests on Definition A.689 and Theorem A.688.
Let \(\pp\Omega\) be a closed piecewise-smooth curve with corners \(c_{1},\ldots,c_{N}\). Then the edge integral Equation (A.1115) may be written
with the effective (Kirchhoff) shear and the corner force
where \(\left[\!\left[M_{ns}\right]\!\right]_{c_{j}}\) is the jump in \(M_{ns}\) across the corner. Consequently, at an edge where nothing is prescribed the natural boundary conditions (Theorem 16.31) are the two statements
together with the vanishing of \(F_{j}\) at every free corner. These are the conditions asserted by Theorem 30.64. Rests on Theorem A.688, Definition A.689 and Theorem 16.31.
Derives Theorem A.691. Resolving the normal derivative. Along the edge, \(\set{\vect{n},\vect{s}}\) is an orthonormal frame of the plane, so \(\pp_{\beta}\delta w =n_{\beta}\,\pp_{n}\delta w+s_{\beta}\,\pp_{s}\delta w\) and
Substituting Equation (A.1122) and Equation (A.1118) into Equation (A.1115) gives the three-term form
which is where Poisson's three conditions appear to come from.
The three are two. The quantities \(\delta w\) and \(\pp_{n}\delta w\) may be prescribed independently on the edge, but \(\pp_{s}\delta w\) may not: it is the derivative of \(\delta w\) along the edge, and is determined by \(\delta w\) there. Integrate that term by parts along the curve. On each smooth arc,
Summing Equation (A.1124) over the arcs of the closed curve, the endpoint terms of consecutive arcs meet at each corner \(c_{j}\) and combine into \(-M_{ns}^{-}\delta w(c_{j})+M_{ns}^{+}\delta w(c_{j})\), where \(\pm\) denote the limits taken along the outgoing and incoming arc; that is \(F_{j}\,\delta w(c_{j})\) with \(F_{j}\) the jump of Equation (A.1120). On a smooth closed curve, \(M_{ns}\) is continuous and every such term cancels. Collecting with the \(Q_{n}\delta w\) term of Equation (A.1123) gives Equation (A.1119).
The conditions. In Equation (A.1119) the three factors \(\pp_{n}\delta w\) on the edge, \(\delta w\) on the edge and \(\delta w\) at each corner are now independently assignable. By Lemma 16.18 applied on the edge and then at the finitely many corners, \(\delta S=0\) for all admissible variations requires each coefficient to vanish wherever its variation is free. Where the plate is clamped, \(\delta w=\pp_{n}\delta w=0\) and nothing is required of \(M_{nn}\) or \(V_{n}\); where it is simply supported, \(\delta w=0\) but \(\pp_{n}\delta w\) is free, giving the single natural condition \(M_{nn}=0\); where it is free, both are free and Equation (A.1121) follows, with \(F_{j}=0\) at a free corner. In every case the count is two conditions per edge, as the order of Equation (A.1109) requires.
∎\(F_{j}\) is not an artefact of the variational bookkeeping. Take a square plate, simply supported on all four edges — resting on knife edges, free to rotate — and load it uniformly. Along each edge \(\delta w=0\), so the edge integral of Equation (A.1119) imposes only \(M_{nn}=0\); but at a corner the tangent turns through a right angle, the roles of \(n\) and \(s\) are exchanged, and \(M_{ns}\) changes sign, so \(F_{j}=\left[\!\left[M_{ns}\right]\!\right]=2M_{ns}\) does not vanish. A downward concentrated force of that magnitude must be supplied at each corner. If the supports cannot pull down — if the plate merely rests on them — the corners lift off, which is exactly what a uniformly loaded square glass pane does, and why a concrete floor slab supported on four walls is reinforced diagonally at its corners. The corner force also explains the sign of the effect: the plate is attempting to become anticlastic, curving up along one diagonal while it curves down along the other, and it is \(M_{ns}\) that measures that twist. Rests on Theorem A.691 and Equation (A.1120).
The fourth-order equation is Germain's [Germain:1821] and the two correct edge conditions are Kirchhoff's [Kirchhoff:1850], obtained exactly as above, from the variation of the energy rather than from a force balance; Poisson's three-condition count is in his memoir on elastic bodies [Poisson:1829]. Three approximations were made and each is of relative order \(\left(h/L\right)^{2}\): the transverse shear strain was set to zero (Equation (A.1093)), the transverse normal stress was set to zero (Proposition A.682), and the rotatory inertia was dropped (Lemma A.687). Two of the three are removed by the Mindlin–Reissner theory, and none of them is a small-deflection assumption: the linearisation of the strain in Definition A.680 is a separate hypothesis, and it is the one that fails first for a thin plate, since a deflection of order \(h\) already stretches the mid-surface and brings in the membrane terms of Remark 30.68.
Reduction of Three-Dimensional Elasticity to the Kirchhoff Plate Equation discharges the derivation owed at Theorem 30.64 of Continuum Mechanics and Elasticity. The three stages announced in that chapter are Plane stress, not plane strain (plane stress, and the birth of the factor \(1-\nu^{2}\) in Equation (30.57)), Integrating the energy across the thickness together with The field equation (the thickness integration and the variation, giving Equation (30.58)), and The free edge: three conditions are one too many (Kirchhoff's resolution of the edge count, with the concentrated corner force). Everything the chapter builds on Equation (30.58) — the Chladni figures of Phenomenon 30.65, the scaling Equation (30.59), and the historical reading of Remark 30.67 — rests on the equation proved here, and the isotropic constitutive law it starts from is Equation (30.21), proved in Isotropic and Cubic Elastic Tensors.
Hertz's Solution for the Elastic Half-Space Under an Axisymmetric Pressure
This appendix completes the derivation of Phenomenon 30.70 of Continuum Mechanics and Elasticity. It is deliberately not a second derivation of that phenomenon. The inline derivation in the chapter already establishes the exponent, and establishes it exactly: from the geometry of two nearly touching spheres, the single strain scale of the contact region and the linearity of the elastic law it obtains \(F\propto E^{*}R^{1/2}\delta^{3/2}\), with \(a=\sqrt{R\delta}\), and nothing about that argument is a guess. What it cannot reach is the pure number in front and the shape of the pressure distribution, because both require actually solving the elastic problem. Those two things — the coefficient \(\tfrac{4}{3}\) of Equation (30.60) and the elliptic profile \(p(r)=p_{0}\sqrt{1-r^{2}/a^{2}}\) — are what this section supplies, together with a proof of the two-body combination rule for \(E^{*}\), which the chapter's derivation asserts in a clause and proves nowhere.
What is quoted here
One result is imported, and everything else is derived from it.
Let an isotropic elastic half-space \(z\ge0\) with Young's modulus \(E\) and Poisson's ratio \(\nu\) (Definition 30.31) be loaded by a normal force \(P\) concentrated at the origin of its free surface, the rest of the surface being free of traction and the displacement vanishing at infinity. Then the normal displacement of the surface at distance \(r\) from the load is
directed into the half-space [Landau:1986]. Here \(P\) carries \(\mathrm{N}\), \(E\) and \(E^{*}\) carry \(\mathrm{Pa}\) and \(u_{z}\) carries \(\mathrm{m}\). Rests on Equation (30.21), Definition 30.31 and Theorem 30.18.
Theorem A.695 is Boussinesq's, published in 1885, and it is the one step of this section that is not carried out from the material of this treatise. Obtaining it needs a general representation of the solutions of the Navier–Cauchy equations of elastostatics in terms of harmonic functions — the Papkovich–Neuber representation — and the potential theory of Section 10.4.3 carries the harmonic functions but not that representation. It is therefore a debt owed by Partial Differential Equations, and it is named here rather than passed over; the statement above is proved in Landau and Lifshitz's treatment of the elastic half-space [Landau:1986], from which the form Equation (A.1125) is taken.
Two features of Equation (A.1125) can be checked here without the representation, and are worth checking, because between them they fix everything except the pure number \(1/\pi\). Dimensions: \(P/(E^{*}r)\) is \(\mathrm{N}/(\mathrm{Pa}\cdot\mathrm{m}) =\mathrm{m}\), as a displacement must be. Scale invariance: the problem has no length in it at all — a half-space is scale-invariant, a point load introduces no length, and linear elasticity has no intrinsic scale — so \(u_{z}\) must be a homogeneous function of \(r\), and linearity in \(P\) together with the dimensions forces the exponent \(-1\). The modulus combination \(E/(1-\nu^{2})\) appears rather than \(E\) because the material beneath a surface load is laterally confined, which is the same statement as the one made for a plate in Remark A.683.
Superposition and the contact conditions
Let a pressure \(p(\rho)\ge0\) act on the disc \(\rho\le a\) of the free surface of the half-space, and vanish outside it. Then the surface normal displacement is
where \(s\) is the distance in the surface plane from the field point to the element \(\dd A'\). The integral converges although the kernel is singular, because \(\dd A'=s\,\dd s\,\dd\varphi\) in polar coordinates centred on the field point. Rests on Theorem A.695.
Derives Lemma A.697. The elastic problem is linear (Equation (30.21)) and the boundary conditions are linear in the applied traction, so displacements superpose. Treating the load on the element \(\dd A'\) as a point force \(P=p\,\dd A'\) and summing Equation (A.1125) gives Equation (A.1126). In polar coordinates \((s,\varphi)\) about the field point the element is \(s\,\dd s\,\dd\varphi\), so the integrand is \(p\,\dd s\,\dd\varphi\) with no singularity at all; the integral exists for any bounded \(p\).
∎Two elastic bodies of radii \(R_{1}\) and \(R_{2}\), with moduli \((E_{1},\nu_{1})\) and \((E_{2},\nu_{2})\), are pressed together along their line of centres by a normal force \(F\), their centres approaching by \(\delta\). Let \(a\) be the radius of the circle of contact and \(p(r)\) the contact pressure. The problem is to find \(p\), \(a\) and the relation between \(F\) and \(\delta\), subject to
-
contact inside the circle: for \(r\le a\) the sum of the two surface displacements closes the initial gap,
\begin{equation}\tag{A.1127} u_{z}^{(1)}(r)+u_{z}^{(2)}(r)=\delta-\frac{r^{2}}{2R}\ec \qquad\frac{1}{R}=\frac{1}{R_{1}}+\frac{1}{R_{2}}\ec \end{equation} -
separation outside: \(p(r)=0\) for \(r>a\), and the surfaces there do not overlap;
-
no adhesion and no edge singularity: \(p(r)\ge0\) throughout, and \(p\) is bounded as \(r\to a\).
The contact is taken to be frictionless, so that the traction transmitted across it is purely normal; and \(a\ll R\), so that near the contact each body is indistinguishable from a half-space and Theorem A.695 applies to it. Rests on Theorem A.695 and Phenomenon 30.70.
Conditions (1) and (2) alone do not determine the answer. The same geometry admits the flat-punch solution \(p(r)\propto\left(1-r^{2}/a^{2}\right)^{-1/2}\), which is a perfectly good solution of the elastic equations, produces a constant surface displacement inside the circle, and diverges at the rim; it is the right answer for a rigid flat cylindrical punch pressed into a half-space, where the contact radius is fixed by the punch and the material at the rim is genuinely singular. What excludes it here is condition (3): the contact circle of two smooth bodies is not fixed externally — it is whatever radius the load produces — so the surfaces must meet the rim tangentially and the pressure must go to zero there, not to infinity. Condition (3) also excludes the tensile solutions that would hold the surfaces together outside the circle, which is the assumption of no adhesion; it fails for very small or very compliant bodies, where surface energy matters, and Hertz's relation is then replaced by an adhesive contact theory. Stating the third condition is what turns the profile below from a lucky guess into the unique answer. Rests on Definition A.698.
The elliptic pressure and its displacement
The whole solution rests on one integral, and it is worth doing in full.
Let \(p(\rho)=p_{0}\sqrt{1-\rho^{2}/a^{2}}\) on \(\rho\le a\) and zero outside. Then for every field point inside the circle, \(r\le a\),
The displacement is therefore quadratic in \(r\) inside the contact circle — which is exactly the shape Equation (A.1127) demands. Rests on Lemma A.697 and Definition A.698.
Derives Lemma A.700. Place the field point at \((r,0)\) with \(0\le r\le a\) and use polar coordinates \((s,\varphi)\) centred on it, so that the element at \((s,\varphi)\) sits at \(\left(r+s\cos\varphi,\ s\sin\varphi\right)\) and is at distance
measured from the centre of the disc. By Lemma A.697 the kernel cancels the Jacobian, so
where \(\ell(\varphi)\) is the distance from the field point to the rim along the direction \(\varphi\), that is the value of \(s\) at which \(\rho=a\).
Completing the square. Put
which is checked by expanding \(A^{2}-u^{2}\). The rim, \(\rho=a\), is \(u=A\); the field point, \(s=0\), is \(u=r\cos\varphi\). Since \(u\) and \(s\) differ by a constant at fixed \(\varphi\), \(\dd u=\dd s\), and
the antiderivative being elementary. At the upper limit the first term vanishes and \(\arcsin1=\pi/2\), giving \(\pi A^{2}/4\). At the lower limit,
a constant independent of \(\varphi\) — the geometric fact that makes the calculation work. Writing \(B=\sqrt{a^{2}-r^{2}}\), the inner integral is
The angular integral. Integrate Equation (A.1134) over \(\varphi\in[0,2\pi)\) term by term. The second term integrates to zero, \(\int_{0}^{2\pi}\cos\varphi \,\dd\varphi=0\). The third also integrates to zero: under \(\varphi\mapsto\pi-\varphi\) the quantity \(A\) is unchanged, since it depends on \(\sin^{2}\varphi\), while \(\cos\varphi\) changes sign and \(\arcsin\) is odd, so the integrand is odd about that involution. The first term gives
using \(\int_{0}^{2\pi}\sin^{2}\varphi\,\dd\varphi=\pi\). Multiplying by the prefactor \(p_{0}/a\) of Equation (A.1130) gives the first of Equation (A.1128), and Equation (A.1126) then gives the second.
∎Matching, and the coefficient $\tfrac{4}{3}$
Under Definition A.698 the two bodies carry the same contact pressure \(p(r)\), so their surface displacements add, and
which is the combination appearing in Equation (30.60). The geometry likewise combines: the gap between the two undeformed surfaces is \(r^{2}/2R\) with \(1/R=1/R_{1}+1/R_{2}\), so the pair is equivalent to a single sphere of radius \(R\) pressed against a rigid plane by the same force. Rests on Lemma A.697 and Definition A.698.
Derives Proposition A.701. The pressures are equal. By Newton's third law the traction each body exerts on the other is equal and opposite at every point of the common contact area, so the same scalar \(p(r)\) appears in Equation (A.1126) for both. Applying that formula to each body with its own modulus and adding,
which is Equation (A.1136). It is the compliances \(1/E^{*}_{i}\) that add and not the moduli, because the two bodies are loaded in series by a common force and their deflections accumulate; that is the clause the chapter's inline derivation asserts without proof, and Equation (A.1137) is the proof. A rigid body has \(E_{i}\to\infty\) and contributes nothing, so a single elastic sphere against a rigid plane has \(E^{*}=E/(1-\nu^{2})\).
The geometry. Measure heights from the plane of first contact. To leading order in \(r/R_{i}\) the surface of a sphere of radius \(R_{i}\) lies at height \(r^{2}/2R_{i}\) below its own pole — the expansion \(R_{i}-\sqrt{R_{i}^{2}-r^{2}}=r^{2}/2R_{i}+O(r^{4}/R_{i}^{3})\) — so the gap between the two undeformed surfaces at distance \(r\) from the axis is \(r^{2}/2R_{1}+r^{2}/2R_{2}=r^{2}/2R\). Both the elastic response and the geometry therefore depend on the pair only through \(E^{*}\) and \(R\).
∎The unique solution of Definition A.698 is the elliptic pressure distribution
with contact radius, approach and force related by
which is Equation (30.60) complete with its coefficient. Rests on Lemma A.700, Proposition A.701 and Definition A.698.
Derives Theorem A.702. Matching. Insert Equation (A.1128) into Equation (A.1136) and set the result equal to the required displacement Equation (A.1127):
Both sides are polynomials in \(r^{2}\) of degree one, so Equation (A.1140) holds identically if and only if the constant terms and the coefficients of \(r^{2}\) agree separately:
That an elliptic pressure produces a displacement of exactly the required quadratic form, with two free constants \(p_{0}\) and \(a\) to match two coefficients, is what makes the problem solvable in closed form; it is the content of Lemma A.700 and it is why Hertz's guess was the right one.
Solving. The second of Equation (A.1141) gives \(p_{0}=2E^{*}a/(\pi R)\), the first of Equation (A.1138). Substituting into the first of Equation (A.1141),
which is \(a=\sqrt{R\delta}\), the contact radius quoted in Phenomenon 30.70.
The force. Integrating the pressure over the contact circle, with the substitution \(t=r^{2}/a^{2}\),
so \(p_{0}=3F/(2\pi a^{2})\): the peak pressure is \(\tfrac{3}{2}\) times the mean pressure \(F/\pi a^{2}\), which is the second equality of Equation (A.1138). Eliminating \(p_{0}\) between Equation (A.1143) and Equation (A.1138), and then \(a\) by Equation (A.1142),
Uniqueness and admissibility. The profile Equation (A.1138) is non-negative and vanishes at \(r=a\), so it satisfies condition (3) of Definition A.698; outside the circle the deformed surfaces separate, since the elastic displacement there falls faster than the parabolic gap grows. That it is the only admissible solution follows from linearity together with the positive definiteness of the elastic energy. Suppose two admissible pressure distributions produced the same \(\delta\) over the same contact circle. Their difference drives an elastic field whose surface traction is purely normal (the contact is frictionless), whose normal displacement vanishes inside the circle, and whose traction vanishes outside it. Its stored energy is \(\tfrac{1}{2}\oint\sigma_{ik}n_{k}u_{i}\,\dd S\) by Equation (30.18) with \(w_{i}=u_{i}\) and no body force, and every term of that integral has one vanishing factor, so the energy is zero. Since the energy density is positive definite (Proposition 30.30), the difference field has zero strain and the two pressures coincide.
The coefficient \(\tfrac{4}{3}\) of Equation (A.1144) is exactly and only what the scaling argument of the chapter left open, and it is assembled from three pure numbers: the \(\pi\) of the Boussinesq kernel Equation (A.1125), the \(\pi^{2}/4\) of the integral Equation (A.1128), and the \(\tfrac{2}{3}\) of Equation (A.1143).
∎Checks and consequences
Example 30.71 presses two steel balls of radius \(10\,\mathrm{mm}\) together with \(1\,\mathrm{N}\), with \(E=210\,\mathrm{GPa}\) and \(\nu=0.29\). Then Equation (A.1136) gives
and \(R=5.0\,\mathrm{mm}\) by Equation (A.1127). Inverting Equation (A.1144) with \(\delta=a^{2}/R\) gives \(a=\left[3FR/(4E^{*})\right]^{1/3}\), so
whence \(\delta=a^{2}/R=2.0\times 10^{-7}\,\mathrm{m}=0.20\,\mu\mathrm{m}\) and, by Equation (A.1138), \(p_{0}=3F/(2\pi a^{2})=4.7\times 10^{8}\,\mathrm{Pa}=470\,\mathrm{MPa}\). All three reproduce Example 30.71, which is the arithmetic check a reader can run in a minute. Note also the ratio \(a/R=6.4\times 10^{-3}\), which is what licenses the half-space approximation of Definition A.698. Rests on Theorem A.702 and Phenomenon 30.70.
The symbol \(E/(1-\nu^{2})\) occurs twice in Continuum Mechanics and Elasticity and means two different things, which is worth stating plainly because the two are easily read as one quantity. In Definition 30.63 it is a single-body stiffness: the plane-strain modulus of the plate material, appearing because a bent plate cannot contract transversely. In Equation (30.60) the same expression is one term of a two-body sum, Equation (A.1136), and the contact modulus \(E^{*}\) is the harmonic-type combination of both bodies' plane-strain moduli. The underlying modulus is the same and its physical reason is the same — lateral confinement — but a contact between two steel bodies has \(E^{*}\) equal to half the plate value for steel, and substituting one for the other misstates the contact stiffness by a factor of two. Rests on Theorem A.702 and Definition 30.63.
The peak pressure is at the centre of the contact, on the surface. The peak shear stress is not: evaluating the subsurface stress field from the same Boussinesq kernel — an elementary but long computation, not carried out here — places the maximum of \(\tfrac{1}{2}(\sigma_{1}-\sigma_{3})\) on the axis at a depth of about \(0.48\,a\) for \(\nu\approx0.3\), where it reaches about \(0.31\,p_{0}\). Since ductile yielding is governed by the maximum shear stress and not by the pressure (Proposition 30.74), first yield under a Hertzian contact begins at that depth, inside apparently undamaged material. For the balls of Example A.703 the peak shear is \(0.31\times470\,\mathrm{MPa}=146\,\mathrm{MPa}\) at a depth of \(15\,\mu\mathrm{m}\). That is the mechanism behind rolling-contact fatigue: a bearing race accumulates plastic damage at a fixed small depth under every pass of a ball, and eventually a subsurface crack reaches the surface and detaches a flake. It is why such a failure is a spall and not wear, and why it is invisible until the moment it is not — the point made at the end of Example 30.71 and taken up in Section 30.8. Rests on Theorem A.702 and Proposition 30.74.
Hertz's Solution for the Elastic Half-Space Under an Axisymmetric Pressure discharges the derivation owed at Phenomenon 30.70 of Continuum Mechanics and Elasticity, and only that part of it which the inline derivation cannot reach: the elliptic pressure profile Equation (A.1138), the coefficient \(\tfrac{4}{3}\) of Equation (30.60), and the two-body compliance rule of Proposition A.701. The exponent \(\tfrac{3}{2}\) and the relation \(a=\sqrt{R\delta}\) are proved in the chapter, by an argument that needs no elasticity solution at all, and are not reproved here. One result is imported and named as such (Remark A.696): Boussinesq's point-force solution Equation (A.1125), a debt owed by Partial Differential Equations, whose Papkovich–Neuber representation this treatise does not carry. Everything else — the potential integral of Lemma A.700 in particular — is carried out in full from the isotropic Hooke law Equation (30.21) proved in Isotropic and Cubic Elastic Tensors.
D'Alembert's Paradox for a Body of Arbitrary Shape
This appendix proves Phenomenon 31.30 of Fluid Dynamics in the generality in which it was stated. The chapter derives the result for a sphere, where the fore-and-aft symmetry of Equation (31.28) makes the cancellation visible in one line; what is owed, and is discharged here, is the statement for a body of any shape, where no such symmetry is available and the vanishing of the force has to be extracted from the structure of the far field instead. The conclusion is stronger than the chapter's phrasing: not only the drag but the whole resultant vanishes, so an arbitrary body in a steady, simply connected, irrotational flow of an ideal fluid feels no lift either. That is what makes Remark 31.31 exhaustive — the only way to buy a transverse force is to break the hypotheses, which is what Theorem 31.68 does.
Statement and hypotheses
Let \(S\) be the smooth closed surface of a finite rigid body held fixed in an unbounded incompressible ideal fluid of uniform density \(\rho\), in \(\mathrm{kg}/\mathrm{m}^{3}\), occupying the exterior region \(\Omega\), and let the motion be steady and irrotational with a single-valued potential \(\phi\) on \(\Omega\), satisfying
Then the resultant force exerted on the body by the fluid pressure vanishes identically,
\(\vect{N}\) being the unit normal on \(S\) directed into the fluid. The statement holds for every shape of \(S\), for every value of \(U\) in \(\mathrm{m}/\mathrm{s}\), and componentwise: there is no drag and no lift. Rests on Proposition 31.28 and Theorem 31.24.
Three hypotheses carry the whole argument and each is named in Remark 31.31 as a way out. Steadiness is what licenses the momentum balance of The momentum theorem on a large sphere. Single-valuedness of \(\phi\) is the simple-connectivity assumption; it is what forbids a circulation, and dropping it is the plane-flow escape of Section 31.7.1. Attachment — that \(\Omega\) really is the whole exterior, with the fluid following the surface everywhere — is what nature declines to provide, and its failure is Section 31.7.2.
The far field carries no source
Write
so that \(\chi\) is harmonic on \(\Omega\) and \(\vect{u}'\), the disturbance velocity, tends to zero at infinity. The first step is that \(\chi\) has no monopole term, and this is a consequence of impermeability alone.
Let \(S_{R}\) be a sphere of radius \(R\) large enough to contain \(S\), with outward unit normal \(\vect{n}\). Then
for every such \(R\). Rests on Theorem 7.101 and Equation (A.1149).
Derives Lemma A.708. Apply the divergence theorem (Theorem 7.101) to \(\vect{u}'\) on the region bounded by \(S\) and \(S_{R}\). Since \(\vect{\nabla}\cdot\vect{u}'=\nabla^{2}\chi=0\) there,
The first integral is zero because \(\pp_{n}\phi=0\) on \(S\): the body is impermeable. The second is zero because the integral of the unit normal over any closed surface vanishes — take \(\vect{a}\) constant in the divergence theorem and read \(\oint_{S}\vect{a}\cdot\vect{N}\,\dd S =\int\vect{\nabla}\cdot\vect{a}\,\dd V=0\).
∎Let \(\chi\) be harmonic in the exterior of a ball of radius \(R_{0}\) in three dimensions and tend to zero at infinity. Then, in spherical coordinates \((r,\theta,\varphi)\) centred in that ball,
the series converging uniformly, together with the series of its term-by-term derivatives of every order, on \(r\ge R_{1}\) for each \(R_{1}>R_{0}\). Here \(Y_{lm}\) is the spherical harmonic Equation (9.163), and the powers \(r^{-l-1}\) are the decaying radial solutions of Equation (10.34) at \(k=0\) (Example 10.33). Rests on Equations (9.163) and (10.34).
Theorem A.709 is the only statement in this section that is not derived. Part II supplies one half of it and not the other. The half it supplies is the basis: separation of Laplace's equation in spherical coordinates, carried out at Example 10.33, produces exactly the radial factors \(r^{l}\) and \(r^{-l-1}\) against the spherical harmonics of Section 9.9.3, and the decay condition selects the second family. The half it does not supply is completeness with convergence — that every exterior harmonic function decaying at infinity is the sum of such a series, uniformly and differentiably. That is a theorem of potential theory, resting on the mean-value property (The Mean-Value Property Characterizes Harmonic Functions) and on the analyticity of harmonic functions, and this treatise does not build it; it is recorded here as a debt against Partial Differential Equations, in the same spirit as the vector identities recorded at Remark 31.5. Its classical use in exactly the present setting — the far field of a body in a stream as a source, a dipole and higher terms — is Lamb's [Lamb:1932]. Nothing below uses more of it than the leading two terms.
There is a constant vector \(\vect{A}\), of dimension \(\mathsf{L}^{4}\mathsf{T}^{-1}\), such that
with \(\vect{n}=\vect{x}/r\). Rests on Lemma A.708 and Theorem A.709.
Derives Corollary A.711. In Equation (A.1151) the \(l=0\) term is \(c_{00}Y_{00}/r\), a point source of strength \(-4\pi c_{00}Y_{00}\); integrating its gradient over \(S_{R}\) gives that strength, and every \(l\ge1\) term integrates to zero over the sphere by the orthogonality of the spherical harmonics (Theorem 9.119). So Lemma A.708 forces \(c_{00}=0\). The three \(l=1\) terms are linear combinations of \(x_{i}/r^{3}\), which is the first expression in Equation (A.1152); the remainder is \(l\ge2\), i.e. \(O(r^{-3})\). Differentiating,
and term-by-term differentiation of the remainder, licensed by Theorem A.709, gives \(O(r^{-4})\).
∎The physical reading of Lemma A.708 is worth one sentence, because it is the only place where the body enters at all: a rigid impermeable body neither creates nor destroys fluid, so it cannot look like a source from far away, and the slowest decay it can produce is that of a dipole. The sphere of Example 31.29 is the worked case, \(\chi=Ua^{3}\cos\theta/2r^{2}\), with \(\vect{A}=\tfrac{1}{2}Ua^{3}\hat{\vect{z}}\).
The momentum theorem on a large sphere
For every \(R\) large enough that \(S_{R}\) encloses \(S\),
with \(\vect{v}=\vect{\nabla}\phi\) and \(p\) the pressure. In particular the right-hand side is independent of \(R\). Rests on Theorems 7.101 and 31.22.
Derives Lemma A.712. Let \(V\) be the fluid region between \(S\) and \(S_{R}\) and let \(\vect{m}\) be the outward normal of \(V\), equal to \(\vect{n}\) on \(S_{R}\) and to \(-\vect{N}\) on \(S\). In steady flow of an ideal fluid with no body force, the Euler equation (Theorem 31.22) combined with \(\vect{\nabla}\cdot\vect{v}=0\) reads \(\pp_{j}\left(\rho v_{i}v_{j}+p\,\delta_{ij}\right)=0\), since \(\rho v_{j}\pp_{j}v_{i}=\pp_{j}(\rho v_{i}v_{j})\). Integrating this over \(V\) and applying the divergence theorem (Theorem 7.101) to each \(i\),
where the momentum flux over \(S\) dropped because \(\vect{v}\cdot\vect{N}=0\) there. The last integral is \(-F_{i}\) by Equation (A.1148), which is the assertion.
∎The pressure is supplied by Bernoulli's theorem (Theorem 31.24) at constant height. With \(q^{2}=\abs{\vect{v}}^{2} =U^{2}+2Uu'_{z}+\abs{\vect{u}'}^{2}\),
By Corollary A.711 the last term is \(O(R^{-6})\) on \(S_{R}\); multiplied by the area element it contributes \(O(R^{-4})\) to Equation (A.1153) and vanishes in the limit. The same estimate disposes of the quadratic part of the momentum flux:
Every surviving term integrates to zero
Assemble Equations (A.1153), (A.1154) and (A.1155) and take the four contributions in turn.
The uniform terms.
\(\oint_{S_{R}}p_{\infty}n_{i}\,\dd S=0\) and \(\rho U^{2}\delta_{i3}\oint_{S_{R}}n_{z}\,\dd S=0\), both because the integral of the unit normal over a closed surface vanishes. A constant pressure and a uniform stream, taken alone, push a closed surface nowhere.
The disturbance flux.
\(\rho U\delta_{i3}\oint_{S_{R}}\vect{u}'\cdot\vect{n}\,\dd S=0\) by Lemma A.708 — exactly, at every \(R\), not merely in the limit.
The two surviving dipole terms.
What is left is
the first term coming from the pressure through Equation (A.1154) and the second from the momentum flux, with the signs of Equation (A.1153) already applied. Insert Equation (A.1152) and write \(\dd S=R^{2}\dd\Omega\) with \(\dd\Omega\) the solid angle:
because the two terms in \(\vect{A}\cdot\vect{n}\) cancel identically — they enter the two integrands as \(-3(\vect{A}\cdot\vect{n})n_{3}n_{i}\) and \(+3(\vect{A}\cdot\vect{n})n_{i}n_{3}\), the same quantity with opposite signs — and because \(\oint n_{i}\,\dd\Omega=0\) over the whole sphere.
That cancellation is the mechanism, and it deserves to be read rather than merely verified. The dipole disturbance contributes to the pressure integral and to the momentum-flux integral separately, and neither contribution is zero; what is zero is their difference, at every \(R\) and for every orientation of \(\vect{A}\), hence for every shape of body.
Derives Theorem A.707. By Lemma A.712 the quantity \(F_{i}\) defined by Equation (A.1153) does not depend on \(R\). By Equations (A.1156) and (A.1157) it equals \(O(R^{-2})\). A constant that is \(O(R^{-2})\) for arbitrarily large \(R\) is zero, so \(\vect{F}=\vect{0}\) — exactly, and for all three components at once.
∎The chapter's proof for the sphere runs on the symmetry of Equation (31.28) about \(\theta=\pi/2\), and a general body has no such symmetry: the pressure distribution over an aerofoil at incidence is grossly asymmetric, and every individual surface element carries a real force. What survives in general is weaker and sufficient — the resultant is fixed by the far field alone, and the far field of any impermeable body is a dipole, whose two contributions to Equation (A.1153) cancel. The near field, where all the physics of shape lives, never enters the calculation. This is also why the theorem is so unforgiving: nothing about the body is used except that it is closed and that the fluid does not pass through it.
D'Alembert's Paradox for a Body of Arbitrary Shape discharges the general case of Phenomenon 31.30 in Fluid Dynamics, where the sphere is done inline and the arbitrary body is deferred here. Read it next to Remark 31.31, whose claim that there are exactly two escapes is now exact rather than rhetorical, because Theorem A.707 kills the resultant and not only its streamwise component. The first escape is multiple connectivity: in the plane, dropping single-valuedness of \(\phi\) admits a circulation and produces the transverse force of Theorem 31.68 — which still leaves the drag zero, so the paradox survives its own resolution for lift, and the appendix computing the circulation itself is The Joukowski Map, the Kutta Condition and the $2\pi$ Lift Slope. The second is that real flow is neither steady nor attached, which is Section 31.7.2, quantified by Theorem 31.71 and measured in Section 36.5.
The Coefficient $6\pi$ in Stokes' Drag Law
This appendix supplies the one thing Phenomenon 31.45 of Fluid Dynamics does not derive. The chapter's inline argument settles the form of Equation (31.39) completely: linearity of Equation (31.38) makes the drag exactly proportional to \(U\), and Theorem D.2 then leaves \(\mu a U\) as the only force that can be built, so \(F=C\mu aU\) with \(C\) a pure number. What is owed is \(C\). It is obtained here by solving the creeping-flow problem for the sphere in closed form and integrating the surface traction, and the calculation returns rather more than a number: it says how the drag divides between pressure and friction, which is the physically informative part and the part the chapter has no room for.
The stream function and the equation it obeys
Work in the frame of the sphere, spherical coordinates \((r,\theta,\varphi)\) with the polar axis along the stream, and take the flow axisymmetric and without swirl, \(v_{\varphi}=0\). Boundary conditions: no slip on the sphere, and a uniform stream at infinity,
The Stokes stream function \(\psi(r,\theta)\), in \(\mathrm{m}^{3}/\mathrm{s}\), is defined by
Rests on Example 31.15 and Equation (7.92).
Incompressibility is then an identity rather than a constraint: in spherical coordinates \(\vect{\nabla}\cdot\vect{v} =r^{-2}\pp_{r}(r^{2}v_{r}) +\left(r\sin\theta\right)^{-1}\pp_{\theta}(\sin\theta\,v_{\theta})\), and substituting Equation (A.1159) gives \(\left(r^{2}\sin\theta\right)^{-1} \left(\pp_{r}\pp_{\theta}\psi-\pp_{\theta}\pp_{r}\psi\right)=0\). This is the axisymmetric counterpart of the plane construction of Example 31.15, and it is why the whole problem collapses to one scalar equation.
Under Definition A.715 the Stokes equations Equation (31.38) are equivalent to
together with a pressure recovered by quadrature. Rests on Definitions 31.43 and A.715.
Derives Lemma A.716. Take the curl of \(\vect{0}=-\vect{\nabla}p+\mu\nabla^{2}\vect{v}\). The pressure disappears because the curl of a gradient vanishes (Equation (7.128)), leaving \(\nabla^{2}\vect{\omega}=\vect{0}\) with \(\vect{\omega}=\vect{\nabla}\times\vect{v}\). In an axisymmetric flow without swirl the vorticity has only an azimuthal component, and substituting Equation (A.1159) into \(\omega_{\varphi}=r^{-1}\left[\pp_{r}(rv_{\theta}) -\pp_{\theta}v_{r}\right]\) gives
The vector Laplacian acting on a purely azimuthal field of this form reduces, term by term in the spherical expression of Equation (7.94), to \(\left(\nabla^{2}\vect{\omega}\right)_{\varphi} =-\left(r\sin\theta\right)^{-1}E^{2}\left(E^{2}\psi\right)\), which is Equation (A.1160). Conversely, given \(\psi\) with \(E^{4}\psi=0\), the momentum equation reads \(\vect{\nabla}p=\mu\nabla^{2}\vect{v} =-\mu\vect{\nabla}\times\vect{\omega}\) by Equation (31.6) and \(\vect{\nabla}\cdot\vect{v}=0\), and its right-hand side is then curl-free, so \(p\) exists and is obtained by integrating along any path.
∎Solution for the sphere
The condition at infinity fixes the angular dependence once and for all. A uniform stream has \(\psi\longrightarrow\tfrac{1}{2}Ur^{2}\sin^{2}\theta\) — read Equation (A.1159) backwards — and both the equation and the boundary conditions are then satisfied by the single separated form
Acting with \(E^{2}\) on Equation (A.1162),
since \(\pp_{\theta}\left(\sin^{-1}\theta\,\pp_{\theta}\sin^{2}\theta \right)=\pp_{\theta}\left(2\cos\theta\right)=-2\sin\theta\). Applying \(E^{2}\) a second time to \(g:=f''-2f/r^{2}\) in the same way and collecting,
The general solution of Equation (A.1163) is \(f=A/r+Br+Cr^{2}+Dr^{4}\), and the conditions Equation (A.1158) select
Rests on Lemma A.716 and Equation (A.1158).
Derives Lemma A.717. Equation (A.1163) is equidimensional: every term scales as \(r^{n-4}\) under \(f=r^{n}\), so the substitution turns it into the algebraic equation
whose left-hand side factorizes as \(\left(n+1\right)\left(n-1\right)\left(n-2\right)\left(n-4\right)\); one checks the four roots directly, \(n=-1\) giving \(24-8-16=0\), \(n=1\) giving \(0-0+8-8=0\), \(n=2\) giving \(0-8+16-8=0\) and \(n=4\) giving \(24-48+32-8=0\). The four powers are independent, so they span the solution space.
The condition at infinity requires \(f\longrightarrow\tfrac{1}{2}Ur^{2}\), which forbids \(r^{4}\) and fixes \(C=U/2\). No slip requires both components of Equation (A.1159) to vanish at \(r=a\), that is \(f(a)=0\) and \(f'(a)=0\):
Multiplying the second by \(a\) and subtracting it from the first eliminates \(B\) and leaves \(2A/a-Ua^{2}/2=0\), whence \(A=Ua^{3}/4\); back-substitution gives \(B=A/a^{2}-Ua=-3Ua/4\).
∎Reading the velocity field off Equations (A.1159) and (A.1164),
Both vanish at \(r=a\) and both tend to the free stream, as they must.
The pressure, and a check
From Equation (A.1164), \(f''-2f/r^{2}=3Ua/2r\), so by Equation (A.1161)
The radial component of \(\vect{\nabla}p=-\mu\vect{\nabla}\times\vect{\omega}\) is then
and integrating from infinity,
It is worth confirming that this is harmonic, since \(\nabla^{2}p=0\) follows from taking the divergence of Equation (31.38): the function \(\cos\theta/r^{2}\) is the \(l=1\) exterior solid harmonic of Example 10.33, and it is. The pressure is high in front (\(\theta=\pi\), taking the stream along \(+\hat{\vect{z}}\) so that \(\theta=\pi\) faces upstream) and low behind, by equal amounts — the fore-and-aft antisymmetry that would give zero drag in an ideal fluid (Phenomenon 31.30). Here it does not, because the pressure is no longer the whole traction.
The traction and its integral
For a Newtonian fluid Equation (31.30) gives, in spherical coordinates,
Evaluate both at \(r=a\) from Equations (A.1165) and (A.1167). For the normal component, \(\pp_{r}v_{r} =U\cos\theta\left(3a/2r^{2}-3a^{3}/2r^{4}\right)\), which vanishes at \(r=a\); the entire normal traction is therefore the pressure,
For the tangential component, \(v_{r}\) vanishes identically on \(r=a\) so its \(\theta\) derivative does too, and \(v_{\theta}/r=-U\sin\theta\left(1/r-3a/4r^{2}-a^{3}/4r^{4}\right)\) gives \(\pp_{r}\left(v_{\theta}/r\right) =-U\sin\theta\left(-1/r^{2}+3a/2r^{3}+a^{3}/r^{5}\right)\), equal to \(-3U\sin\theta/2a^{2}\) at \(r=a\); hence
Both are in \(\mathrm{Pa}\), as \(\mu U/a\) has the dimensions \(\mathsf{M}\mathsf{L}^{-1}\mathsf{T}^{-2}\).
The force exerted on the sphere by the fluid is directed along the stream and has magnitude
of which \(2\pi\mu aU\) — one third — comes from the pressure and \(4\pi\mu aU\) — two thirds — from the tangential viscous traction. Rests on Lemma A.717 and Theorem 31.36.
Derives Theorem A.718. The component along the stream of the traction on the surface element \(\dd S\) is \(\left(\sigma_{rr}\cos\theta-\sigma_{r\theta}\sin\theta\right)\dd S\), the projections of the radial and meridional directions on \(\hat{\vect{z}}\). Substituting Equations (A.1169) and (A.1170),
The uniform pressure integrates to zero over the closed surface, and what remains is a constant traction \(3\mu U/2a\) over the whole sphere:
For the split, keep the two sources separate. The pressure contributes
and the tangential traction contributes
using \(\int_{0}^{\pi}\cos^{2}\theta\sin\theta\,\dd\theta=2/3\) and \(\int_{0}^{\pi}\sin^{3}\theta\,\dd\theta=4/3\). Their sum is Equation (A.1171), which is Equation (31.39) with \(C=6\pi\) [Stokes:1851].
∎The one-third–two-thirds division is a fact about creeping flow that survives no other regime and is easy to get backwards. At high Reynolds number the drag of a bluff body is almost entirely pressure drag, because the wake destroys the rear pressure recovery (Phenomenon 31.55); at low Reynolds number the pressure distribution Equation (A.1167) is exactly the fore-and-aft antisymmetric one, and it contributes only because it is multiplied by \(\cos\theta\) and integrated — it is not the residue of a broken symmetry but the symmetric distribution itself doing work against the motion. The larger part of the resistance is skin friction, which is the sense in which Stokes flow is the opposite of the ideal flow of Section 31.3.3: there the friction was absent and the pressure did nothing, here the friction dominates and the pressure does a third.
Equation (A.1165) decays only as \(a/r\): the term \(-3aU\cos\theta/2r\) in \(v_{r}\) is the leading disturbance, in contrast with the \(a^{3}/r^{3}\) dipole of the ideal-flow solution Equation (31.27). A sphere in creeping flow drags a very large volume of fluid with it, which is why suspensions interact hydrodynamically at many diameters' separation, and why the drag is proportional to \(a\) rather than to \(a^{2}\): the resisted region has the size of the disturbance, not the size of the body. The same slow decay is the origin of the non-uniformity discussed at Remark 31.47 — the neglected inertial term falls off as \(r^{-3}\) while the retained viscous term falls off as \(r^{-4}\), so beyond a distance of order \(a/\mathrm{Re}\) the approximation inverts itself. Whitehead's paradox is the failure of the naive second approximation to satisfy the condition at infinity, and Oseen's repair, Equation (31.41), keeps the advection by the uniform stream in the outer region from the start.
The Coefficient $6\pi$ in Stokes' Drag Law discharges the numerical coefficient of Phenomenon 31.45 in Fluid Dynamics, where Equation (31.39) is derived up to the pure number \(C\) from linearity and Theorem D.2 and the calculation of \(C\) is deferred here. With \(C=6\pi\) in hand, Example 31.46 converts an observed fall speed into a radius, which is how the oil droplets of [Millikan:1913] were sized, and Remark 31.47 states the first correction. Nothing in this section is valid outside \(\mathrm{Re}\ll1\) (Definition 31.43), and the honest statement of its range is Equation (31.41): the leading correction to Equation (A.1171) is a relative \(\tfrac{3}{16}\mathrm{Re}\), so the law is good to one percent only up to \(\mathrm{Re}\approx0.05\).
The Kármán Spacing Ratio of the Vortex Street
This appendix proves Equation (31.54) of Phenomenon 31.61 in Fluid Dynamics: among all staggered double rows of point vortices, exactly one ratio of row separation to streamwise spacing is not destroyed by its own induced motion, namely
[Karman:1911]. What the calculation settles and what it does not is stated in the chapter at Remark 31.62, and the reader should have that remark in view throughout: nothing here bears on the Strouhal number Equation (31.53), and the stability obtained is neutral, not asymptotic. The closing Remark A.731 returns to the point with the calculation in hand.
Throughout, the fluid is ideal, unbounded and two-dimensional; the vortices are point vortices in the sense of Section 31.7.1, each of circulation \(\pm\Gamma\) in \(\mathrm{m}^{2}/\mathrm{s}\), moving with the local velocity induced by all the others. The complex variable is \(z=x+\ii y\), the complex potential \(w\) is that of Equation (31.57), and a vortex of circulation \(\Gamma\) counterclockwise at \(z_{0}\) has \(w=-\left(\ii\Gamma/2\pi\right)\log\left(z-z_{0}\right)\), so that
the second equation being the statement that each vortex is carried by the flow of the others.
One summation formula does all the work
Every lattice sum below is a special case of a single identity, which is proved from the Fourier series of an elementary function and needs nothing else.
Let \(\zeta\in\C\setminus\Z\) and \(0<\varphi<2\pi\). Then
and consequently, differentiating with respect to \(\zeta\),
whose limit as \(\varphi\rightarrow0^{+}\) is \(\Sigma\left(0,\zeta\right)=\pi^{2}/\sin^{2}\pi\zeta\). Rests on Theorem 8.6 and Equation (8.3).
Derives Lemma A.722. Let \(F(\varphi):=\left(\pi/\sin\pi\zeta\right) \ee^{\ii\left(\pi-\varphi\right)\zeta}\) on \((0,2\pi)\) and compute its Fourier coefficients:
the elementary integral of \(\ee^{-\ii\left(\zeta+n\right)\varphi}\). Since \(n\) is an integer, \(\ee^{-2\pi\ii\left(\zeta+n\right)} =\ee^{-2\pi\ii\zeta}\), and \(1-\ee^{-2\pi\ii\zeta} =\ee^{-\ii\pi\zeta}\left(\ee^{\ii\pi\zeta}-\ee^{-\ii\pi\zeta}\right) =2\ii\,\ee^{-\ii\pi\zeta}\sin\pi\zeta\) by Equation (8.3). The two exponentials and the two sines cancel, leaving \(c_{n}=1/(n+\zeta)\), which is Equation (A.1174).
The series Equation (A.1175) converges absolutely and uniformly in \(\zeta\) on compact subsets of \(\C\setminus\Z\), so it may be obtained from Equation (A.1174) by differentiating term by term, using \(\dd\left(n+\zeta\right)^{-1}/\dd\zeta=-\left(n+\zeta\right)^{-2}\) and the product rule on the right-hand side. For the limit, set \(\varphi=0\) in Equation (A.1175): the bracket becomes \(\pi\left(\cos\pi\zeta-\ii\sin\pi\zeta\right)/\sin^{2}\pi\zeta =\pi\,\ee^{-\ii\pi\zeta}/\sin^{2}\pi\zeta\), which cancels the prefactor \(\ee^{\ii\pi\zeta}\) exactly.
∎The complex potential of vortices of circulation \(\Gamma\) placed at \(z=na\) for every \(n\in\Z\) is
and it is the only such potential whose velocity is bounded as \(\abs{\operatorname{Im}z}\rightarrow\infty\). The induced velocity tends to \(\mp\Gamma/2a\) in the \(x\) direction as \(\operatorname{Im}z\rightarrow\pm\infty\). Rests on Equation (A.1173) and Theorem 8.18.
Derives Corollary A.723. The function \(\cot\left(\pi z/a\right)\) is meromorphic with simple poles exactly at \(z=na\), where \(\sin\left(\pi z/a\right)\) has its simple zeros, and near \(z=na\) its principal part is \(a/\pi\left(z-na\right)\). Hence \(\dd w/\dd z\) in Equation (A.1176) behaves near \(z=na\) as \(-\left(\ii\Gamma/2a\right)\cdot a/\pi\left(z-na\right) =-\ii\Gamma/2\pi\left(z-na\right)\), which by Equation (A.1173) is precisely a vortex of circulation \(\Gamma\) there, and it is holomorphic elsewhere. As \(\operatorname{Im}z\rightarrow\pm\infty\), \(\cot\left(\pi z/a\right)\rightarrow\mp\ii\), so \(\dd w/\dd z\rightarrow\mp\Gamma/2a\); since \(\dd w/\dd z=v_{x}-\ii v_{y}\), the row drives the fluid backwards above it and forwards below it, at the stated speed. Any other candidate differs from Equation (A.1176) by a function that is entire (the singularities coincide) and bounded (both velocities are), hence constant by Liouville's theorem (Theorem 8.18).
∎The street and its rigid translation
Fix \(a>0\) and \(h>0\), both in \(\mathrm{m}\). The staggered vortex street consists of the rows
for all \(n\in\Z\). Write \(s:=\pi h/a\), \(c:=\cosh s\) and \(t:=\tanh s\) throughout. Rests on Corollary A.723.
The configuration Equation (A.1177) is an exact solution of Equation (A.1173): every vortex moves with the same velocity
parallel to the rows, so the street is carried downstream without change of shape. Rests on Definition A.724 and Corollary A.723.
Derives Proposition A.725. Take the vortex \(A_{0}=0\). Its own row induces nothing there: the contributions of \(A_{n}\) and \(A_{-n}\) in Equation (A.1173) are equal and opposite. Row \(B\), of circulation \(-\Gamma\) and origin \(a/2-\ii h\), contributes by Equation (A.1176)
using \(\cot\left(\theta-\pi/2\right)=-\tan\theta\) and \(\tan\left(\ii s\right)=\ii\tanh s\). The result is real and positive, so \(v_{x}=\Gamma t/2a\) and \(v_{y}=0\). Repeating at \(B_{0}=a/2-\ii h\) under row \(A\) gives \(-\left(\ii\Gamma/2a\right)\cot\left(\pi/2-\ii s\right) =-\left(\ii\Gamma/2a\right)\tan\left(\ii s\right) =\Gamma t/2a\), the same velocity; and again the vortex's own row contributes nothing. Every vortex therefore moves with Equation (A.1178).
∎Equation (A.1178) is worth reading before the stability calculation, because it is directly visible in the photographs: the street moves more slowly than the stream that made it, since \(U_{\mathrm{s}}\) is the velocity relative to the fluid at infinity and it is directed against the free stream in the frame of the cylinder. In the frame translating with the street the base state is an equilibrium, and since the base velocity is a constant it drops out of the linearized equations altogether; the analysis below may be read in either frame.
Linearization
Displace every vortex,
with \(\abs{\alpha_{n}},\abs{\beta_{n}}\ll a\). Expanding Equation (A.1173) to first order with \(\left(D+\delta\right)^{-1}=D^{-1}-\delta D^{-2}+O(\delta^{2})\), and using Proposition A.725 to cancel the zeroth-order terms against the rigid motion, gives for each \(k\)
the signs following from the circulations \(\pm\Gamma\) carried by the inducing vortices.
Equations (A.1180) and (A.1181) are linear over \(\R\) but not over \(\C\): the left-hand sides carry the complex conjugates of the unknowns while the right-hand sides carry the unknowns themselves. The single-mode ansatz \(\alpha_{n}=\alpha\,\ee^{\ii n\varphi}\) is therefore inconsistent, since \(\bar{\alpha}_{n}=\bar{\alpha}\,\ee^{-\ii n\varphi}\) belongs to the mode \(-\varphi\): matching the \(k\) dependence of the two sides would force \(\varphi\equiv0\). The wavenumbers \(\varphi\) and \(-\varphi\) close on each other and on nothing else, so the smallest closed system is four-dimensional. Overlooking this is the one way the calculation can be got wrong while looking right, and it changes the answer: the two-dimensional system obtained by ignoring it makes the street unstable at every spacing, Equation (A.1172) included.
Accordingly set, for a fixed \(\varphi\in(0,2\pi)\),
Three lattice sums appear. With \(m=n-k\) and \(\zeta=\tfrac{1}{2}-\ii h/a\), so that \(\pi\zeta=\pi/2-\ii s\) and therefore \(\sin\pi\zeta=\cosh s=c\) and \(\cos\pi\zeta=\ii\sinh s\):
where the closed forms come from Lemma A.722 and
For Equation (A.1183) use \(\sum_{m\ge1}\cos m\varphi/m^{2} =\pi^{2}/6-\pi\varphi/2+\varphi^{2}/4\) on \([0,2\pi]\); for Equation (A.1185) substitute \(\sin\pi\zeta=c\), \(\cos\pi\zeta=\ii\sinh s\) and \(\ee^{\ii\psi\zeta}=\ee^{\ii\psi/2}\ee^{\psi s/\pi}\) in Equation (A.1175). The companion sum with \(\varphi\) replaced by \(-\varphi\), i.e. \(\psi\) by \(-\psi\), is
Two further sums reduce to these: replacing \(n\) by \(-n\) shows that \(\sum_{n}\left(\left(n-\tfrac{1}{2}\right)a+\ii h\right)^{-2}=S_{0}\) and \(\sum_{n}\ee^{\ii n\varphi} \left(\left(n-\tfrac{1}{2}\right)a+\ii h\right)^{-2}=\tilde{S}_{1}\).
Substituting Equation (A.1182) into Equations (A.1180) and (A.1181) and matching the coefficients of \(\ee^{\ii k\varphi}\) and \(\ee^{-\ii k\varphi}\) separately — noting \(P(-\varphi)=P(\varphi)\) and \(S_{1}(-\varphi)=\tilde{S}_{1}(\varphi)\) — yields, with \(K:=\Gamma/2\pi\) and
the closed four-dimensional system
in the variables \(p:=\alpha\), \(q:=\beta\), \(r:=\bar{\gamma}\), \(u:=\bar{\delta}\), with
The growth rates
Solutions of Equation (A.1189) proportional to \(\ee^{\sigma t}\) have \(\sigma^{2}\) an eigenvalue of \(\mathsf{B}\mathsf{C}\). Three real quantities suffice to write it down:
the last being real because the two phases in Equations (A.1185) and (A.1187) cancel and \(\left(\ii\pi\right)^{2}=-\pi^{2}\); note \(XY=\Pi^{2}\).
The eigenvalues \(\lambda=\sigma^{2}\) of \(\mathsf{B}\mathsf{C}\) satisfy
Derives Lemma A.727. Multiplying the matrices of Equation (A.1190),
whose trace is \(-K^{2}\left(-2Q^{2}+Y+X\right)\), i.e. the coefficient displayed. Its determinant is the product of the two determinants. Now \(\overline{S_{1}}=\ee^{\ii\varphi}S_{1}\) — replace \(n\) by \(-n-1\) in the series defining \(\overline{S_{1}}\) and compare — and likewise \(\overline{\tilde{S}_{1}}=\ee^{-\ii\varphi}\tilde{S}_{1}\), so the determinant of \(\mathsf{B}\) is \(-K^{2}\left(-Q^{2} +\overline{\tilde{S}_{1}}\,\overline{S_{1}}\right) =K^{2}\left(Q^{2}-\Pi\right)\), and that of \(\mathsf{C}\) is the same. The characteristic polynomial of a \(2\times2\) matrix is \(\lambda^{2}-\left(\text{trace}\right)\lambda +\left(\text{determinant}\right)\), which is Equation (A.1192).
∎The mode \(\varphi\) neither grows nor decays — all four roots \(\sigma\) are purely imaginary — if and only if
and it grows otherwise. Rests on Lemma A.727.
Derives Proposition A.728. Since \(\sigma^{2}=\lambda\), the four exponents are purely imaginary exactly when both roots of Equation (A.1192) are real and non-positive; if either root is complex, or real and positive, some \(\sigma\) has positive real part. Write \(\mathcal{T}=2Q^{2}-X-Y\) and \(\mathcal{D}=\left(Q^{2}-\Pi\right)^{2}\ge0\) for the two coefficients divided by \(K^{2}\) and \(K^{4}\). The discriminant factorizes,
because \(\mathcal{T}\pm2\left(Q^{2}-\Pi\right)\) are the two brackets. Put \(\pi_{+}:=\sqrt{X}\), \(\pi_{-}:=\sqrt{Y}\), so \(\Pi=\pm\pi_{+}\pi_{-}\) by \(XY=\Pi^{2}\), the sign being that of \(-MM'\). If \(MM'>0\) then \(\Pi=-\pi_{+}\pi_{-}\) and the factorization reads \(-\left(\pi_{+}+\pi_{-}\right)^{2} \left[4Q^{2}-\left(\pi_{+}-\pi_{-}\right)^{2}\right]\), non-negative exactly when \(2\abs{Q}\le\abs{\pi_{+}-\pi_{-}}\); if \(MM'<0\) then \(\Pi=+\pi_{+}\pi_{-}\) and it reads \(-\left(\pi_{+}-\pi_{-}\right)^{2} \left[4Q^{2}-\left(\pi_{+}+\pi_{-}\right)^{2}\right]\), non-negative exactly when \(2\abs{Q}\le\pi_{+}+\pi_{-}\). Both cases are covered by \(2\abs{Q}\le\abs{EM-M'/E}\pi/a^{2}\), since \(EM\) and \(M'/E\) have the same sign as \(M\) and \(M'\) respectively; and the closed form on the right of Equation (A.1193) follows from Equation (A.1186),
Finally the second requirement, \(\mathcal{T}\le0\), is implied by the first: in either case \(4Q^{2}\le\left(\pi_{+}\pm\pi_{-}\right)^{2} \le2\left(X+Y\right)\), so \(2Q^{2}\le X+Y\).
∎Only one spacing survives
The staggered street of Definition A.724 is neutrally stable to every infinitesimal disturbance Equation (A.1182) if and only if
which is Equation (31.54). For every other ratio some mode grows exponentially. Rests on Proposition A.728 and Definition A.724.
Derives Theorem A.729. Necessity. Take \(\varphi=\pi\), that is \(\psi=0\): neighbouring vortices of a row are then displaced in opposite senses, the disturbance of longest wavelength that is not a rigid translation. Then \(u=0\), \(E=1\) and \(M=M'=\pi t/c\), so the right-hand side of Equation (A.1193) vanishes identically and the criterion demands \(Q=0\). By Equation (A.1188) with \(\psi=0\) this is \(\pi^{2}/2=\pi^{2}/c^{2}\), i.e.\ Equation (A.1194). Solving, \(\sinh s=\sqrt{c^{2}-1}=1\) and \(s=\log\left(c+\sinh s\right)=\log\left(1+\sqrt{2}\right)\), whence Equation (A.1172).
Sufficiency. Let \(c=\sqrt{2}\), so \(t=1/\sqrt{2}\), \(\sinh s=1\) and \(s=\log\left(1+\sqrt{2}\right)\). Then \(Q=-\psi^{2}/2a^{2}\) by Equation (A.1188), and Equation (A.1193) becomes, after multiplying by \(a^{2}\) and using \(c=\sqrt{2}\),
Both sides are even in \(\psi\), so take \(0\le\psi\le\pi\), where the bracket is negative and Equation (A.1195) is the assertion \(G(u)\ge0\) for \(0\le u\le s\), with \(\psi=\pi u/s\) and
obtained by dividing Equation (A.1195) by \(\pi^{2}\) and substituting \(\sqrt{2}=\cosh s\), \(1=\sinh s\). Now \(G(0)=0\) and \(G(s)=\cosh^{2}s-\sinh s-1=2-1-1=0\); moreover
Since \(2\cosh s/s>1\), every \(u\)-dependent term of \(G''\) is strictly increasing on \([0,s]\), so \(G''\) is; and \(G''(0)=-2/s^{2}<0\) while \(G''(s)=2\sqrt{2}/s+1-2/s^{2}>0\) numerically (\(s=0.88137\) gives \(3.209+1-2.575\)). Hence \(G''\) changes sign exactly once, so \(G'\) decreases and then increases. As \(G'(0)=\cosh s/s-1>0\) and \(G'(s)=\cosh^{2}s/s+\cosh s\sinh s-\cosh s-2/s=0\) — the two terms \(\cosh^{2}s/s=2/s\) and \(-2/s\) cancel, as do \(\cosh s\sinh s=\sqrt{2}\) and \(-\cosh s=-\sqrt{2}\) — the function \(G'\) is positive on \((0,u^{*})\) and negative on \((u^{*},s)\) for a single \(u^{*}\). Therefore \(G\) rises from \(G(0)=0\) and falls back to \(G(s)=0\) without crossing zero in between, so \(G\ge0\) on \([0,s]\), which is Equation (A.1195). Equality holds only at \(\psi=0\) and \(\psi=\pi\), the mode of Equation (A.1194) and the rigid translation.
Instability elsewhere. If \(c\neq\sqrt{2}\) then \(Q\neq0\) at \(\psi=0\), while \(X=Y=-\Pi=\pi^{4}t^{2}/a^{4}c^{2}\) there, so the discriminant computed in Proposition A.728 equals \(\left[4\Pi\right]\left[4Q^{2}\right]=16\Pi Q^{2}<0\): the roots \(\lambda\) are a complex conjugate pair, their square roots are not purely imaginary, and the mode \(\varphi=\pi\) grows.
∎If the second row is placed at \(B_{n}=na-\ii h\), directly beneath the first, the mode \(\varphi=\pi\) grows for every \(h>0\). Rests on Theorem A.729.
Derives Proposition A.730. The only change is in the offset, so Lemma A.722 is applied at \(\zeta=-\ii h/a\) instead of \(\tfrac{1}{2}-\ii h/a\), giving \(\sin\pi\zeta=-\ii\sinh s\) and \(\cos\pi\zeta=\cosh s\); hence \(S_{0}=-\pi^{2}/a^{2}\sinh^{2}s\), and \(S_{1}\), \(\tilde{S}_{1}\) come out real. At \(\psi=0\) they are equal, with \(\sqrt{X}=\pi^{2}\cosh s/a^{2}\sinh^{2}s\) and \(\Pi=+X\), so the second case of Proposition A.728 applies and neutrality would require \(\abs{Q}\le\sqrt{X}\). Here \(Q=\left(\pi^{2}/a^{2}\right)\left(\tfrac{1}{2} +1/\sinh^{2}s\right)\), so the requirement reads \(\tfrac{1}{2}\sinh^{2}s+1\le\cosh s\), i.e.\ \(\left(\cosh s-1\right)^{2}\le0\) after using \(\sinh^{2}s=\cosh^{2}s-1\). That holds only at \(s=0\), where the two rows coincide. For every real separation the unstaggered arrangement therefore grows, and no spacing rescues it — which is why only the staggered pattern is ever photographed.
∎Four limitations are built into the statement proved above and none of them is repaired by any refinement of the algebra.
First, the stability is neutral. Theorem A.729 says that at the critical ratio no linear mode grows; it does not say that any decays, and it says nothing at second order, where the street is in fact unstable. The observed consequence is that the photographed spacing drifts slowly downstream instead of locking to Equation (A.1172), and the agreement with the photographs — close to a few percent — is better than a neutral result has any right to expect.
Second, the vortices are points in an unbounded ideal fluid. Real vortices have cores of finite size, diffuse by Equation (31.52), and decay; the calculation has no viscosity in it anywhere.
Third, and most important, the cylinder never appears. The configuration Equation (A.1177) is postulated, not derived: nothing here explains why two boundary layers separating from a bluff body should roll up into two staggered rows, nor with what strength \(\Gamma\), nor at what rate. In particular nothing here predicts the Strouhal number, and Remark 31.62 says so at length: Equation (31.53) is a measurement, disciplined by the similarity argument of Phenomenon 31.49 into the form \(\mathrm{St}=\mathrm{St}(\mathrm{Re})\) and no further.
Fourth, the analysis is of a rigid infinite pattern already formed. The selection it performs is therefore of the same kind as the selection of a preferred wavelength in Section 31.8.1: it says which configurations can persist, not which one nature will build.
The Kármán Spacing Ratio of the Vortex Street discharges Equation (31.54) of Phenomenon 31.61 in Fluid Dynamics, and only that equation: Equation (31.53), the Strouhal number that shares the phenomenon with it, is reported there as measurement and is not derived here or anywhere else in this treatise. The two halves of that phenomenon are of different epistemic kinds and Remark 31.62 is where the difference is set out; the reader who has followed the algebra above should return to it, because the calculation just performed makes concrete how much idealization purchased the one number that could be derived. The measured record for the cylinder wake is collected in Experiment: Fluid Flow and Turbulence.
The Joukowski Map, the Kutta Condition and the $2\pi$ Lift Slope
This appendix proves Proposition 31.70 of Fluid Dynamics — that the Kutta condition selects Equation (31.60) from the one-parameter family of irrotational flows past a thin profile, giving a lift-curve slope of exactly \(2\pi\) per radian. The chapter proves Theorem 31.68, which converts a circulation into a force and says nothing about what the circulation is; Remark 31.69 names the physical selection principle and defers the computation here. Three things are established below: the flow past a circle with arbitrary circulation (Flow past a circle with circulation), the conformal transfer to the profile with the trailing edge as the distinguished point (The Joukowski map and the cusped edge), and the value of \(\Gamma\) that the finiteness of the velocity there enforces (The Kutta condition and the lift).
The sign convention is the chapter's throughout: \(\Gamma\) is the circulation measured counterclockwise, Equation (31.55), and the transverse force per unit span is \(Y=-\rho U\Gamma\), Equation (31.59). The complex velocity is Equation (31.57), \(\dd w/\dd z=v_{x}-\ii v_{y}\).
Flow past a circle with circulation
Let \(U>0\) in \(\mathrm{m}/\mathrm{s}\), \(R>0\) in \(\mathrm{m}\), \(\alpha\in\left(-\pi/2,\pi/2\right)\) and \(\Gamma\in\R\) in \(\mathrm{m}^{2}/\mathrm{s}\). The function
is holomorphic on \(\abs{\zeta}>R\) apart from the branch cut of the logarithm, has \(\abs{\zeta}=R\) as a streamline, tends to the uniform stream \(U\ee^{\ii\alpha}\) at infinity, and carries circulation \(\Gamma\) around the circle. Its derivative on the circle is
so the stagnation points on the circle are at the angles \(\theta\) with
Rests on Proposition 31.28 and Equation (31.57).
Derives Lemma A.733. On \(\zeta=R\ee^{\ii\theta}\) the first bracket of Equation (A.1196) is \(UR\left(\ee^{\ii\left(\theta-\alpha\right)} +\ee^{-\ii\left(\theta-\alpha\right)}\right) =2UR\cos\left(\theta-\alpha\right)\), which is real, and \(\log\zeta=\log R+\ii\theta\), so \(\operatorname{Im}w=-\left(\Gamma/2\pi\right)\log R\), a constant: the circle is a streamline. As \(\zeta\rightarrow\infty\), \(\dd w/\dd\zeta\rightarrow U\ee^{-\ii\alpha}\), which by Equation (31.57) is the stream \(\vect{v}=U\left(\cos\alpha,\sin\alpha\right)\). The circulation is \(\oint\left(\dd w/\dd\zeta\right)\dd\zeta\) around the circle; the first bracket is single valued and contributes nothing, and \(-\left(\ii\Gamma/2\pi\right)\oint\dd\zeta/\zeta =-\left(\ii\Gamma/2\pi\right)2\pi\ii=\Gamma\) by Theorem 8.24.
For Equation (A.1197), differentiate Equation (A.1196) and set \(\zeta=R\ee^{\ii\theta}\):
and the bracket is \(2\ii\sin\left(\theta-\alpha\right)\) by Equation (8.3). Factoring out \(\ii\ee^{-\ii\theta}\) gives Equation (A.1197), which vanishes exactly at the angles Equation (A.1198).
Uniqueness among flows with the stated data is the Neumann uniqueness of Proposition 31.28: two such potentials differ by a function whose velocity vanishes at infinity, whose normal derivative vanishes on the circle, and around which the circulation is zero, hence by the maximum principle (Theorem 10.77) a constant.
∎Note that Equation (A.1198) has a solution only for \(\abs{\Gamma}\le4\pi UR\); larger circulations lift the stagnation points off the circle entirely. Nothing in the potential problem picks a value — every \(\Gamma\) gives a legitimate irrotational flow that does not penetrate the body, which is exactly the indeterminacy stated at Remark 31.69.
The Joukowski map and the cusped edge
For \(b>0\) in \(\mathrm{m}\), the map
carries \(\abs{\zeta}>b\) bijectively onto the complement of the segment \(\left[-2b,2b\right]\) of the real axis, with inverse
the branch of the square root behaving as \(z\) at infinity. Rests on Theorem 8.6.
That Equation (A.1200) inverts Equation (A.1199) is the quadratic formula applied to \(\zeta^{2}-z\zeta+b^{2}=0\), whose two roots have product \(b^{2}\); the branch named is the one of modulus greater than \(b\), and it is holomorphic off the segment because \(z^{2}-4b^{2}\) vanishes only at \(z=\pm2b\). The circle \(\abs{\zeta}=b\) itself maps two-to-one onto the segment: \(\zeta=b\ee^{\ii\theta}\) gives \(z=2b\cos\theta\). The chord is therefore \(c=4b\), and the point \(\zeta=b\) maps to the trailing edge \(z=2b\).
Let \(w\) be as in Equation (A.1196) with \(R=b\), and put \(W(z):=w\bigl(\zeta(z)\bigr)\) with Equation (A.1200). Then \(W\) is a complex potential for the flow past the segment: it is holomorphic off the segment, the segment is a streamline, the free stream at infinity is \(U\ee^{\ii\alpha}\), the circulation around the segment is \(\Gamma\), and the physical velocity is
Rests on Definition A.734 and Lemma A.733.
Derives Lemma A.735. A composition of holomorphic functions is holomorphic and its derivative is given by the chain rule; both statements follow from Theorem 8.6 and Definition 8.5 exactly as for real functions, and \(\dd z/\dd\zeta=1-b^{2}/\zeta^{2}\) is non-zero on \(\abs{\zeta}>b\), which gives Equation (A.1201). Since \(W\) takes the same values as \(w\), its imaginary part is constant on the image of the circle, i.e. on the segment, and its behaviour at infinity is that of \(w\) because \(\zeta(z)=z+O\left(z^{-1}\right)\) there. The circulation is unchanged for the same reason: it is the real part of \(\oint_{C_{z}}\dd W\), and \(\dd W=\dd w\) under the substitution \(z\mapsto\zeta(z)\), which carries a circuit around the segment to a circuit around the circle.
∎The step just taken is the only one where a general theorem might have been invoked, and it has been avoided. What is usually quoted here is that a holomorphic bijection carries harmonic functions to harmonic functions and streamlines to streamlines, so a solved problem in one plane is a solved problem in the other; the general statement needs the holomorphic inverse function theorem, and beyond it, to know that a suitable map exists at all, the Riemann mapping theorem. Part II's complex-analysis chapter carries neither: Complex Analysis ends at Theorem 8.24 and has no conformal-mapping material, which is a real gap and is recorded as such. Nothing above needs it, because the map Equation (A.1199) and its inverse Equation (A.1200) are both written down explicitly, and the only property used is that a composition of holomorphic functions is holomorphic — which is Theorem 8.6 and the chain rule.
The transformation is degenerate at \(\zeta=\pm b\), where \(\dd z/\dd\zeta=0\). This is not an accident of the map but the definition of a sharp edge: the interior angle of the profile at \(z=2b\) is zero, and a smooth curve is being folded onto itself. Equation (A.1201) then says that the physical velocity at the trailing edge is infinite — unless the numerator vanishes there too.
The Kutta condition and the lift
Of the one-parameter family of flows Equation (A.1196) with \(R=b\), exactly one has a finite velocity at the trailing edge \(z=2b\), namely that with
and for it Equation (31.59) gives the lift per unit span and lift coefficient
which is Equation (31.60). Rests on Lemma A.735 and Theorem 31.68.
Derives Theorem A.737. By Equation (A.1201) the velocity at \(z=2b\) is the quotient of \(\dd w/\dd\zeta\) and \(1-b^{2}/\zeta^{2}\) as \(\zeta\rightarrow b\) along the circle. The denominator has a simple zero there: with \(\zeta=b\ee^{\ii\theta}\), \(1-\ee^{-2\ii\theta}=2\ii\ee^{-\ii\theta}\sin\theta\), which vanishes linearly in \(\theta\). The numerator is Equation (A.1197) with \(R=b\). The quotient is therefore
and its limit as \(\theta\rightarrow0\) is finite if and only if the numerator vanishes at \(\theta=0\), that is \(-2U\sin\alpha=\Gamma/2\pi b\), which is Equation (A.1202); every other value of \(\Gamma\) gives a velocity diverging like \(1/\theta\). (When the condition holds, l'Hôpital's rule gives the finite edge velocity \(U\cos\alpha\), the component of the stream along the plate — the flow leaves the edge smoothly, along the chord.)
For the force, Lemma A.735 places a circulation \(\Gamma\) around a body in a uniform stream of speed \(U\), so Theorem 31.68 applies in axes aligned with the stream: the drag vanishes and the transverse force is \(-\rho U\Gamma=\pi\rho U^{2}c\sin\alpha\), directed perpendicular to the stream and, since \(\Gamma<0\) for \(\alpha>0\), upwards. Dividing by \(\tfrac{1}{2}\rho U^{2}c\) gives Equation (A.1203), and for small incidence \(\sin\alpha=\alpha+O\left(\alpha^{3}\right)\), so \(\Gamma\rightarrow-\pi cU\alpha\) and \(C_{\mathrm{L}}\rightarrow2\pi \alpha\) as in Equation (31.60).
∎The step that makes the calculation possible is that \(\Gamma\) is the same number in the two planes. It is worth isolating, because it is what licenses computing the circulation where the geometry is easy — a circle — and applying Theorem 31.68 where the geometry is the one that matters. The reason is that the circulation is a contour integral of an exact differential, \(\oint\dd W\), and a change of variable in a contour integral is a substitution and nothing more: no property of the map enters beyond its being holomorphic and one-to-one. The same remark explains why the Blasius formula Equation (31.58) may be evaluated on a large circle rather than on the profile — Corollary 8.15, used already in the chapter's proof of Lemma 31.67.
Equation (A.1203) is a genuine prediction with no adjustable content, and measured slopes for thin two-dimensional sections at small incidence do come within a few percent of \(2\pi\) per radian. Four qualifications are owed.
Thickness and camber. The circle of Lemma A.733 was concentric with the origin, which is what degenerated the profile to a segment. Displacing its centre gives thickness (a shift along the real axis) and camber (a shift along the imaginary axis), and the same argument then yields \(C_{\mathrm{L}}=2\pi\sin\left(\alpha-\alpha_{0}\right)\): camber moves the zero-lift angle \(\alpha_{0}\) and leaves the slope alone, which is why the number \(2\pi\) is so much more robust than anything else in aerofoil theory.
Finite span. A real wing sheds trailing vorticity, whose downwash reduces the effective incidence, and the slope falls with aspect ratio. This is a three-dimensional effect and nothing in a plane flow can see it.
The Kutta condition is a viscous fact. It has been imposed here, not derived. Its justification is that the boundary layer of Section 31.7.2 cannot negotiate the infinite adverse pressure gradient implied by an infinite edge velocity and separates there instead — Proposition 31.73 — and its consistency with the conservation of circulation Corollary 31.65 is secured by the starting vortex described at Remark 31.69. An inviscid theory that selects its solution by a viscous criterion is not a closed theory, and it should not be presented as one.
No stall. Equation (A.1203) rises without bound with \(\alpha\), which is false: beyond ten to fifteen degrees the flow separates from the upper surface and the lift collapses. The theory contains no mechanism for this whatever, because it contains no separation. The measured lift curve, its linear range and its breakdown are in Experiment: Fluid Flow and Turbulence.
The Joukowski Map, the Kutta Condition and the $2\pi$ Lift Slope discharges Proposition 31.70 of Fluid Dynamics, whose statement Equation (31.60) is proved here as Equation (A.1203), and completes Remark 31.69 by supplying the value of \(\Gamma\) that the chapter's discussion of the Kutta condition leaves open. It should be read against D'Alembert's Paradox for a Body of Arbitrary Shape: the circulation is the first of the two escapes from d'Alembert's theorem catalogued at Remark 31.31, and what this section computes is precisely how much of it a sharp trailing edge is worth. The drag is still exactly zero, here as there — Theorem 31.68 gives \(X=0\) — so the escape buys lift and nothing else. Kutta's own computation is [Kutta:1902]; the general circulation theorem is Joukowski's [Joukowski:1910].
The Kármán–Howarth Relation and the Four-Fifths Law
This appendix proves Equation (31.74) of Remark 31.86 in Fluid Dynamics:
in the inertial range of homogeneous isotropic turbulence, with no adjustable constant. The chapter calls this the one nontrivial statement about turbulence that is derived from the Navier–Stokes equations rather than assumed and tested, and sets it deliberately against the dimensional argument of Phenomenon 31.84, which is not derived. The whole weight of that claim rests on the pages below, and the section is written accordingly: everything is proved, the symmetry hypotheses are used explicitly at the four places where they are needed, and Remark A.750 states at the end exactly how much the word exact is buying. It is buying less than it sounds, and more than anything else in the subject.
Setting
Let \(\vect{v}(\vect{x},t)\) be a random velocity field on \(\R^{3}\) obeying the incompressible Navier–Stokes equations Equation (31.32) with kinematic viscosity \(\nu\) in \(\mathrm{m}^{2}/\mathrm{s}\) and a body force \(\vect{F}\) per unit mass, and let \(\avg{\,\cdot\,}\) denote the ensemble average. The turbulence is assumed
-
homogeneous: every joint moment is invariant under translations of \(\R^{3}\);
-
isotropic in the strong sense: every joint moment is invariant under the full orthogonal group \(\Ogrp(3)\), rotations and reflections alike;
-
statistically steady: every moment is independent of \(t\);
-
forced at large scales only: \(\vect{F}\) is a random field with the same symmetries whose correlation with \(\vect{v}\) varies on a single length \(L\), the integral scale.
Write \(\vect{v}:=\vect{v}(\vect{x},t)\), \(\vect{v}':=\vect{v}(\vect{x}+\vect{r},t)\), \(\delta\vect{v}:=\vect{v}'-\vect{v}\), \(r=\abs{\vect{r}}\), \(\vect{n}=\vect{r}/r\), and \(\delta u_{\parallel}:=\delta\vect{v}\cdot\vect{n}\) for the longitudinal increment. Define
and let
be the mean dissipation rate per unit mass, in \(\mathrm{W}/\mathrm{kg}\)\(=\)\(\mathrm{m}^{2}/\mathrm{s}^{3}\). Rests on Theorem 31.37 and Definition 11.4.
Hypothesis (2) is stronger than rotational invariance alone and the difference matters: a turbulence with a mean helicity is invariant under \(\SO(3)\) but not under reflection, and the pseudo-tensor terms that reflection invariance removes below would then survive.
Isotropic tensor functions of one vector
Everything in this section rests on one algebraic lemma, which is proved here rather than quoted because the whole reduction from a field theory to two ordinary differential equations is contained in it.
Let \(T\) be a Cartesian tensor field on \(\R^{3}\setminus\set{0}\), depending on the single vector \(\vect{r}\), and covariant under \(\Ogrp(3)\): \(T_{i\ldots}(\mathsf{O}\vect{r}) =\mathsf{O}_{ii'}\cdots T_{i'\ldots}(\vect{r})\) for every orthogonal \(\mathsf{O}\). Then, with \(\vect{n}=\vect{r}/r\),
with scalar functions of \(r\) alone. A constant tensor with these symmetries and no \(\vect{r}\) dependence is zero at odd rank and a multiple of \(\delta_{ij}\) at rank two. Rests on Theorem 5.128 and Definition 14.26.
Derives Lemma A.742. Fix \(r>0\). Covariance under the rotations that move \(\vect{r}\) shows that the components in a frame adapted to \(\vect{r}\) depend on \(r\) only, so it suffices to work at one point. Choose axes with \(\vect{n}=\hat{\vect{e}}_{3}\) and let Latin indices \(a,b,c\) run over the transverse values \(1,2\). The subgroup of \(\Ogrp(3)\) fixing \(\vect{r}\) is the group \(\Ogrp(2)\) acting on the transverse plane, and each block of components must be an invariant tensor of that group.
At rank one: \(T_{a}\) is an invariant vector of \(\Ogrp(2)\), hence zero (the element \(-\mathsf{I}\) of \(\Ogrp(2)\) reverses it), and \(T_{3}\) is free. This is Equation (A.1207).
At rank two, symmetric: \(T_{3a}\) is again an invariant transverse vector, hence zero; \(T_{ab}\) is an invariant symmetric two-dimensional tensor, hence \(\phi_{2}\delta_{ab}\), since the traceless part carries the spin-two representation of \(\SO(2)\), which has no invariant; and \(T_{33}=\phi_{1}\) is free. Writing \(\delta_{ab}=\delta_{ij}-n_{i}n_{j}\) in covariant form gives Equation (A.1208).
At rank three, symmetric in the first two indices, four blocks occur. \(T_{33a}\) is an invariant transverse vector: zero. \(T_{ab3}\) is an invariant symmetric transverse two-tensor: \(\mathcal{B}\,\delta_{ab}\). \(T_{3ab}=T_{a3b}\) is an invariant transverse two-tensor, not required symmetric, hence again \(\mathcal{C}\,\delta_{ab}\) — the antisymmetric \(\epsilon_{ab}\) is invariant under \(\SO(2)\) but changes sign under a reflection of the plane, and hypothesis (2) of Definition A.741 excludes it. \(T_{abc}\) is an invariant transverse tensor of odd rank, and \(-\mathsf{I}\in\SO(2)\) reverses it: zero. Assembling with \(\delta_{ab}=\delta_{ij}-n_{i}n_{j}\) and renaming the free coefficient of \(n_{i}n_{j}n_{k}\) gives Equation (A.1209).
For a constant tensor the same argument applies at every \(\vect{r}\) simultaneously, so the surviving coefficients must be independent of the direction \(\vect{n}\); at odd rank every term in Equations (A.1207) and (A.1209) carries an odd number of factors \(n\), and no such expression is direction independent unless it vanishes.
∎Write \(B_{ij}(\vect{r}):=\avg{v_{i}v'_{j}}\). Then there is a single scalar \(f\), with \(f(0)=1\), such that
and \(R(r)=B_{ii}=u^{2}\left(3f+rf'\right)\); moreover \(\avg{v_{i}v_{j}v'_{k}}\) is determined by a single scalar \(k\), namely \(\avg{v_{\parallel}^{2}v'_{\parallel}}=u^{3}k(r)\), and \(\avg{v_{i}v_{j}v_{k}}=0\). Rests on Lemma A.742 and Corollary 31.13.
Derives Corollary A.743. \(B_{ij}\) is symmetric under the simultaneous exchange \(i\leftrightarrow j\), \(\vect{r}\rightarrow-\vect{r}\), and by Equation (A.1208) it takes the displayed form with \(f=B_{\parallel\parallel}/u^{2}\) and \(g=B_{\perp\perp}/u^{2}\), the longitudinal and transverse correlations; \(f(0)=1\) by the definition of \(u^{2}\). Incompressibility acts through \(\pp B_{ij}/\pp r_{j} =\avg{v_{i}\,\pp'_{j}v'_{j}}=0\). Writing \(h:=\left(f-g\right)/r^{2}\) so that the first term is \(u^{2}h\,r_{i}r_{j}\),
using \(\pp_{j}r=r_{j}/r\) and \(\pp_{j}\left(r_{i}r_{j}\right)=4r_{i}\). Their sum vanishes for all \(\vect{r}\), so \(h'r+4h+g'/r=0\); substituting \(h\) and clearing \(r^{2}\) leaves \(rf'+2\left(f-g\right)=0\), which is the second relation in Equation (A.1210). Contracting the first relation gives \(R=u^{2}\left(f-g+3g\right)=u^{2}\left(f+2g\right) =u^{2}\left(3f+rf'\right)\).
For the triple correlation put \(B_{ij,k}:=\avg{v_{i}v_{j}v'_{k}}\); it is symmetric in \(i,j\), so Equation (A.1209) applies with coefficients \(\mathcal{A},\mathcal{B},\mathcal{C}\). Incompressibility at the far point gives \(\pp B_{ij,k}/\pp r_{k}=\avg{v_{i}v_{j}\pp'_{k}v'_{k}}=0\); carrying out the differentiation as above,
and since \(n_{i}n_{j}\) and \(\delta_{ij}\) are independent the two coefficients must vanish separately:
Two equations for three functions leave one free function, which may be taken to be \(u^{3}k:=B_{\parallel\parallel,\parallel} =\mathcal{A}+\mathcal{B}+2\mathcal{C}\). Finally \(\avg{v_{i}v_{j}v_{k}}\) is a constant isotropic tensor of rank three, so it vanishes by the last clause of Lemma A.742.
∎The pressure drops out
This is the step that makes the result exact rather than a closure, and it is the one a reader should doubt until it is shown, since a correlation involving the pressure is otherwise as intractable as anything in the subject.
Under Definition A.741, \(\avg{p\,v'_{i}}=0\) identically, and likewise \(\avg{\abs{\vect{v}}^{2}v'_{i}}=0\). Rests on Lemma A.742 and Corollary 31.13.
Derives Lemma A.744. Both are isotropic vector functions of \(\vect{r}\), so by Equation (A.1207) each equals \(\phi(r)n_{i}\) for some scalar \(\phi\). Both are divergence free in \(\vect{r}\): differentiating with respect to \(r_{i}\) is differentiating the far point, and
because the field is incompressible (Corollary 31.13) and the near point is held fixed. For a radial field \(\phi(r)n_{i}\) the divergence is \(r^{-2}\left(r^{2}\phi\right)'\), so \(r^{2}\phi\) is constant; the constant is zero because \(\phi\) is bounded at the origin, both quantities being finite one-point moments there. Hence \(\phi\equiv0\).
∎The two-point energy equation
Under Definition A.741, and without any closure hypothesis whatever,
and \(\mathcal{F}(0)=2\varepsilon\) in a steady state. Rests on Theorem 31.37 and Lemma A.744.
Derives Theorem A.745. By homogeneity every correlation depends on \(\vect{r}=\vect{x}'-\vect{x}\) alone, so \(\pp/\pp x_{i}=-\pp/\pp r_{i}\) and \(\pp/\pp x'_{i}=+\pp/\pp r_{i}\) when applied to one. Write the Navier–Stokes equation Equation (31.32) at \(\vect{x}\) and at \(\vect{x}'\), multiply the first by \(v'_{i}\) and the second by \(v_{i}\), add and average:
The pressure bracket is \(\pp_{i}\avg{p\,v'_{i}}-\pp_{i}\avg{v_{i}p'}\) after moving the derivatives outside the averages, and both terms vanish by Lemma A.744 — the second is the first with \(\vect{r}\) reversed. The viscous bracket is \(2\nu\nabla_{r}^{2}R\), since each Laplacian acts on one factor and becomes \(\nabla_{r}^{2}\) on the correlation.
For the nonlinear terms use incompressibility to write \(v_{j}\pp_{j}v_{i}=\pp_{j}\left(v_{j}v_{i}\right)\); then
so the nonlinear contribution is \(\pp_{r_{j}}\bigl[\avg{v_{j}v_{i}v'_{i}} -\avg{v_{i}v'_{j}v'_{i}}\bigr]\). It remains to identify the bracket. Expand
The terms \(\avg{\abs{\vect{v}'}^{2}v'_{j}}\) and \(\avg{\abs{\vect{v}}^{2}v_{j}}\) are equal by homogeneity and cancel; \(\avg{\abs{\vect{v}}^{2}v'_{j}}\) and \(\avg{\abs{\vect{v}'}^{2}v_{j}}\) vanish by Lemma A.744; and what is left is
which is twice the bracket. This gives Equation (A.1212).
At \(\vect{r}=\vect{0}\): the divergence term vanishes, because \(\avg{\abs{\delta\vect{v}}^{2}\delta\vect{v}}=O(r^{3})\) there; the viscous term is \(2\nu\nabla_{r}^{2}R|_{0}=2\nu\avg{v_{i}\nabla^{2}v_{i}} =-2\nu\avg{\pp_{j}v_{i}\pp_{j}v_{i}}=-2\varepsilon\) by Equation (A.1206) and homogeneity; and \(\pp_{t}R(0)=0\) in a steady state. Hence \(\mathcal{F}(0)=2\varepsilon\), which is the statement that in a steady state the force injects energy at exactly the rate viscosity destroys it.
∎For unforced (decaying) homogeneous isotropic turbulence, Equation (A.1212) is equivalent to
with \(f\) and \(k\) the scalars of Corollary A.743. Rests on Theorem A.745 and Corollary A.743.
Derives Proposition A.746. Write \(T(r)\) for the scalar in \(\avg{\abs{\delta\vect{v}}^{2}\delta v_{j}}=T(r)n_{j}\), legitimate by Equation (A.1207); then \(\vect{\nabla}_{r}\cdot\avg{\abs{\delta\vect{v}}^{2}\delta\vect{v}} =r^{-2}\left(r^{2}T\right)'\) and \(\nabla_{r}^{2}R=r^{-2}\left(r^{2}R'\right)'\), so Equation (A.1212) unforced reads
By Corollary A.743, \(R=u^{2}\left(3f+rf'\right)=u^{2}\left(r^{3}f\right)'/r^{2}\), and Structure functions below gives \(T=u^{3}\left(8k+2rk'\right)\). Apply the operator \(\mathcal{L}[\Phi]:=r^{-2}\pp_{r}\left(r^{3}\Phi\right)\) to each term of Equation (A.1213). On the left, \(\mathcal{L}\bigl[\pp_{t}(u^{2}f)\bigr]=\pp_{t}R\). On the right,
the last step by direct differentiation of \(r^{2}T=u^{3}\left(8r^{2}k+2r^{3}k'\right)\); and
since \(r^{2}R'=u^{2}\left(4r^{2}f'+r^{3}f''\right)\). So Equation (A.1213) maps term by term onto Equation (A.1214); and \(\mathcal{L}\) is injective on functions regular at the origin, since \(\mathcal{L}[\Phi]=0\) forces \(r^{3}\Phi\) constant, hence \(\Phi\propto r^{-3}\), hence \(\Phi=0\).
∎Equation (A.1213) is the Kármán–Howarth relation, obtained by von Kármán and Howarth in 1938 (Proceedings of the Royal Society of London A 164, 192–215). That paper is not in the bibliography of this treatise, and the statement is not being taken from it: everything above is derived here from Equation (31.32) and Definition A.741, so the attribution is historical and nothing rests on it. The equation is exact and it does not close: one equation relates the two unknown functions \(f\) and \(k\), and nothing determines \(k\). What Kolmogorov saw is that in one limit the undetermined function drops out.
Structure functions
With \(S_{ijk}:=\avg{\delta v_{i}\delta v_{j}\delta v_{k}}\) and \(T(r)\) as above,
Rests on Corollary A.743 and Lemma A.742.
Derives Lemma A.747. Expand \(S_{ijk}\) in the eight products of \(v\) and \(v'\). The two single-point terms \(\avg{v_{i}v_{j}v_{k}}\) and \(\avg{v'_{i}v'_{j}v'_{k}}\) vanish by Corollary A.743. Of the remaining six, three have two factors at \(\vect{x}\) and three have two at \(\vect{x}+\vect{r}\); using homogeneity to write the latter as \(B\) evaluated at \(-\vect{r}\), and the fact that \(B_{ij,k}\) is odd — every term of Equation (A.1209) carries an odd number of factors \(\vect{n}\) — the two groups add rather than cancel:
the second equality by adding the three copies of Equation (A.1209) with indices permuted. Write \(W:=\mathcal{A}+\mathcal{B}+2\mathcal{C}=u^{3}k\) and \(V:=\mathcal{B}+2\mathcal{C}\). Contracting with \(n_{i}n_{j}n_{k}\) gives \(S_{3}=6\mathcal{A}+6V=6W\), the first relation in Equation (A.1215); contracting on \(i=j\) and then with \(n_{k}\) gives \(T=S_{iik}n_{k}=6\mathcal{A}+10V=S_{3}+4V\).
It remains to express \(V\) through \(W\). Adding the two relations Equation (A.1211) and using \(\mathcal{A}+\mathcal{B}=W-2\mathcal{C}\),
and substituting \(\mathcal{B}=V-2\mathcal{C}\) into the second of Equation (A.1211) gives \(V'+2V/r=2\mathcal{C}'+2\mathcal{C}/r\). With the expression just obtained for \(\mathcal{C}\), the right-hand side is \(\left[rW''+4W'+2W/r\right]/2\), so
the last step because \(\left(r^{3}W'\right)'=3r^{2}W'+r^{3}W''\) and \(\left(r^{2}W\right)'=2rW+r^{2}W'\). Integrating and discarding the \(r^{-2}\) homogeneous solution by regularity at the origin, \(V=\tfrac{1}{2}\left(rW\right)'\). Hence \(T=S_{3}+2\left(rW\right)'=S_{3}+\tfrac{1}{3}\left(rS_{3}\right)'\), and in terms of \(k\), \(T=6u^{3}k+2u^{3}\left(k+rk'\right) =u^{3}\left(8k+2rk'\right)\).
∎\(D(r):=\avg{\abs{\delta\vect{v}}^{2}}=2\left[R(0)-R(r)\right] =3S_{2}+rS_{2}'\). Rests on Corollary A.743.
Derives Lemma A.748. \(\avg{\abs{\delta\vect{v}}^{2}} =\avg{\abs{\vect{v}}^{2}}+\avg{\abs{\vect{v}'}^{2}} -2\avg{\vect{v}\cdot\vect{v}'}\), and the first two are \(R(0)=3u^{2}\) each by homogeneity. For the second equality, Equation (A.1210) gives \(S_{2}=2u^{2}\left(1-f\right)\) and the transverse structure function \(\avg{\left(\delta u_{\perp}\right)^{2}}=2u^{2}\left(1-g\right) =S_{2}+\tfrac{1}{2}rS_{2}'\), using \(g=f+\tfrac{1}{2}rf'\); summing one longitudinal and two transverse contributions gives \(D=3S_{2}+rS_{2}'\).
∎The inertial range and the four-fifths law
Under Definition A.741, for every separation \(r\ll L\),
and consequently, in the inertial range \(\eta\ll r\ll L\) where the viscous term is negligible,
which is Equation (31.74). Here \(\eta=\left(\nu^{3}/\varepsilon\right)^{1/4}\) is the Kolmogorov length Equation (31.72). Rests on Theorem A.745, Lemma A.747 and Lemma A.748.
Derives Theorem A.749. In a steady state the left-hand side of Equation (A.1212) vanishes, and in radial form
Multiply by \(r^{2}\) and integrate from \(0\) to \(r\). The first term gives \(\tfrac{1}{2}r^{2}T\), the boundary term at the origin vanishing because \(T=O(r^{3})\). The second gives \(2\nu r^{2}R'(r)\), likewise. For the third, \(\mathcal{F}\) varies on the scale \(L\) by hypothesis (4), so for \(r\ll L\) it may be replaced by \(\mathcal{F}(0)=2\varepsilon\), giving \(\tfrac{2}{3}\varepsilon r^{3}\). Hence
By Lemma A.748, \(R=R(0)-\tfrac{1}{2}D\), so \(R'=-\tfrac{1}{2}D'\) and
This is already exact and already the whole content; the rest is translation. Substituting \(T=\left(r^{4}S_{3}\right)'/3r^{3}\) — which is Equation (A.1215) rewritten, since \(\left(r^{4}S_{3}\right)'=r^{3}\left(4S_{3}+rS_{3}'\right) =3r^{3}\left[S_{3}+\tfrac{1}{3}\left(rS_{3}\right)'\right]\) — and \(D=\left(r^{3}S_{2}\right)'/r^{2}\) from Lemma A.748, one verifies that Equation (A.1216) solves Equation (A.1218): with \(S_{3}=6\nu S_{2}'-\tfrac{4}{5}\varepsilon r\),
because \(D'=\left(3S_{2}+rS_{2}'\right)'=4S_{2}'+rS_{2}''\). Since Equation (A.1218) determines \(S_{3}\) given \(S_{2}\) — it is a first-order linear equation with the regularity condition \(S_{3}=O(r^{3})\) at the origin fixing the integration constant — Equation (A.1216) is the solution.
For Equation (A.1217): the viscous term \(6\nu S_{2}'\) is of order \(\nu S_{2}/r\), and in the inertial range \(S_{2}\sim\left(\varepsilon r\right)^{2/3}\), so the ratio of the viscous to the inertial term is of order \(\nu\varepsilon^{2/3}r^{-4/3}/\varepsilon r =\left(\eta/r\right)^{4/3}\), which is negligible for \(r\gg\eta\); and \(r\ll L\) was used already. Dropping it leaves Equation (A.1217).
∎The direction of the inequality is the physics. \(S_{3}<0\) says that the odd moment of the longitudinal increment is negative: the fluid at \(\vect{x}+\vect{r}\) is, on average, moving towards the fluid at \(\vect{x}\) more strongly than away from it, in the cubic sense. That asymmetry is the cascade, and Equation (A.1217) measures it against \(\varepsilon\) with no free constant. Reversed sign would be an inverse cascade, which is what two-dimensional turbulence does and three-dimensional turbulence does not.
The chapter calls Equation (31.74) the one exact nontrivial result about turbulence, and the reader is owed the bill.
Homogeneity was used at every step where \(\pp/\pp x_{i}=-\pp/\pp r_{i}\) was applied, and in Lemma A.748. No real flow is homogeneous — there is always a boundary and always a forcing region — so the statement is about an idealized ensemble that laboratory turbulence approximates locally.
Isotropy, including reflection, was used three times: in Lemma A.742 to reduce tensor fields to scalars, in Lemma A.744 to kill the pressure term, and in Lemma A.747. Anisotropy at the large scales survives into the inertial range more persistently than K41 supposes, and the four-fifths law is correspondingly harder to observe cleanly than its exactness suggests.
Stationarity set \(\pp_{t}R=0\). For decaying turbulence the same calculation retains a term in \(\pp_{t}\) and Equation (A.1217) acquires a correction that is small only when the decay time exceeds the eddy turnover time at separation \(r\).
The ordering of the limits. Two neglects were made, \(r\ll L\) and \(r\gg\eta\), and they are compatible only when \(L/\eta\gg1\), that is at large Reynolds number by Equation (31.73). The four-fifths law is a statement about an asymptotic regime, and the inertial range at attainable Reynolds numbers is short.
Finite dissipation as \(\nu\rightarrow0\). The whole argument treats \(\varepsilon\) as a fixed quantity while the viscous term is dropped. That \(\varepsilon\) tends to a non-zero limit as \(\nu\rightarrow0\) — the dissipation anomaly — is an assumption, not a theorem; it is strongly supported by measurement and it is what makes the cascade picture consistent, but it is precisely the kind of statement whose proof would require knowing that the equations have the solutions Remark 31.48 says are not known to exist.
What survives all five qualifications is still remarkable, and the comparison with Phenomenon 31.84 is the point of stating it: the exponent \(-5/3\) follows from hypotheses that measurement has falsified in detail — the intermittency corrections are real — while the coefficient \(-4/5\) follows from the equations themselves, and measurement has found no correction to it. Remark 31.87 sets the whole position out, and Experiment: Fluid Flow and Turbulence collects the measurements a future theory will be asked to reproduce.
Equation (A.1217) is Kolmogorov's, and the paper that contains it is the companion of [Kolmogorov:1941] — the dissipation paper, Doklady Akademii Nauk SSSR 32 (1941) 16–18, recorded in the bibliographic note of that entry. The entry [Kolmogorov:1941] itself is the local-structure paper, which states the similarity hypotheses used for Equation (31.71). Remark 31.86 already warns that the K41 attributions must be read carefully, the spectral form being Obukhov's [Obukhov:1941]; the same care applies within Kolmogorov's own output, and the two 1941 papers are not interchangeable.
The Kármán–Howarth Relation and the Four-Fifths Law discharges Equation (31.74) inside Remark 31.86 of Fluid Dynamics, and with it the chapter's claim that exactly one nontrivial consequence of Equation (31.32) about turbulence is derived rather than assumed. The reader returning there should carry back two things this section supplies and the chapter cannot: that the Kármán–Howarth relation Equation (A.1213) is exact but does not close, one equation in two unknown functions, which is the closure problem of Remark 31.83 in its sharpest form; and that the four-fifths law survives the non-closure only because the undetermined function is exactly the one the inertial-range limit eliminates. Everything else in Section 31.9 is dimensional analysis and measurement, as Remark 31.87 says.
The Korteweg–de Vries Equation from the Free-Surface Problem
This appendix derives Equation (31.81) of Fluid Dynamics from the exact equations of an ideal free-surface layer. The chapter, in the derivation of Phenomenon 31.96, takes the Korteweg–de Vries equation as given, integrates it for a travelling wave, and obtains the profile Equation (31.82) with Russell's speed Equation (31.80); what it owes, and what is supplied here, is the equation itself [Korteweg:1895]. The travelling-wave integration is not repeated below — Equation (31.82) stands where it is.
The derivation is an expansion in two small parameters, and the single most important thing about it is that the relation between them is chosen, not deduced. That choice is stated up front in Two small parameters, and a choice and revisited at Remark A.753; a reader who takes it for a consequence has misunderstood the whole calculation.
The exact free-surface problem
Take an ideal, incompressible, irrotational layer of undisturbed depth \(h\) over a flat rigid bottom at \(y=0\), with the free surface at \(y=h+\eta(x,t)\) and gravity \(g\) acting downwards. Surface tension is set to zero here and restored in Remark A.758. By Proposition 31.28 the velocity is \(\vect{v}=\vect{\nabla}\phi\) with
the second being impermeability of the bottom. At the surface two conditions hold, both nonlinear and both imposed on a boundary whose position is itself unknown:
Equation (A.1220) says the surface is material — a particle on it stays on it — and Equation (A.1221) is the unsteady Bernoulli equation Equation (31.21) evaluated at the surface, where the pressure equals the constant atmospheric value, absorbed into \(\phi\).
Two small parameters, and a choice
Let \(a\) be a typical wave amplitude and \(\ell\) a typical horizontal length, and set
Nondimensionalize by
the scale for \(\phi\) being the one that makes Equation (A.1221) balance at leading order. Substituting and dropping tildes, Equations (A.1219), (A.1220) and (A.1221) become exactly
No approximation has yet been made: the three coefficients \(\epsilon\), \(1/\delta^{2}\) and \(\epsilon/\delta^{2}\) are what the substitution produces, using \(ga/c_{0}^{2}=a/h=\epsilon\) and \(\ell^{2}/h^{2}=1/\delta^{2}\).
Nothing so far relates \(\epsilon\) to \(\delta^{2}\), and different relations give different equations. Taking \(\epsilon\ll\delta^{2}\) and keeping only \(\delta^{2}\) gives the linear dispersive equation whose relation is Equation (A.1239) below; taking \(\delta^{2}\ll\epsilon\) and keeping only \(\epsilon\) gives the nonlinear non-dispersive equation, whose solutions steepen and break in finite time. The Korteweg–de Vries equation is the distinguished limit
in which the two effects enter at the same order and can balance. This is the entire physical content of the solitary wave — steepening against spreading, as Phenomenon 31.96 says — and it is a hypothesis about the class of initial data being described, not a consequence of the equations. Data not satisfying Equation (A.1227) are described by one of the other two limits and do not produce solitary waves. The dimensionless group in Equation (A.1227) is the Ursell number; the solitary wave of Equation (31.82) has \(\ell^{2}=4h^{3}/3a\), i.e. Ursell number \(4/3\) exactly, which is the self-consistency of the whole construction.
Both parameters are treated as first order below: terms of order \(\epsilon\) and \(\delta^{2}\) are retained, and terms of order \(\epsilon^{2}\), \(\epsilon\delta^{2}\) and \(\delta^{4}\) are discarded.
Eliminating the depth
Every solution of Equation (A.1224) analytic in \(y\) is
which for the retained orders reads \(\phi=F-\tfrac{1}{2}\delta^{2}y^{2}\pp_{x}^{2}F +\tfrac{1}{24}\delta^{4}y^{4}\pp_{x}^{4}F+O(\delta^{6})\). Rests on Equation (A.1224) and Theorem 7.38.
Derives Lemma A.754. Write \(\phi=\sum_{m\ge0}y^{m}\phi_{m}(x,t)/m!\). The bottom condition kills \(\phi_{1}\), and Equation (A.1224) gives the recursion \(\phi_{m+2}=-\delta^{2}\pp_{x}^{2}\phi_{m}\); hence every odd coefficient vanishes and \(\phi_{2n}=\left(-\delta^{2}\right)^{n}\pp_{x}^{2n}\phi_{0}\), which is Equation (A.1228) with \(\phi_{0}=F\).
∎The lemma is exact and is what makes the problem one-dimensional: the whole vertical structure is slaved to the single function \(F\), and the long-wave parameter \(\delta^{2}\) is precisely the bookkeeping parameter of that slaving. Introduce the horizontal velocity at the bottom,
and read off from Equation (A.1228)
To first order in \(\epsilon\) and in \(\delta^{2}\), the free-surface conditions Equations (A.1225) and (A.1226) reduce to the closed system
Rests on Lemma A.754 and Equation (A.1225).
Derives Proposition A.755. Evaluate Equation (A.1230) at \(y=1+\epsilon\eta\) and substitute in Equation (A.1225). The right-hand side is
while the left-hand side is \(\pp_{t}\eta+\epsilon u\,\pp_{x}\eta\). Collecting and using \(\epsilon u\,\pp_{x}\eta+\epsilon\eta\,\pp_{x}u =\epsilon\,\pp_{x}\left(\eta u\right)\) gives Equation (A.1231).
In Equation (A.1226) the term \(\epsilon\left(\pp_{y}\phi\right)^{2}/2\delta^{2}\) is \(O\left(\epsilon\delta^{2}\right)\) by Equation (A.1230) and is discarded, while \(\epsilon\left(\pp_{x}\phi\right)^{2}/2=\epsilon u^{2}/2+O(\epsilon \delta^{2})\). Evaluating \(\pp_{t}\phi\) at \(y=1+\epsilon\eta\) to the retained order,
and differentiating once in \(x\), with \(\pp_{x}F=u\), gives Equation (A.1232).
∎Unidirectional reduction
At \(\epsilon=\delta^{2}=0\) the pair Equations (A.1231) and (A.1232) is \(\pp_{t}\eta+\pp_{x}u=0\), \(\pp_{t}u+\pp_{x}\eta=0\), whence \(\pp_{t}^{2}\eta=\pp_{x}^{2}\eta\): two waves, one running each way, with \(u=\eta\) on the right-running one. Restrict to that branch and correct it.
A right-running solution of Equations (A.1231) and (A.1232), correct to first order in \(\epsilon\) and \(\delta^{2}\), has
and \(\eta\) obeys
In dimensional variables this is
which is Equation (31.81). Rests on Proposition A.755 and Equation (A.1227).
Derives Theorem A.756. Write \(u=\eta+\epsilon A+\delta^{2}B\) with \(A\) and \(B\) functions of \(\eta\) and its \(x\) derivatives, to be determined. Substituting into Equations (A.1231) and (A.1232) and discarding second-order terms,
Inside a term already carrying a factor \(\epsilon\) or \(\delta^{2}\) the leading-order relation \(\pp_{t}=-\pp_{x}\) may be used, since the error is of second order; so \(\pp_{t}A=-\pp_{x}A\), \(\pp_{t}B=-\pp_{x}B\) and \(\pp_{x}^{2}\pp_{t}\eta=-\pp_{x}^{3}\eta\), and Equation (A.1237) becomes
The two equations must be consistent: a single function \(\eta\) cannot obey two different evolution laws. Subtracting them removes the leading part and leaves
and since \(\epsilon\) and \(\delta^{2}\) are independent parameters each bracket must vanish. Integrating, and taking the constants of integration to be zero so that \(u\) vanishes with \(\eta\),
which is Equation (A.1233). Adding the two equations instead, the terms in \(A\) and \(B\) cancel identically and what survives is
which is Equation (A.1234).
For Equation (A.1235), undo Equation (A.1223): multiplying Equation (A.1234) by \(ac_{0}/\ell\) and using \(\pp_{\tilde{t}}\tilde{\eta} =\left(\ell/ac_{0}\right)\pp_{t}\eta\), \(\pp_{\tilde{x}}\tilde{\eta}=\left(\ell/a\right)\pp_{x}\eta\), \(\epsilon\tilde{\eta}\pp_{\tilde{x}}\tilde{\eta} =\left(\ell/ah\right)\eta\,\pp_{x}\eta\) and \(\delta^{2}\pp_{\tilde{x}}^{3}\tilde{\eta} =\left(h^{2}\ell/a\right)\pp_{x}^{3}\eta\), every occurrence of \(a\) and \(\ell\) cancels and Equation (A.1235) results — as it must, since \(a\) and \(\ell\) were bookkeeping devices and no physical statement may depend on them.
∎The step just carried out is what is usually called the removal of secular terms, and it is often presented as an application of the method of multiple scales. Part II carries no such method — its asymptotic material is the WKB expansion of Section 9.6 and the matched expansions named at Definition 9.86 — and none is needed. The reason is visible in the proof: because Equations (A.1231) and (A.1232) are two equations for the same pair of unknowns, consistency alone fixes \(A\) and \(B\), and the freedom that a general theory would have to organize into slow and fast variables has already been used up. Had the reduction been attempted on a single equation the situation would be different, and the apparatus would be owed.
Two checks
The chapter asserts, without demonstration, that the two nontrivial terms of Equation (31.81) are exactly the steepening and the dispersion. Both halves can be verified against results already proved.
Drop the third derivative.
Equation (A.1235) becomes \(\pp_{t}\eta+c_{0}\left(1+3\eta/2h\right)\pp_{x}\eta=0\), a quasilinear first-order equation whose solution is constant along the characteristics \(\dd x/\dd t=c_{0}\left(1+3\eta/2h\right)\). Higher parts of the profile therefore travel faster and overtake lower ones, the front steepens, and after a finite time the solution becomes multivalued: this is the simple-wave steepening of Remark 31.90, which in a compressible gas ends in a shock. Note that the amplitude-dependent speed \(c_{0}\left(1+3\eta/2h\right)\) is consistent with Russell's measured \(c^{2}=g\left(h+a\right)\) only through the factor \(3/2\) that the expansion supplies and dimensional analysis cannot.
Drop the nonlinear term.
Setting \(\eta\propto\ee^{\ii\left(kx-\omega t\right)}\) in Equation (A.1235) without the \(\eta\,\pp_{x}\eta\) term gives \(-\ii\omega+\ii c_{0}k+\left(c_{0}h^{2}/6\right)\left(\ii k\right)^{3} =0\), that is
Expanding the exact gravity-wave relation Equation (31.78) at \(\gamma=0\) with \(\tanh\left(kh\right)=kh-\tfrac{1}{3}\left(kh\right)^{3}+\cdots\),
which is Equation (A.1239) exactly, coefficient included. The third-derivative term of Equation (31.81) is therefore the first dispersive correction to the shallow-water speed and nothing else, as the chapter says.
Four idealizations were made and each has an observable consequence.
No viscosity. Equation (A.1219) is the ideal-fluid problem, so the solitary wave of Equation (31.82) propagates without loss. Real waves in a channel decay, chiefly through the boundary layer on the bed and walls (Section 31.7.2), and Russell's own observation ended when his wave died away.
No surface tension. Restoring it means replacing the constant surface pressure in Equation (A.1221) by the Young–Laplace jump Equation (31.18), and the expansion then delivers the same equation with the dispersive coefficient \(\left(c_{0}h^{2}/2\right)\left(\tfrac{1}{3}-\tau\right)\), \(\tau:=\gamma/\rho gh^{2}\) the Bond parameter. The check above confirms the coefficient independently: expanding Equation (31.78) with \(\gamma\) retained gives \(\omega=c_{0}k\bigl[1-\tfrac{1}{2}\left(kh\right)^{2} \left(\tfrac{1}{3}-\tau\right)\bigr]\). The sign of the dispersion therefore reverses at \(\tau=1/3\), that is for \(h<\sqrt{3\gamma/\rho g}\), about \(4.7\,\mathrm{mm}\) for clean water with \(\gamma=7.3\times 10^{-2}\,\mathrm{N}/\mathrm{m}\) and \(\rho=998\,\mathrm{kg}/\mathrm{m}^{3}\). In a layer thinner than that there are no elevation solitary waves, and Equation (31.82) does not apply — a limitation invisible in the chapter's statement.
Irrotational and one-directional. The reduction of Unidirectional reduction discards the left-running wave entirely. An initial disturbance at rest splits into two, and Equation (A.1235) describes one of them.
Small amplitude. The expansion is in \(\epsilon=a/h\) with \(\epsilon=O\left(\delta^{2}\right)\), so it is valid only for \(a\ll h\) — Russell's regime, as the chapter notes in the derivation of Phenomenon 31.96. Solitary waves of large amplitude exist but are not described by Equation (31.81): they have a limiting height of about \(0.8h\) and a corner at the crest, neither of which this equation knows about.
The Korteweg–de Vries Equation from the Free-Surface Problem discharges Equation (31.81) in the derivation of Phenomenon 31.96 in Fluid Dynamics, which takes the equation as given and proceeds from it to the profile Equation (31.82) and to Russell's speed Equation (31.80) [Russell:1845]. The reader returning there now knows what the two nontrivial terms are and where each comes from — the first from the \(\epsilon\) ordering of the surface conditions, the second from the \(\delta^{2}\) ordering of the depth expansion — and, more importantly, that their coexistence was imposed by the choice Equation (A.1227) rather than discovered. The reading of Equation (31.82) as a soliton, with the elastic collisions and the infinite hierarchy of conserved quantities that follow, belongs to Nonlinear Dynamics and Chaos; nothing of that is visible in the derivation above, which is why Korteweg and de Vries did not see it [Korteweg:1895].
The Lorenz System from Boussinesq Convection
This appendix proves Definition 32.62 of Nonlinear Dynamics and Chaos: that Equation (32.36), with the three parameters Equation (32.37), is what the equations of a convecting fluid layer become when they are truncated to one velocity mode and two temperature modes. The point of writing it out is stated in Remark 32.2: the Lorenz system is almost always presented as a dimensionless curiosity, and it is only by carrying the reduction through in SI variables that \(\sigma\), \(r\) and \(b\) can be read back as ratios of quantities a laboratory measures. Along the way the critical Rayleigh number \(\mathrm{Ra}_{\mathrm{c}}=27\pi^{4}/4\) and the critical wavenumber \(a^{2}=\tfrac{1}{2}\), both quoted in Definition 32.62 and in Equation (31.68), are derived.
Nothing is imported from outside the treatise. The starting equations are those of Fluid Dynamics — the incompressible Navier–Stokes equations Equation (31.32) and the advection–diffusion equation for heat — specialised by the Boussinesq approximation, which is that chapter's business and is discussed at Phenomenon 31.78; the linear-stability calculation and its comparison with measurement are Section 36.4. The attribution of the mode set is to Saltzman, whose 1962 study of finite convective amplitude in the Journal of the Atmospheric Sciences supplied the expansion Lorenz truncated; that paper has no entry in this bibliography, so the attribution is made in words, and the truncation and its consequences are Lorenz's [Lorenz:1963].
The layer, in SI units
A horizontal layer of fluid occupies \(0\le z\le d\), with \(z\) measured upwards against gravity \(\vect{g}=-g\hat{\vect{z}}\). The lower surface is held at temperature \(T_{0}+\Delta T\) and the upper at \(T_{0}\), so that \(\Delta T>0\) means heating from below. The fluid is characterised by four measured properties, all functions of the working substance and of its mean state:
the kinematic viscosity, the thermal diffusivity, the coefficient of thermal expansion and the reference density; \(d\) is in \(\mathrm{m}\), \(\Delta T\) in \(\mathrm{K}\) and \(g\) in \(\mathrm{m}/\mathrm{s}^{2}\). Rests on Theorem 31.37 and Definition 31.51.
Under the Boussinesq approximation — the density is treated as the constant \(\rho_{0}\) everywhere except in the buoyancy force, where it is \(\rho=\rho_{0}\left[1-\alpha_{T}\left(T-T_{0}\right)\right]\), and the flow is treated as incompressible — the motion of the layer of Definition A.760 obeys
with \(p'\) the departure of the pressure from its hydrostatic value. The state of rest, \(\vect{u}=\vect{0}\), is a solution with the conducting temperature profile
Rests on Definition A.760, Theorem 31.37 and Phenomenon 31.78.
Derives Proposition A.761. Equation (A.1241) is Equation (31.32) of Theorem 31.37 with the body force \(\vect{g}\) retained and the density in that force alone allowed to vary. Writing \(\rho\vect{g}=-\rho_{0}g\hat{\vect{z}} +\rho_{0}g\alpha_{T}(T-T_{0})\hat{\vect{z}}\) and absorbing the first, constant, piece into the pressure — which is what \(p'\) means — leaves the buoyancy term displayed. Equation (A.1242) is the advection–diffusion equation for the temperature of an incompressible fluid, viscous heating being of order \(\nu U^{2}/(\kappa\Delta T)\) relative to the terms kept and negligible in every laboratory convection experiment. For the conducting state, put \(\vect{u}=\vect{0}\) and seek a steady \(T\): Equation (A.1242) becomes \(\nabla^{2}T=0\), whose solution with the imposed boundary values is the linear profile Equation (A.1243), and Equation (A.1241) is then satisfied by a \(p'\) depending on \(z\) alone.
∎The Boussinesq approximation is an approximation and not a limit theorem: it is justified when \(\alpha_{T}\Delta T\ll1\) and when the layer is thin against the density scale height, and the honest statement of its domain belongs with the convection material of Fluid Dynamics, not here. This is an internal dependency of Part III — Classical Mechanics on itself rather than a debt owed by Part II — Mathematical Methods, and it is the only thing in this section taken from elsewhere. For air at room temperature \(\alpha_{T}\approx3.4\times 10^{-3}\,/\mathrm{K}\), so a layer driven by \(\Delta T=10\,\mathrm{K}\) has \(\alpha_{T}\Delta T=0.034\): the approximation is good to a few per cent, which is the accuracy at which Section 36.4 tests the threshold.
Rolls, the stream function, and the elimination of the pressure
Attention is restricted to motions independent of the horizontal coordinate \(y\) — rolls with axes along \(y\). The velocity is then written in terms of a stream function \(\psi(x,z,t)\) of SI dimension \(\mathrm{m}^{2}/\mathrm{s}\),
and the temperature by its departure from the conducting profile,
For any pair of functions the Jacobian is written
so that \(\left(\vect{u}\cdot\vect{\nabla}\right)h=J(\psi,h)\) for the field Equation (A.1244). Rests on Proposition A.761 and Equation (31.11).
Definition A.763 restricts attention to solutions of the three-dimensional equations Equation (A.1241) and Equation (A.1242) that happen not to depend on one of the three spatial coordinates. It is a symmetry ansatz inside the observed \(3+1\) spacetime, exactly like the axisymmetric ansatz used for Stokes flow, and it is not a model of a two-dimensional universe; the fluid, the layer and the rolls are all three-dimensional objects. The ansatz is a genuine restriction — real convection at large \(\mathrm{Ra}\) is not roll-like — and that restriction, not the dimensionality of space, is what has to be watched: it is one of the approximations listed in Remark A.771.
For motions of the form Equation (A.1244), the system Equation (A.1241)–Equation (A.1242) is equivalent to the pair
in which the pressure has disappeared and \(\nabla^{2}\), \(\nabla^{4}\) act in the \((x,z)\) plane. Rests on Definition A.763 and Proposition A.761.
Derives Proposition A.765. Incompressibility. \(\vect{\nabla}\cdot\vect{u} =-\pp_{x}\pp_{z}\psi+\pp_{z}\pp_{x}\psi=0\) identically, so Equation (A.1244) solves the constraint of Equation (A.1242) and no Lagrange multiplier is needed.
Vorticity. For a \(y\)-independent flow the vorticity has only a \(y\) component,
Taking the \(y\) component of the curl of Equation (A.1241) annihilates the pressure gradient, since the curl of a gradient vanishes; the buoyancy term \(g\alpha_{T}\theta\hat{\vect{z}}\) — the \(z\)-dependent part of \(T-T_{0}\) having been absorbed into \(p'\) along with the hydrostatic piece — contributes \(-g\alpha_{T}\pp_{x}\theta\); and the advective term contributes \(\left(\vect{u}\cdot\vect{\nabla}\right)\omega_{y}=J(\psi,\omega_{y})\), the vortex-stretching term \(\left(\vect{\omega}\cdot\vect{\nabla} \right)\vect{u}\) vanishing because \(\vect{\omega}\) points along \(y\) and nothing depends on \(y\). Hence
and substituting Equation (A.1249) and multiplying through by \(-1\) gives Equation (A.1247).
Temperature. Substituting \(T=T_{\mathrm{c}}+\theta\) into Equation (A.1242), and using \(\nabla^{2}T_{\mathrm{c}}=0\) because Equation (A.1243) is linear in \(z\),
and \(u_{z}=\pp_{x}\psi\) with \(\dd T_{\mathrm{c}}/\dd z=-\Delta T/d\) turns the third term into \(-\left(\Delta T/d\right)\pp_{x}\psi\), which moved to the right is Equation (A.1248). The physical content of that term is the whole instability: fluid rising through the mean gradient (\(\pp_{x}\psi>0\)) arrives warmer than its surroundings, which increases \(\theta\), which by Equation (A.1247) drives more rising motion.
∎The scaling, and the two dimensionless groups
Scale lengths by \(d\), time by \(d^{2}/\kappa\), the stream function by \(\kappa\) and the temperature departure by \(\nu\kappa/\left(g\alpha_{T}d^{3}\right)\),
Then Equation (A.1247)–Equation (A.1248) become, with the tildes dropped,
and exactly two dimensionless groups survive, the Prandtl and Rayleigh numbers
which are Equation (32.37) and Equation (32.38). Rests on Proposition A.765 and Definition 31.51.
Derives Proposition A.766. Under Equation (A.1252), \(\nabla^{2}=d^{-2}\tilde{\nabla}^{2}\), \(\pp_{t}=\left(\kappa/d^{2}\right)\pp_{\tilde{t}}\) and \(J=d^{-2}\tilde{J}\) acting on the scaled fields. The four terms of Equation (A.1247) then carry the prefactors
so dividing by \(\kappa^{2}/d^{4}\) leaves the factor \(\nu/\kappa=\sigma\) on the right-hand pair and gives Equation (A.1253). The four terms of Equation (A.1248) carry
and dividing by the first leaves the third with the factor
which is Equation (A.1254).
The dimension check. It is worth carrying out in full, because it is what allows \(\sigma\) and \(r\) to be read back as measurements:
and \(\sigma=\nu/\kappa\) is a ratio of two quantities in \(\mathrm{m}^{2}/\mathrm{s}\), so both are pure numbers, as Definition 31.51 requires. That no third group appears is the substance of the reduction: the four material properties Equation (A.1240) and the three imposed quantities \(g,\Delta T,d\) enter the scaled equations only through \(\sigma\) and \(\mathrm{Ra}\), and \(\rho_{0}\) has dropped out entirely, having been absorbed into \(p'\).
∎Linear stability and the critical Rayleigh number
Let the boundaries \(z=0,1\) be stress-free and perfectly conducting, so that \(\psi=\pp_{z}^{2}\psi=\theta=0\) there. Then the conducting state is marginally stable to the roll disturbance of dimensionless horizontal wavenumber \(\pi a\) precisely when
and the minimum of Equation (A.1260) over \(a\) is attained at \(a^{2}=\tfrac{1}{2}\), where
the value quoted in Definition 32.62 and in Equation (31.68). Rests on Proposition A.766 and Phenomenon 31.78.
Derives Theorem A.767. Linearisation. Drop the Jacobians from Equation (A.1253)–Equation (A.1254); they are quadratic in the disturbance and cannot affect the threshold.
The modes. Take
which satisfy every boundary condition: \(\sin(\pi z)\) and its second derivative vanish at \(z=0,1\). Both are eigenfunctions of the Laplacian, \(\nabla^{2}\to-k^{2}\) with
the sum of the squared horizontal wavenumber \(\pi a\) and the squared vertical wavenumber \(\pi\). Marginal stability means a steady disturbance, so set \(\pp_{t}=0\).
The two-by-two system. With \(\pp_{x}\theta=-\pi a\Theta\sin(\pi ax)\sin(\pi z)\) and \(\pp_{x}\psi=\pi a\Psi\cos(\pi ax)\sin(\pi z)\), the two linearised equations reduce to
A nonzero solution exists exactly when the determinant vanishes, \(k^{4}\cdot k^{2}=\mathrm{Ra}\,\pi^{2}a^{2}\), that is
which is Equation (A.1260).
The minimum. Write \(X=a^{2}>0\) and \(F(X)=\left(1+X\right)^{3}/X\). Then
which is negative for \(X<\tfrac{1}{2}\) and positive for \(X>\tfrac{1}{2}\): the unique minimum is at \(X=a^{2}=\tfrac{1}{2}\), and substituting gives Equation (A.1261). Numerically \(27\pi^{4}/4=657.51\). The corresponding horizontal wavelength is \(2\pi/(\pi a)=2/a=2\sqrt{2}\) in units of \(d\), which is the \(\lambda_{\mathrm{c}}=2\sqrt{2}\,d\) of Equation (31.68): convection cells are about \(1.4\) times as wide as the layer is deep, a prediction with no free parameter that Section 36.4 tests.
∎Fluid Dynamics writes the marginal curve as \(\mathrm{Ra}(a)=\left(\pi^{2}+a^{2}\right)^{3}/a^{2}\) and Nonlinear Dynamics and Chaos writes it as Equation (A.1260); the two are the same curve in different variables, and a reader comparing them will otherwise think one of them wrong. In Fluid Dynamics the symbol \(a\) is the horizontal wavenumber itself, in units of \(1/d\); here and in Definition 32.62 it is that wavenumber in units of \(\pi/d\), so that \(a_{\text{ch.\,}14}=\pi a\). Substituting turns one expression into the other, and the minimum sits at \(a^{2}=\pi^{2}/2\) in the first convention and \(a^{2}=\tfrac{1}{2}\) in the second. The second is used here because it is the one in which Equation (32.37) reads \(b=4/(1+a^{2})\), giving Lorenz's \(b=8/3\). Rests on Theorem A.767 and Equation (31.68).
The Galerkin truncation
Fix a horizontal wavenumber \(a\). Write \(\tilde{x}\) for the dimensionless horizontal coordinate and \(z\) for the dimensionless vertical one, reserving \(x(s)\), \(y(s)\) and \(z_{\ast}(s)\) for the three amplitudes — these last are the \(x\), \(y\) and \(z\) of Equation (32.36), decorated here only so that no symbol does two jobs. Substitute into Equation (A.1253)–Equation (A.1254) the three-mode ansatz
with the dimensionless time
and project the resulting equations onto the three modes. Discarding the single term that leaves the retained set, the amplitudes obey
with the dot denoting \(\dd/\dd s\) and
These are Equation (32.36) and Equation (32.37); at the critical wavenumber \(a^{2}=\tfrac{1}{2}\) they give \(b=8/3\). Rests on Proposition A.766, Theorem A.767 and Definition 32.62.
Derives Theorem A.769. Write \(\tilde{x}\) for the horizontal coordinate throughout, to keep it apart from the amplitude \(x(s)\), and \(z_{\ast}\) for the third amplitude, to keep it apart from the vertical coordinate \(z\). Put
with constants \(A,B,C\) to be determined and the abbreviations
All three are mutually orthogonal over one horizontal period and over \(0\le z\le1\), and all satisfy the boundary conditions of Theorem A.767.
Step 1: the vorticity equation has no nonlinearity. Since \(\nabla^{2}\psi=-k^{2}\psi\) with \(k^{2}\) from Equation (A.1263), the Jacobian \(J\left(\psi,\nabla^{2}\psi\right)=-k^{2}J(\psi,\psi)=0\) identically: a single roll does not advect its own vorticity. Using \(\pp_{\tilde{t}}=k^{2}\dd/\dd s\) from Equation (A.1269), Equation (A.1253) becomes
because \(\pp_{\tilde{x}}C_{1}=-\pi aS_{1}\) and the \(S_{2}\) term of \(\theta\) is independent of \(\tilde{x}\). Dividing by \(-k^{4}A/\sigma\),
which is the first of Equation (A.1270) provided
Step 2: the Jacobian in the temperature equation, in full. This is the step that decides whether three modes close, and it cannot be seen without doing it. From Equation (A.1272),
Assembling \(J(\psi,\theta)=\pp_{\tilde{x}}\psi\,\pp_{z}\theta -\pp_{z}\psi\,\pp_{\tilde{x}}\theta\) gives three products. The two involving \(y\) combine, since
using \(\sin\phi\cos\phi=\tfrac{1}{2}\sin2\phi\): the horizontal dependence cancels completely, and what is left is a mode with twice the vertical wavenumber and no horizontal structure. That is the physical origin of the third mode — a roll carrying warm fluid up on one side and cool fluid down on the other distorts the horizontally averaged temperature profile — and it is why \(z_{\ast}\) must be kept. The remaining product is
by \(\sin\phi\cos2\phi=\tfrac{1}{2} \left[\sin3\phi-\sin\phi\right]\). Its \(\sin(\pi z)\) half is \(+\pi^{2}aAC\,xz_{\ast}C_{1}\), a term in the retained set; its \(\sin(3\pi z)\) half is not, and it is the only term produced anywhere in the calculation that leaves the three-mode set. Discarding it is the entire content of the truncation.
Step 3: projection. Insert Equation (A.1278), the retained half of Equation (A.1279), and \(\pp_{\tilde{x}}\psi=\pi aA\,x\,C_{1}\) into Equation (A.1254), use \(\nabla^{2}C_{1}=-k^{2}C_{1}\) and \(\nabla^{2}S_{2}=-4\pi^{2}S_{2}\), and collect the coefficients of the two orthogonal modes. The \(C_{1}\) component gives
and the \(S_{2}\) component, remembering the minus sign carried by \(z_{\ast}\) in Equation (A.1272), gives
Step 4: fixing the constants. Dividing Equation (A.1280) by \(k^{2}B\) and Equation (A.1281) by \(-k^{2}C\),
Comparison with Equation (A.1270) imposes three conditions, to be read together with Equation (A.1276):
The last two give \(C=k^{2}B/\left(\pi^{2}aA\right)\) and \(C=\pi^{2}aAB/\left(2k^{2}\right)\); equating them yields \(\pi^{4}a^{2}A^{2}=2k^{4}\), hence
by Equation (A.1263), which is the coefficient in Equation (A.1267). Then Equation (A.1276) gives
the last equality being Equation (A.1265), and \(C=k^{2}B/(\pi^{2}aA)=B/\sqrt{2} =\mathrm{Ra}_{\mathrm{c}}(a)/\pi\) after substituting Equation (A.1284). With \(\theta=By\,C_{1}-Cz_{\ast}S_{2}\) this is exactly Equation (A.1268). Finally the first of Equation (A.1283) reads, using Equation (A.1276),
and the last term of Equation (A.1282) gives
which completes Equation (A.1271). At \(a^{2}=\tfrac{1}{2}\), \(b=4/(3/2)=8/3\).
∎Reading the parameters back
Undoing the scaling Equation (A.1252) makes every symbol of Equation (32.36) a statement about the layer. The variables. \(x\) is the amplitude of the convective roll: the physical stream function is \(\kappa\,A\,x\,S_{1}\), so the fluid speed is of order \(\left(\kappa/d\right)\left(1+a^{2}\right)\abs{x}\) — \(2\times 10^{-5}\,\mathrm{m}^{2}/\mathrm{s}\) divided by a centimetre, of order \(1\,\mathrm{mm}/\mathrm{s}\) for air at \(\abs{x}\) of order unity. \(y\) measures the horizontal temperature variation at the roll's own wavenumber, and \(z_{\ast}\) the horizontally averaged distortion of the vertical temperature profile. Their common temperature scale is not arbitrary: substituting Equation (A.1285) into Equation (A.1252) gives
where \(\Delta T_{\mathrm{c}}\) is the temperature difference at which that wavenumber first becomes unstable — so \(y\) and \(z_{\ast}\) are temperature departures measured in units of the critical driving difference divided by \(\pi\). The parameters. \(\sigma=\nu/\kappa\) is the Prandtl number of the working fluid alone: about \(0.71\) for air, \(7\) for water at \(20\,\mathrm{^\circ\mathrm{C}}\), of order \(0.02\) for liquid sodium and \(10^{2}\) or more for oils. Lorenz's \(\sigma=10\) is therefore a plausible fluid and not a dial setting. \(r=\mathrm{Ra}/ \mathrm{Ra}_{\mathrm{c}}\) is the imposed temperature difference in units of the one that starts convection, so \(r=28\) means a layer driven \(28\) times past onset. \(b=4/(1+a^{2})\) is a pure function of the horizontal wavenumber of the roll and of nothing else at all. Rests on Theorem A.769 and Remark 32.2.
Five approximations were made, and they are of quite different standing. One is the Boussinesq approximation (Remark A.762), good to a few per cent in ordinary laboratory convection. Two is the restriction to rolls (Remark A.764), which is a fact about the observed flow only near onset. Three is stress-free boundaries, chosen because they make \(\sin\pi z\) an exact eigenfunction; the laboratory case of two rigid plates gives \(\mathrm{Ra}_{\mathrm{c}}\approx1708\) instead of \(657.5\) (Phenomenon 31.78) and no closed-form modes at all, so the tidy arithmetic above is bought with a boundary condition no experiment has. Four is the fixing of a single horizontal wavenumber \(a\), which suppresses every interaction between cells of different sizes. Five, and the serious one, is the discarding of the \(\sin\left(3\pi z\right)\) term at Equation (A.1279). That discarded term is small only while the convection is weak. At Lorenz's \(r=28\) it is not small, the neglected modes are not small, and Lorenz said so plainly in the paper that introduced the system [Lorenz:1963]. The projection is exact on the linear terms — which is why Theorem A.767 is a genuine result about fluids — and an approximation on the nonlinear ones, which is why Equation (32.36) beyond onset is a model of convection and not convection. Remark 32.63 makes the same point in the chapter, and the two statements are meant to stand together: what Equation (32.36) is good for is the study of a low-dimensional dissipative flow that happens to have been derived from a fluid, not the prediction of a heat flux. Rests on Theorem A.769 and Remark 32.63.
The Lorenz System from Boussinesq Convection discharges the derivation owed at Definition 32.62 of Nonlinear Dynamics and Chaos. The four stages announced there are Rolls, the stream function, and the elimination of the pressure (the stream function, which eliminates the pressure), The scaling, and the two dimensionless groups (the explicit scaling and the two dimensionless groups Equation (32.38)), Linear stability and the critical Rayleigh number (the linear stability calculation giving \(\mathrm{Ra}_{\mathrm{c}}=27\pi^{4}/4\) at \(a^{2}=\tfrac{1}{2}\) for stress-free boundaries [Rayleigh:1916]), and The Galerkin truncation (the Galerkin truncation producing Equation (32.36) with the parameters Equation (32.37)). Everything the chapter builds on Equation (32.36) — the volume contraction, the fixed points, the strange attractor — is derived there and does not depend on this section; what this section supplies is the licence to call \(\sigma\), \(r\) and \(b\) measurements rather than dials, which is the demand Remark 32.2 makes of every model in that chapter.
Rayleigh Acoustic Streaming Above a Vibrating Surface
This appendix supplies the derivation owed at Phenomenon 35.8 of Experiment: Waves and Acoustics: the steady second-order flow that an oscillating boundary drives in the fluid above it, and the direction of its outer branch, which is what carries a light powder to the antinodes of a sounding plate where sand collects at the nodes [Faraday:1831] [Chladni:1787].
The scope is fixed in advance by Remark 35.9 of the chapter and is not widened here. What is derived is the streaming velocity, its coefficient, and the sense of the circulation. What is not derived is a threshold grain radius separating sand from lycopodium: that comparison turns on the grain's adhesion to the plate, which this treatise does not model. The reader should hold the chapter's bound in view throughout, because the calculation below is exact enough to invite being pushed further than it can go.
What is quoted here
No theorem. The whole argument is an expansion of the incompressible Navier–Stokes equations Equation (31.32) of Fluid Dynamics in the amplitude of the motion, and every step — the linear boundary-layer solution, the time averages, and five elementary integrals — is carried out here.
Three things are assumed rather than proved, and all three are statements about the physical regime rather than about mathematics. The fluid is treated as incompressible within a plate wavelength, which requires the acoustic wavelength in air to be much the longer of the two (Lemma A.777 states the condition and Example A.779 checks it). The motion is assumed to have settled into a state that is steady in the mean, so that \(\avg{\pp_{t}\vect{v}_{2}}=0\). And the expansion is assumed to be asymptotic in the amplitude, which is the usual and unproved status of such expansions in fluid mechanics.
The result derived below is Rayleigh's, published in 1884 in the
Philosophical Transactions under the title “On the circulation
of air observed in Kundt's tubes, and on some allied acoustical
problems”. That paper has no key in references.bib —
Rayleigh:1883 there is the paper on the equilibrium of a
heavy fluid of variable density, a different work, and citing it here
would be exactly the misattribution the editorial rules warn against
— so the attribution is made in words, on the footing of Darboux's
memoir in Remark A.81. Nothing below rests on
it.
Why the effect cannot appear before second order
Write \(\avg{\cdot}\) for the average over one period \(2\pi/\omega\) of the driving, and expand every field in the dimensionless amplitude \(\epsilon\) of the boundary motion:
At first order the equations are linear with coefficients independent of \(t\) and the forcing is \(\propto\cos\omega t\), so every first-order field is a harmonic function of time of the same frequency and
A tracer carried by the fluid therefore returns, at first order, to where it started after each period: there is no steady transport whatever. Steady transport first appears at order \(\epsilon^{2}\), and it appears there because the average of a product of two oscillating quantities of the same frequency need not vanish even though each factor averages to zero. That is the entire physical content of the phenomenon, and it is why Phenomenon 35.8 cannot be obtained from any first-order acoustic quantity, however cleverly combined.
The first-order Stokes layer
Let a plane wall occupy \(y=0\) with fluid in \(y>0\), and let the outer acoustic field impose, just outside the layer, the tangential velocity \(u_{\infty}=U(x)\cos\omega t\), with \(U\) varying on a length scale much greater than the layer thickness. Then the first-order velocity satisfying no slip at the wall is
with the Stokes-layer thickness
in \(\mathrm{m}\), and the wall-normal component follows from continuity as
where \(h(\eta):=1-E(\eta)\) and \(E(\eta):=\ee^{-\left(1+\ii\right)\eta}\), so that \(u_{1}=U\,\Re\left[h\,\ee^{\ii\omega t}\right]\). Rests on Equations (31.32) and (A.1289).
Derives Proposition A.773. At first order Equation (31.32) linearizes to \(\pp_{t}u_{1}=-\rho^{-1}\pp_{x}p_{1}+\nu\,\pp_{y}^{2}u_{1}\), the convective terms being of second order and the streamwise viscous term smaller than the transverse one by the square of the ratio of the layer thickness to the scale of \(U\). The outer field carries no shear, so it satisfies \(\pp_{t}u_{\infty}=-\rho^{-1}\pp_{x}p_{1}\); and by Equation (A.1289) applied to the pressure together with the thinness of the layer, \(p_{1}\) has the same value inside the layer as just outside it. Subtracting,
since \(\pp_{y}^{2}u_{\infty}=0\). Seek \(u_{1}-u_{\infty} =\Re\left[A(x)\,\ee^{\ii\omega t}\ee^{-\gamma y}\right]\): then \(\ii\omega=\nu\gamma^{2}\), so \(\gamma=\sqrt{\ii\omega/\nu}=\left(1+\ii\right)/\delta\) with \(\delta\) as in Equation (A.1292), the root with positive real part being the one that decays. No slip at \(y=0\) requires \(u_{1}=0\) there, so \(A=-U(x)\) and
and expanding the real part with \(\ee^{-\left(1+\ii\right)\eta}\ee^{\ii\omega t} =\ee^{-\eta}\ee^{\ii\left(\omega t-\eta\right)}\) gives Equation (A.1291).
For \(v_{1}\), integrate continuity \(\pp_{x}u_{1}+\pp_{y}v_{1}=0\) upward from the wall, where \(v_{1}=0\):
and
which is Equation (A.1293). Note that \(v_{1}\) carries the factor \(\dd U/\dd x\): a spatially uniform oscillation drives no normal motion, and it will drive no streaming either.
∎The second-order mean flow and the slip velocity
Let \(\avg{u_{2}}\) be the time-averaged second-order streamwise velocity. Then, to leading order in the thinness of the layer,
with \(\avg{u_{2}}=0\) at \(y=0\) and \(\pp_{y}\avg{u_{2}}\to0\) as \(\eta\to\infty\). Consequently
Rests on Proposition A.773, Equation (31.32) and Equation (A.1290).
Derives Lemma A.774. Take the streamwise component of Equation (31.32) at order \(\epsilon^{2}\) and average over a period. The unsteady term averages to zero in a state steady in the mean; the convective terms contribute \(\avg{u_{1}\pp_{x}u_{1}+v_{1}\pp_{y}u_{1}}\), since the second-order velocity's own convection is of fourth order; and of the viscous term only \(\nu\,\pp_{y}^{2}\avg{u_{2}}\) survives, by the same comparison of transverse with streamwise gradients used in Proposition A.773. Hence
Now evaluate the same equation just outside the layer, where the field is the irrotational \(u_{\infty}\) and there is no mean shear. There
that is, the Reynolds stress of an irrotational oscillation is balanced entirely by the second-order mean pressure and drives nothing. The layer is thin, so \(\avg{p_{2}}\) is the same function of \(x\) inside it as immediately outside; subtracting Equation (A.1301) from Equation (A.1300) eliminates the pressure and leaves Equation (A.1298). By construction \(F\to0\) as \(\eta\to\infty\), and exponentially, since \(u_{1}\to u_{\infty}\) exponentially.
For the boundary conditions: no slip holds at every order, so \(\avg{u_{2}}(0)=0\); and the mean shear at the outer edge of the layer belongs to the outer streaming, which varies on the acoustic length scale rather than on \(\delta\) and is therefore negligible here. Then integrating Equation (A.1298) downward from infinity, \(\mu\,\pp_{y}\avg{u_{2}}(y)=-\int_{y}^{\infty}F\,\dd y'\), and integrating that upward from the wall,
the exchange of the order of integration being legitimate because \(F\) decays exponentially. That is Equation (A.1299).
∎With \(U\) and \(\delta\) as above,
in \(\mathrm{m}/\mathrm{s}\). The viscosity has cancelled: \(u_{s}\) does not contain \(\nu\), and the thickness \(\delta\) of the layer that produced it does not appear. Rests on Lemma A.774, Equation (A.1291) and Equation (A.1293).
Derives Theorem A.775. Write \(g(\eta,t):=\Re\left[h\,\ee^{\ii\omega t}\right]\) and \(G(\eta,t):=\Re\left[H\,\ee^{\ii\omega t}\right]\), so that by Proposition A.773
with \(U'=\dd U/\dd x\). The two Reynolds-stress terms are therefore
the factor \(\delta\) cancelling between \(v_{1}\) and \(\pp_{y}u_{1}\) — which is already the reason the answer will not contain the layer thickness. Since \(\avg{u_{\infty}\pp_{x}u_{\infty}}=UU'\avg{\cos^{2}\omega t} =\tfrac{1}{2}UU'\),
Evaluating \(\Psi\). For two quantities written as \(\Re\left[A\ee^{\ii\omega t}\right]\) and \(\Re\left[B\ee^{\ii\omega t}\right]\) the period average of the product is \(\tfrac{1}{2}\Re\left[A\bar{B}\right]\). With \(\pp_{\eta}g=\Re\left[h'\ee^{\ii\omega t}\right]\) and \(h'=\left(1+\ii\right)E\),
Now \(h=1-E\) with \(E=\ee^{-\eta}\ee^{-\ii\eta}\), so
and, using \(\overline{h'}=\left(1-\ii\right)\bar{E}\) together with \(\left(1-\ii\right)/\left(1+\ii\right)=-\ii\),
Taking real parts with \(\bar{E}=\ee^{-\eta} \left(\cos\eta+\ii\sin\eta\right)\),
the last term of Equation (A.1309) being purely imaginary. Assembling Equations (A.1308) and (A.1310),
The moment. By Equations (A.1299) and (A.1306), with \(y=\delta\eta\),
since \(\rho\delta^{2}/\mu=\delta^{2}/\nu=2/\omega\) by Equation (A.1292). The required integrals follow from
with \(\left(1-\ii\right)^{2}=-2\ii\) and \(\left(1-\ii\right)^{3}=-2-2\ii\). For \(n=1\) the right-hand side is \(1/\left(-2\ii\right)=\ii/2\), so
and for \(n=2\) it is \(2/\left(-2-2\ii\right)=-1/\left(1+\ii\right) =\left(-1+\ii\right)/2\), so
while \(\int_{0}^{\infty}\eta\,\ee^{-2\eta}\dd\eta=\tfrac{1}{4}\). Multiplying Equation (A.1311) by \(\eta\) and integrating term by term,
so \(\int_{0}^{\infty}\eta\Psi\,\dd\eta=\tfrac{3}{8}\) and Equation (A.1312) gives \(u_{s}=-\left(2/\omega\right)\left(3/8\right)UU'\), which is the first member of Equation (A.1303). The second member follows from \(\avg{u_{\infty}^{2}}=\tfrac{1}{2}U^{2}\), whose \(x\) derivative is \(UU'\).
∎The mean flow at the wall itself runs the other way. Its shear there is
which has the sign of \(U\,\dd U/\dd x\), opposite to that of \(u_{s}\) in Equation (A.1303). The mean profile therefore reverses once within the layer, at a height of order \(\delta\). Rests on Lemma A.774, Equation (A.1311) and Theorem A.775.
Derives Proposition A.776. From the proof of Lemma A.774, \(\mu\left(\pp_{y}\avg{u_{2}}\right)_{y=0} =-\int_{0}^{\infty}F\,\dd y =-\rho UU'\delta\int_{0}^{\infty}\Psi\,\dd\eta\). The zeroth moments needed are \(\int_{0}^{\infty}\ee^{-\eta}\cos\eta\,\dd\eta =\int_{0}^{\infty}\ee^{-\eta}\sin\eta\,\dd\eta=\tfrac{1}{2}\), from Equation (A.1313) at \(n=0\), together with \(\int_{0}^{\infty}\ee^{-2\eta}\dd\eta=\tfrac{1}{2}\) and Equation (A.1314). Integrating Equation (A.1311),
so \(\int_{0}^{\infty}\Psi\,\dd\eta=-\tfrac{1}{4}\) and \(\mu\left(\pp_{y}\avg{u_{2}}\right)_{y=0} =\tfrac{1}{4}\rho\delta UU'\). Dividing by \(\mu=\rho\nu\) and using \(\nu=\omega\delta^{2}/2\) gives Equation (A.1317). Since \(\avg{u_{2}}\) starts from zero at the wall with a shear of one sign and ends at \(u_{s}\) of the other, it changes sign once, at a height that the explicit profile places within a small multiple of \(\delta\).
∎The direction of the circulation over a plate
Everything so far holds for any oscillating tangential field. To reach Phenomenon 35.8 one further step is needed, and it is the step at which the answer is decided: the relation between the tangential air velocity \(U(x)\) and the plate's normal motion.
Let the plate vibrate with normal velocity \(v\left(x,0,t\right)=V_{0}\sin\left(k_{p}x\right)\cos\omega t\), so that its nodal lines are at \(\sin\left(k_{p}x\right)=0\) and its antinodes midway between them. Suppose the bending wavelength is much shorter than the acoustic wavelength in the fluid, \(k_{p}\gg\omega/c\), and much longer than the Stokes layer, \(k_{p}\delta\ll1\). Then the air motion within a bending wavelength is incompressible and irrotational, with
and the tangential velocity imposed on the layer is
Its magnitude is therefore greatest over the plate's nodal lines and vanishes over its antinodes. Rests on Proposition 31.28, Equation (31.25) and Equation (A.1292).
Derives Lemma A.777. When \(k_{p}\gg\omega/c\) the compressible wave equation for the velocity potential, \(\nabla^{2}\phi=c^{-2}\pp_{t}^{2}\phi\), reduces within a bending wavelength to Equation (31.25), because the term on the right is smaller than \(\pp_{x}^{2}\phi\) by \(\left(\omega/ck_{p}\right)^{2}\). Separating with the \(x\) dependence imposed by the plate and requiring decay as \(y\to\infty\) gives \(\phi\propto\sin\left(k_{p}x\right)\ee^{-k_{p}y}\), and \(\nabla^{2}\phi=\left(-k_{p}^{2}+k_{p}^{2}\right)\phi=0\) confirms it. Fixing the constant by \(\pp_{y}\phi=v\left(x,0,t\right)\) at \(y=0\) gives Equation (A.1319), whence
Since \(k_{p}\delta\ll1\), the exponential is indistinguishable from unity across the Stokes layer, so the field Equation (A.1321) is what Proposition A.773 sees as \(u_{\infty}\), and Equation (A.1320) follows.
The physical reading is worth one sentence, because the result is the opposite of the naive guess. Over an antinode the plate pushes air straight up and down and the horizontal motion vanishes by symmetry; over a nodal line the plate does not move at all, but the air expelled from the antinode on one side must cross to the other, and it is there that the horizontal velocity is greatest.
∎With Equation (A.1320), Equation (A.1303) gives
which is directed from each nodal line towards the antinodes on either side of it. The outer branch of the circulation therefore sweeps the fluid just above the layer from the nodes to the antinodes, rising there and returning above; a grain carried by that branch collects at the antinodes, where sand does not, which is Phenomenon 35.8. Rests on Theorem A.775, Lemma A.777 and Phenomenon 35.5.
Derives Corollary A.778. \(U^{2}=V_{0}^{2}\cos^{2}\left(k_{p}x\right)\), so \(\dd\left(U^{2}\right)/\dd x =-V_{0}^{2}k_{p}\sin\left(2k_{p}x\right)\) and \(u_{s}=-\left(3/8\omega\right)\dd\left(U^{2}\right)/\dd x\) is Equation (A.1322). Take a nodal line at \(x=0\) and the neighbouring antinode at \(x=\pi/2k_{p}\): on the interval between them \(\sin\left(2k_{p}x\right)>0\), so \(u_{s}>0\) and the flow is directed towards the antinode. By the reflection symmetry about the nodal line the same holds on the other side, so the nodal lines are lines of divergence of the slip flow and the antinodes are lines of convergence. Mass conservation then requires the fluid to rise above each antinode and to descend over each node, closing the circulation with an outward branch aloft.
The statement Equation (A.1303) makes about \(\abs{U}\) is general; what is special to the plate is Lemma A.777, which places the maxima of \(\abs{U}\) over the nodes. The two together are what invert the figure.
∎Take a plate sounding at \(f=500\,\mathrm{Hz}\) with a bending wavelength \(\lambda_{p}=0.10\,\mathrm{m}\), so \(\omega=3.14\times 10^{3}\,/\mathrm{s}\) and \(k_{p}=63\,/\mathrm{m}\), in air of kinematic viscosity \(1.5\times 10^{-5}\,\mathrm{m}^{2}/\mathrm{s}\) and sound speed \(343\,\mathrm{m}/\mathrm{s}\). Then
against an acoustic wavelength \(c/f=0.69\,\mathrm{m}\). The two hypotheses of Lemma A.777 are met with room to spare in one case and by about a factor of seven in the other: \(k_{p}\delta=6.3\times 10^{-3}\ll1\), and \(k_{p}/\left(\omega/c\right)=63/9.2=6.8\gg1\). The second ratio is the one to watch, since it fails at the coincidence frequency where the bending wave becomes sonic, and above it the plate radiates and the near field is no longer evanescent.
With a plate displacement amplitude \(A_{p}=50\,\mu\mathrm{m}\), so \(V_{0}=\omega A_{p}=0.16\,\mathrm{m}/\mathrm{s}\), Equation (A.1322) gives a peak slip velocity
which moves a grain across a five-centimetre cell in a few minutes — slow, steady and unmistakable, which is what the observation reports. The same amplitude gives a peak plate acceleration \(\omega^{2}A_{p}=4.9\times 10^{2}\,\mathrm{m}/\mathrm{s}^{2}\), some fifty times \(g\), so sand is thrown vigorously; the two mechanisms are operating simultaneously and on different powders, exactly as the chapter says. Rests on Equations (A.1292) and (A.1322).
The same mechanism in Kundt's tube
In a gas column carrying a standing wave of wavenumber \(k\) along a tube, the tangential velocity at the wall is \(U(x)=U_{0}\sin\left(kx\right)\) with its maxima at the displacement antinodes. Equation (A.1303) then gives
directed from each displacement antinode towards the displacement nodes on either side. A powder swept along the wall by that flow therefore accumulates at the displacement nodes, spaced by half a wavelength — which is Phenomenon 35.12. Rests on Theorem A.775, Equation (A.1303) and Equation (35.17).
Derives Corollary A.780. In a standing wave the velocity amplitude is greatest at the displacement antinodes and zero at the nodes, and the wall is parallel to the axis, so the tangential field imposed on the Stokes layer is \(U_{0}\sin\left(kx\right)\) with \(U_{0}\) the velocity amplitude. Then \(\dd\left(U^{2}\right)/\dd x=U_{0}^{2}k\sin\left(2kx\right)\) and Equation (A.1303) gives Equation (A.1325), the second form using \(\omega=ck\). Between a node at \(x=0\) and the antinode at \(x=\pi/2k\), \(\sin\left(2kx\right)>0\) and \(u_{s}<0\): the flow runs from the antinode back to the node. So the displacement nodes are lines of convergence and the ridges form there.
Phenomenon 35.12 asserts this collection and uses the ridge spacing to measure a wavelength; the chapter takes the location of the ridges from the observation. It is worth noting that the mechanism is the same as the plate's and the apparent reversal between them is not in the streaming at all but in Lemma A.777: in the tube the maxima of \(\abs{U}\) coincide with the visible antinodes of the wave, whereas over a plate they sit above its nodal lines. In both cases the powder goes where \(\abs{U}\) is least.
∎What is not derived here
Which branch a given grain samples. By Proposition A.776 the mean flow reverses within the layer, so a grain lying in contact with the plate is not in the same current as one riding above it. A lycopodium spore is some \(30\,\mu\mathrm{m}\) across against the \(\delta\approx0.1\,\mathrm{mm}\) of Equation (A.1323): it protrudes into the layer but does not clear it, so which branch dominates the drag on it is a quantitative question about a body of comparable size to the layer, and neither Theorem A.775 nor Proposition A.776 settles it. What is derived here is the outer branch and its direction, which is precisely the bound Remark 35.9 sets.
Adhesion. The competition the chapter names — between the streaming drag and the residual contact friction — requires a model of the grain's adhesion to the plate. There is none in this treatise, so no threshold radius separating the two powders is predicted, and none should be read into Equation (A.1324).
The figure itself. Equation (A.1322) was computed for a one-dimensional mode shape. A real Chladni figure is two-dimensional, and the streaming above it is a two-dimensional pattern of cells whose stagnation structure is not obtained by writing \(\sin\left(2k_{p}x\right)\) twice. That the powder ends up on the antinodal regions follows from the sign of Equation (A.1303) along any line crossing a nodal curve, which is all that is claimed.
Rayleigh Acoustic Streaming Above a Vibrating Surface discharges the derivation owed at Phenomenon 35.8 of Experiment: Waves and Acoustics, in the Chladni experiment Section 35.3. The reader returning there should carry back the reason the chapter keeps the phenomenon at all. The sand figure and the lycopodium figure are traced by the same mode of the same plate at the same amplitude, and they are complementary: one records where the plate is still, the other where the air is still. The first is a threshold in \(\omega^{2}A\) against \(g\), obtained in Continuum Mechanics and Elasticity; the second is Equation (A.1303), a second-order effect in a fluid that the plate merely stirs. Neither figure is the eigenfunction, and an experiment that images a field through a tracer measures the tracer as well as the field — which is the standing warning Experiment: Waves and Acoustics attaches to Phenomenon 35.8, and which this derivation makes quantitative.
The Free-Streamline Jet: Kirchhoff's Contraction Coefficient
This appendix proves Proposition 36.3 of Experiment: Fluid Flow and Turbulence: the two-dimensional efflux of an ideal incompressible liquid through a sharp-edged slit of width \(a\) in a plane wall contracts to an asymptotic jet width \(a_{j}\) with
which is Equation (36.6), and does so with no empirical input whatever [Helmholtz:1868] [Kirchhoff:1869].
The chapter's paragraph after the prooflink explains why the
problem is hard and what the hodograph does about it; that discussion
is not repeated here. This section begins where it stops: it builds the
hodograph region explicitly, maps it, and integrates back.
What is quoted here
One theorem is imported, and it is the only one.
Let \(D\subsetneq\C\) be a simply connected domain whose boundary on the Riemann sphere is a Jordan curve. Then there is a holomorphic bijection of \(D\) onto the upper half plane \(\mathbb{H}:=\set{t\in\C:\Im t>0}\), it extends to a homeomorphism of the closures, and it is unique once three boundary points and their cyclic order are prescribed. Consequently two such maps that agree on three boundary points agree everywhere. Rests on Definition 8.5 and Theorem 8.6.
Complex Analysis builds holomorphy (Definition 8.5), the Cauchy–Riemann equations (Theorem 8.6), Cauchy's theorem and the residue calculus, but no theory of conformal mapping: it has no Riemann mapping theorem, no boundary-correspondence theorem and no Schwarz–Christoffel formula. Theorem A.783 is therefore imported, and it is used for exactly one purpose — to identify the parameter \(t\) reached from the hodograph plane with the parameter \(t\) reached from the potential plane. Everything else below is elementary and is verified here: the mapping properties of \(\exp\), of \(\chi\mapsto \chi^{2}\) and of the Joukowski map are established by tracking their boundary values corner by corner, so no general mapping theorem is needed for them, and the Schwarz–Christoffel formula, which the standard treatments reach for at this point, is not used at all.
The imported theorem has no key in references.bib — the only
Riemann and Carathéodory entries there are a paper on electrodynamics
and one on the foundations of thermodynamics — so its attribution is
made in words, on the footing of Darboux's memoir in
Remark A.81: Riemann's inaugural dissertation of
1851 for the mapping, Carathéodory's papers of 1913 for the extension
to the boundary. A reader who grants
Theorem A.783 is granted everything else
in this section.
Two further limitations are honest to state here rather than at the end. The construction below is a verification: it exhibits a flow satisfying every condition of the problem and computes its contraction. That the free-boundary problem has no other solution is not proved, and the treatise does not carry the machinery to prove it. And gravity is neglected within the jet, so that the speed on the free streamline is the single constant \(q_{0}\); this is exact only in the limit \(a\ll h\), and the chapter's measurements are made in that regime.
Formulation
Let the wall occupy the plane \(x=0\) with a slit \(\abs{y}<a/2\), the reservoir the region \(x<0\), and the jet issue into \(x>0\); the flow is two-dimensional, steady, incompressible and irrotational, so by Proposition 31.28 it has a harmonic velocity potential Equation (31.25) and, with the stream function Equation (31.12), a holomorphic complex potential \(w(z)=\phi+\ii\psi\) of \(z=x+\ii y\), whose derivative is the complex velocity Equation (31.57),
with \(q\ge0\) the speed in \(\mathrm{m}/\mathrm{s}\) and \(\theta\) the direction of the velocity.
Three conditions define the problem.
-
Impermeability on the wall: \(\psi\) is constant there.
-
Free surface: the jet is bounded by two streamlines on which the pressure equals the atmospheric pressure, so by Theorem 31.24 and Equation (36.2) the speed on them is the constant
\begin{equation}\tag{A.1328} q_{0}=\sqrt{2gh}\ec \end{equation}\(h\) being the head in \(\mathrm{m}\) and \(g\) the acceleration of free fall.
-
Position of the free surface: unknown. It is a streamline whose location is part of the answer.
The third condition is what makes this a free-boundary problem, and it is the reason no direct solution of Equation (31.25) is available: the domain on which the equation is to be solved is not given.
The configuration is symmetric about \(y=0\), so it suffices to treat the upper half. Its boundary consists of three arcs, on each of which the stream function is constant:
-
the axis \(y=0\), \(-\infty<x<+\infty\), on which \(\psi=0\);
-
the wall \(x=0\), \(y\ge a/2\), on which \(\psi=m\);
-
the free streamline, leaving the edge \(E=(0,a/2)\) and running downstream to \(x=+\infty\), \(y=a_{j}/2\), on which also \(\psi=m\).
Here \(m\) is half the volume flux per unit span; evaluating it far downstream, where the jet is uniform at speed \(q_{0}\) and width \(a_{j}\),
in \(\mathrm{m}^{2}/\mathrm{s}\). The image of the flow region in the \(w\) plane is therefore the strip \(0<\psi<m\): the reservoir at infinity is the end \(\phi\to-\infty\), the jet at infinity the end \(\phi\to+\infty\), and \(\phi\) is monotone along every streamline.
The hodograph region is a half-strip
Put
Then the flow region maps into the semi-infinite strip
with the three boundary arcs going to the three sides as follows: the free streamline to \(\sigma=0\), the wall to \(\theta=-\pi/2\), the axis to \(\theta=0\), and the reservoir at infinity to \(\sigma=+\infty\). Rests on Equation (A.1327), Equation (A.1328) and Definition 8.3.
Derives Lemma A.785. Take the boundary arcs in turn. On the free streamline \(q=q_{0}\) by Equation (A.1328), so \(\sigma=0\) exactly; the direction turns from \(\theta=-\pi/2\) at the edge, where the fluid is still moving down the inner face of the wall, to \(\theta=0\) far downstream, where the jet is parallel to the axis. On the wall \(x=0\), \(y>a/2\), the velocity is tangent to the wall and directed towards the slit, that is in the \(-y\) direction, so \(\theta=-\pi/2\) throughout, while the speed rises from \(0\) deep in the reservoir to \(q_{0}\) at the edge — so \(\sigma\) falls from \(+\infty\) to \(0\). On the axis \(\theta=0\) by symmetry, and the speed rises from \(0\) far upstream to \(q_{0}\) far downstream, so \(\sigma\) falls from \(+\infty\) to \(0\).
Everywhere in the flow \(0<q<q_{0}\), since by Theorem 31.24 the speed is greatest where the pressure is least and the least pressure in the field is the atmospheric pressure on the free surface; and \(-\pi/2<\theta<0\), since the flow turns monotonically from the wall direction to the axis direction between the two. Hence the image lies in \(S\), and the three sides of \(S\) carry the three boundary arcs as listed. The corner \(\zeta=0\) is the edge \(E\), and the corner \(\zeta=-\ii\pi/2\) at \(\sigma=0\) does not occur: the wall and the free streamline meet only at \(E\).
∎This is the step the chapter's prose promises, and it is worth naming what it has achieved. The unknown of the problem was the shape of the free surface. In the \(\zeta\) plane that surface is the segment \(\sigma=0\), \(-\pi/2\le\theta\le0\) — a known set, fixed before anything has been solved. The price is that the map from the physical plane to the hodograph plane is itself unknown; but that map is what the potential will supply.
Two maps onto the same half plane
Define
Then \(\zeta\mapsto t\) carries \(S\) holomorphically and bijectively onto \(\mathbb{H}\), with the boundary correspondence
the edge \(E\) going to \(t=1\), the jet at infinity to \(t=-1\) and the reservoir at infinity to \(t=\infty\). On the free streamline the parametrization is
Rests on Lemma A.785, Proposition 8.4 and Definition 8.3.
Derives Lemma A.786. Each map is holomorphic where used, and each is checked on the boundary.
The exponential. \(\chi=\ee^{-\zeta}=\ee^{-\sigma}\ee^{-\ii\theta}\) has \(\abs{\chi}=\ee^{-\sigma}\le1\) and \(\arg\chi=-\theta\in[0,\pi/2]\), and \(\zeta\mapsto\chi\) is injective on \(S\) because the imaginary part of \(\zeta\) spans an interval of length \(\pi/2<2\pi\). So \(\chi\) ranges over the open quarter disc \(\set{0<\abs{\chi}<1,\ 0<\arg\chi<\pi/2}\), and the three sides of \(S\) go to the two radii and the arc.
The square. On the quarter disc \(\chi\mapsto s=\chi^{2}\) is injective, since \(\arg\chi\) spans an interval of length \(\pi/2\) and doubling it spans \(\pi<2\pi\). Its image is the open half disc \(\set{0<\abs{s}<1,\ 0<\arg s<\pi}\); the arc \(\abs{\chi}=1\) goes to the arc \(\abs{s}=1\), the radius \(\arg\chi=0\) to \(s\in(0,1)\) and the radius \(\arg\chi=\pi/2\) to \(s\in(-1,0)\).
The Joukowski map. On the half disc, write \(s=\varrho\ee^{\ii \vartheta}\) and
For \(0<\varrho<1\) the bracket \(\varrho-1/\varrho\) is negative and \(\sin\vartheta>0\), so \(\Im t>0\): the half disc goes into \(\mathbb{H}\). Injectivity is the observation that \(s\) and \(1/s\) are the two preimages of a given \(t\) and that exactly one of them has modulus less than one. Surjectivity onto \(\mathbb{H}\) follows because for each \(t\in\mathbb{H}\) the quadratic \(s^{2}+2ts+1=0\) has roots with product \(1\) and neither on the unit circle, since a root of modulus one would make \(t\) real by Equation (A.1335).
On the boundary: \(\abs{s}=1\), \(s=\ee^{\ii\vartheta}\), gives \(t=-\cos\vartheta\in[-1,1]\); \(s\in(0,1)\) gives \(t=-\tfrac{1}{2}\left(s+1/s\right)\le-1\); and \(s\in(-1,0)\) gives \(t\ge1\). Tracing back through the two earlier maps, \(s\in(0,1)\) is the axis and \(s\in(-1,0)\) is the wall, which is Equation (A.1333). On the free streamline \(\Omega=q_{0}\ee^{-\ii\theta}\), so \(\chi=\ee^{\ii\alpha}\) with \(\alpha=-\theta\), \(s=\ee^{2\ii\alpha}\) and \(t=-\cos2\alpha\), which is Equation (A.1334). Its endpoints are \(t=1\) at \(\alpha=\pi/2\) (the edge) and \(t=-1\) at \(\alpha=0\) (downstream).
∎Fix the origin of the velocity potential at the edge \(E\) and define
Then \(w\mapsto t\) carries the strip \(0<\psi<m\) holomorphically and bijectively onto \(\mathbb{H}\), with the same boundary correspondence Equation (A.1333), and
Rests on Equation (A.1329), Definition 8.3 and Lemma A.786.
Derives Lemma A.787. For \(0<\psi<m\) the argument of \(\ee^{-\pi w/m}\) is \(-\pi\psi/m\in(-\pi,0)\), so the argument of \(-2\ee^{-\pi w/m}\) lies in \((0,\pi)\) and \(\Im t>0\); the map is injective because the imaginary part of \(-\pi w/m\) spans an interval of length \(\pi<2\pi\), and it is onto \(\mathbb{H}\) because \(\ee^{-\pi w/m}\) covers the lower half plane.
On \(\psi=0\), \(t=-1-2\ee^{-\pi\phi/m}\) decreases through \((-\infty,-1)\) as \(\phi\) runs from \(-\infty\) to \(+\infty\): the reservoir goes to \(t=\infty\) and the jet at infinity to \(t=-1\). On \(\psi=m\), \(\ee^{-\pi\left(\phi+\ii m\right)/m}=-\ee^{-\pi\phi/m}\), so \(t=-1+2\ee^{-\pi\phi/m}\), which is \(+\infty\) deep in the reservoir, equals \(1\) at \(\phi=0\), and tends to \(-1\) far downstream. The choice of origin therefore puts the edge at \(t=1\), matching Lemma A.786 arc for arc.
Differentiating \(t+1=-2\ee^{-\pi w/m}\) gives \(\dd t=-\left(\pi/m\right)\left(t+1\right)\dd w\), which is Equation (A.1337).
∎Lemmas A.786 and A.787 are two conformal maps of the flow region onto \(\mathbb{H}\) agreeing at the three boundary points \(E\), the jet at infinity, and the reservoir at infinity. By Theorem A.783 they are the same map. Hence \(\Omega\) and \(w\) are both explicit functions of one parameter \(t\), and
Rests on Theorem A.783, Lemma A.786 and Lemma A.787.
This is the whole solution. The physical plane has been eliminated in favour of \(t\), and Equation (A.1338) recovers it by quadrature.
Integrating back along the free streamline
The free streamline leaves the edge of the slit at height \(a/2\) and approaches the height \(a_{j}/2\) downstream, with
so that \(a=a_{j}\left(\pi+2\right)/\pi\) and Equation (A.1326) holds. Rests on Corollary A.788, Equation (A.1334) and Equation (A.1329).
Derives Theorem A.789. Parametrize the free streamline by \(\alpha\) as in Equation (A.1334), so that \(\alpha=\pi/2\) at the edge and \(\alpha=0\) downstream. Then
and Equation (A.1337) gives
using \(\sin2\alpha=2\sin\alpha\cos\alpha\). The sign is right: along the streamline \(\alpha\) decreases and \(\cot\alpha>0\), so \(\dd\phi>0\) and the potential increases downstream.
On the free streamline \(\Omega=q_{0}\ee^{-\ii\theta} =q_{0}\ee^{\ii\alpha}\), so
whose imaginary part is
Integrating from the edge to the far downstream station,
and Equation (A.1329) turns \(2m/q_{0}\) into \(a_{j}\), giving Equation (A.1339). Since the streamline starts at \(y=a/2\) and ends at \(y=a_{j}/2\),
which is Equation (A.1326). Numerically \(\pi/\left(\pi+2\right)=0.61101\).
Note what has and has not entered. The head \(h\) appears only through \(q_{0}\), and \(q_{0}\) cancels between Equation (A.1329) and Equation (A.1344); so does the density, which never appeared at all. The result is a pure number, as Proposition 36.3 claims, and it is the same number at every head.
∎The \(\pi\) and the \(2\) of Equation (A.1326) have different origins, and the calculation is worth reading once with that in mind. The \(\pi\) is the width of the potential strip: it enters through the exponential map Equation (A.1336), whose exponent is \(\pi w/m\) precisely because the strip has height \(m\) and must be opened into a half plane. It is therefore a statement about the flux. The \(2\) is \(\int_{0}^{\pi/2}\cos\alpha\,\dd\alpha=1\) counted for the two halves of the jet — a statement about the turning of the velocity through a right angle at the edge. Nothing else survives. (No arctangent occurs anywhere in the evaluation, and no Schwarz–Christoffel integral is needed: the one quadrature to be done is Equation (A.1344), and it is elementary.)
The edge. At \(\alpha=\pi/2\) the speed is \(q_{0}\), finite. A free-streamline solution is admissible only if the velocity is finite where the free surface leaves the solid boundary, and here that condition holds automatically, because \(\abs{\Omega}=q_{0}\) on the whole free streamline including its endpoint. Nothing had to be imposed to secure it, which is a feature of the sharp-edged geometry and is not true of a rounded mouthpiece.
The asymptote. The real part of Equation (A.1342) is \(\dd x=-\left(2m/\pi q_{0}\right) \left(\cos^{2}\alpha/\sin\alpha\right)\dd\alpha\), whose integrand behaves as \(1/\alpha\) as \(\alpha\to0\). The streamline therefore runs off to \(x=+\infty\) logarithmically in \(\alpha\), while Equation (A.1343) shows the remaining drop in \(y\) is of order \(\alpha^{2}\): the jet approaches its asymptotic width exponentially in \(x\), so “the” contraction is reached within a few slit widths and is a measurable quantity rather than a limit that is never attained.
What the theorem does not cover
The hodograph is a two-dimensional device. It works because the velocity of a plane irrotational flow is a holomorphic function of position, so that the map \(z\mapsto\Omega\) can be treated as a conformal change of variable and an unknown boundary in one plane becomes a known boundary in another. In three dimensions there is no such structure: the velocity is a harmonic vector field, the “hodograph” is a map from a three-dimensional region to another and carries no conformal group beyond the Möbius transformations, and no closed solution of the axisymmetric free-boundary problem is known. The measured contraction coefficient for a sharp-edged circular orifice is near \(0.61\) — close enough to Equation (A.1326) that the coincidence is often quoted as though the theorem covered it. It does not, and Proposition 36.3 says so.
Proposition 36.2 obtains \(C_{c}=1/2\) exactly for the re-entrant mouthpiece, and it does so in half a page with no complex analysis at all. The contrast is instructive rather than embarrassing. What made the flush slit hard was that the pressure over the wall around the orifice is unknown, so the momentum balance carries an undetermined term; the whole apparatus above exists to determine it by finding the flow. The re-entrant geometry removes the term instead: the tube is wetted on both faces, so it transmits no net force, and every other wetted surface carries a hydrostatic pressure. The unknown was arranged out of the problem rather than computed.
The two results are consistent and they bracket the observations: \(\tfrac{1}{2}\) for the re-entrant mouthpiece, \(\pi/(\pi+2)\) for the flush slit, and measurements in between for intermediate geometries. Neither is an approximation to the other.
The Free-Streamline Jet: Kirchhoff's Contraction Coefficient discharges the derivation owed at Proposition 36.3 of Experiment: Fluid Flow and Turbulence, in the efflux experiment Section 36.1. The reader returning there should carry back the division the chapter draws in its Interpretation: Bernoulli's theorem gives the speed Equation (36.2) exactly and gives the discharge not at all, because it says nothing about the area of the bundle of streamlines that reaches the jet. This section supplies that area for one geometry — and the reason it can is that the geometry is two-dimensional and sharp-edged, so the free surface has a known image in the hodograph plane. That is a narrower statement than “the ideal theory predicts the discharge”, and the narrowness is the point.
Prandtl's Boundary-Layer Equations and the Blasius Similarity Solution
This appendix proves Proposition 36.15 of Experiment: Fluid Flow and Turbulence: that at large Reynolds number the Navier–Stokes system Equation (31.32) reduces, within the layer against a body, to Prandtl's equations Equation (36.26); that for a flat plate at zero incidence those equations admit a similarity solution governed by the Blasius ordinary differential equation Equation (36.27); and that the three constants quoted in Phenomenon 36.12 — the \(5.0\) of the thickness, the \(0.664\) of the local friction coefficient and the \(1.328\) of the plate drag — are consequences of two numbers read off the numerical solution of that equation [Prandtl:1904] [Blasius:1908].
The chapter reaches Equation (36.23) by a dominant-balance estimate. This section derives what that estimate guessed: the exponent \(-1/2\) is obtained rather than assumed, and the estimate's silence about the pressure is replaced by a statement of where the pressure comes from.
What is quoted here
Two things, of very different kinds.
The matching principle. The reduction below is a matched asymptotic expansion, and this treatise builds no theory of matched asymptotics. Section 9.6 of Ordinary Differential Equations and Sturm–Liouville Theory carries the WKB approximation, which is a singular perturbation of the same family — a small parameter multiplying the highest derivative, an outer solution that cannot satisfy every boundary condition, and an inner solution on a stretched variable — and its connection formulae play the part that matching plays here. But WKB is a theorem about one linear second-order equation, not a general principle, and nothing in Part II licenses what follows. The principle is therefore stated outright and used as stated:
Let a problem carry a small parameter \(\varepsilon\), let \(F_{\text{out}}\) be the leading term of an expansion valid away from the wall in the original variable \(\hat y\), and let \(F_{\text{in}}\) be the leading term of an expansion valid near the wall in a stretched variable \(Y=\hat y/\varepsilon^{p}\). Then the two expansions are assumed to possess a common domain of validity in which
applied separately to each dependent variable. Rests on Equation (31.32) and Theorem 9.72.
Two numbers. Equation (36.27) has no solution in closed form. The values
are quadratures of a numerical integration, quoted from [Blasius:1908] and not computed here. Every algebraic manipulation that turns them into the three constants of Phenomenon 36.12 is carried out below, so what is imported is exactly two decimal numbers and nothing structural.
Postulate A.795 is an assumption about
the solution, not a theorem: it asserts that the two expansions
overlap. For the flat plate the assertion can be checked a
posteriori — the similarity solution constructed in
The flat plate: similarity and the Blasius equation approaches its free-stream
value exponentially in \(\eta\), so the overlap region is genuine and
wide — but that is a verification in one case, not a proof in
general. The debt is Part II's, and it is named in the chapter's prose
after the prooflink as well as here.
The two numbers of Equation (A.1347) are the only external inputs to the arithmetic. Note in particular that \(f''(0)\) enters the friction coefficient linearly and the drag coefficient linearly, so the relation \(1.328=2\times0.664\) between the last two constants is exact and independent of the numerics, as The three constants shows.
Scaling, and why the inviscid limit is singular
Take steady, two-dimensional, incompressible flow past a body of streamwise extent \(L\) in a stream of speed \(U\), with gravity absorbed into the pressure. Equation (31.32) reads
Scale on the body: \(x=L\hat x\), \(y=L\hat y\), \(u=U\hat u\), \(v=U\hat v\), \(p=p_{\infty}+\rho U^{2}\hat p\). Every hatted quantity is dimensionless and, by construction, of order unity in the outer flow. Dividing Equation (A.1348) by \(U^{2}/L\) and Equation (A.1349) likewise,
with \(\mathrm{Re}=UL/\nu\) the Reynolds number Equation (36.9). The small parameter is \(\varepsilon=1/\mathrm{Re}\), and it multiplies the highest derivative — which is the definition of a singular perturbation and the whole source of the difficulty.
Setting \(\varepsilon=0\) in Equations (A.1351) and (A.1352) gives the Euler equations Equation (31.19), whose solution for flow past a body is the potential flow of Proposition 31.28. That solution satisfies impermeability, \(\hat v=0\) on the surface, and cannot in general satisfy no slip, \(\hat u=0\) there. The limit \(\mathrm{Re}\to\infty\) is therefore not uniform in \(\hat y\). Rests on Theorem 31.22, Proposition 31.28 and Equation (A.1351).
Derives Proposition A.797. With \(\varepsilon=0\) the system is first order in \(\hat y\) and admits one condition on the normal component; the potential solution of Equation (31.25) with \(\pp\phi/\pp n=0\) supplies it and leaves the tangential velocity at the wall determined — and non-zero, by Equation (31.57) and the geometry, except at stagnation points. Since the true solution has \(\hat u=0\) at the wall for every \(\mathrm{Re}\), however large, \(\lim_{\mathrm{Re}\to\infty}\hat u\) is discontinuous at \(\hat y=0\) and the convergence cannot be uniform. This is d'Alembert's error stated as an analytic fact rather than a physical one, and Phenomenon 36.12 is its resolution.
∎The inner expansion: Prandtl's equations
Introduce the stretched wall-normal coordinate and the rescaled normal velocity
with \(Y\) and \(V\) of order unity. Then, to leading order in \(\mathrm{Re}^{-1}\),
Restoring dimensions, with \(\delta\sim L\,\mathrm{Re}^{-1/2}\) the layer thickness of Equation (36.23), this is Equation (36.26). The exponent \(\tfrac{1}{2}\) in Equation (A.1354) is forced: it is the only choice for which the viscous term survives the limit without dominating it. Rests on Equations (36.23), (A.1351) and (A.1353).
Derives Theorem A.798. The exponent. Try \(Y=\hat y\,\mathrm{Re}^{\,p}\) with \(p>0\). In Equation (A.1351) the convective term \(\hat u\,\pp_{\hat x}\hat u\) is of order unity, while the surviving viscous term is \(\mathrm{Re}^{-1}\pp_{\hat y}^{2}\hat u =\mathrm{Re}^{2p-1}\pp_{Y}^{2}\hat u\). If \(2p-1<0\) the viscous term vanishes in the limit and the no-slip condition is again lost; if \(2p-1>0\) it dominates and the leading-order equation is \(\pp_{Y}^{2}\hat u=0\), whose solution is linear in \(Y\) and cannot match a bounded outer flow at \(Y\to\infty\) while vanishing at the wall, except trivially. Only \(p=\tfrac{1}{2}\) retains both, which is Equation (A.1354). Note that this derives \(\delta/L\sim\mathrm{Re}^{-1/2}\), which Equation (36.23) obtained by comparing orders of magnitude.
The normal velocity. With \(Y\) as above, Equation (A.1353) reads \(\pp_{\hat x}\hat u+\mathrm{Re}^{1/2}\pp_{Y}\hat v=0\). For continuity to survive at leading order, \(\hat v\) must itself be of order \(\mathrm{Re}^{-1/2}\), which is the second member of Equation (A.1354), and then Equation (A.1357) follows exactly. The layer is thin and the flow through it is weak, in the same ratio.
Streamwise momentum. Substituting into Equation (A.1351),
every term of which is of order unity except \(\mathrm{Re}^{-1}\pp_{\hat x}^{2}\hat u\), which is smaller by \(\mathrm{Re}^{-1}\) and is dropped. That is Equation (A.1355), the pressure gradient being written as a total derivative in anticipation of Equation (A.1356).
Normal momentum, and the pressure. This is the step usually waved through, and it carries the most important consequence, so it is done in full. Substituting \(\hat v=\mathrm{Re}^{-1/2}V\) and \(\pp_{\hat y}=\mathrm{Re}^{1/2}\pp_{Y}\) into Equation (A.1352), the convective terms are
the pressure term is \(-\pp_{\hat y}\hat p=-\mathrm{Re}^{1/2}\pp_{Y}\hat p\), and the viscous term is
Multiplying the equation through by \(\mathrm{Re}^{-1/2}\) collects the orders as
so \(\pp_{Y}\hat p=O\!\left(\mathrm{Re}^{-1}\right)\) and at leading order Equation (A.1356) holds. In words: the pressure is uniform across the layer. It is therefore not an unknown of the layer at all — it is whatever the outer flow imposes at the wall — and this is exactly the sentence Equation (36.26) makes when it writes \(\pp_{y}p=0\).
∎Matching hands over the pressure
Let \(U_{e}(x)\) be the tangential velocity of the outer potential solution evaluated at the surface. Then Postulate A.795 gives
and, applied to the pressure,
so the coefficient in Equation (A.1355) is a known function of \(x\) before the layer problem is posed. Rests on Postulate A.795, Theorem A.798 and Theorem 31.24.
Derives Corollary A.799. Equation (A.1362) is Equation (A.1346) applied to \(\hat u\): the inner solution at large \(Y\) must agree with the outer solution at small \(\hat y\), and the latter is the wall value \(U_{e}\) of the potential flow, since the potential solution is smooth up to the wall. Applied to the pressure and combined with Equation (A.1356), which makes \(\hat p\) independent of \(Y\), it gives \(\hat p\left(\hat x\right) =\hat p_{\text{out}}\left(\hat x,0\right)\). In the outer flow Theorem 31.24 holds along the wall streamline, so \(p+\tfrac{1}{2}\rho U_{e}^{2}\) is constant there, which on differentiating is Equation (A.1363).
The logical shape is worth stating plainly, because it is what makes the boundary-layer method a method and not merely an approximation. Nothing has been solved simultaneously. The outer problem is solved first, in ignorance of the layer; its wall pressure is then handed to the inner problem as a coefficient; and the inner problem is solved afterwards. The coupling is one-way at this order, and it is Equation (A.1356) — the pressure being constant across the layer — that makes it so.
∎The flat plate: similarity and the Blasius equation
For a flat plate aligned with a uniform stream the outer flow is undisturbed, \(U_{e}=U\) constant, and Equation (A.1363) gives \(\dd p/\dd x=0\). Restoring dimensions, Equation (36.26) becomes
with \(u=v=0\) at \(y=0\) for \(x>0\) and \(u\to U\) as \(y\to\infty\).
Put
with \(\psi\) the stream function Equation (31.12). Then Equation (A.1364) is satisfied identically in \(x\) if and only if \(f\) obeys the Blasius equation Equation (36.27),
Here \(\eta\) is dimensionless and \(\psi\) carries \(\mathrm{m}^{2}/\mathrm{s}\), so \(f\) is dimensionless too. Rests on Equation (A.1364), Equation (31.12) and Theorem A.798.
Derives Theorem A.800. Continuity is satisfied identically by the use of \(\psi\), with \(u=\pp_{y}\psi\) and \(v=-\pp_{x}\psi\). Compute the four derivatives that occur, using
the second because \(\eta\) carries \(x^{-1/2}\). Then
Now assemble the two convective terms. The first is
and the second, using \(\sqrt{\nu U/x}\,\sqrt{U/\nu x}=U/x\),
The terms in \(\eta f'f''\) cancel between Equations (A.1372) and (A.1373) — this is the cancellation on which the whole reduction turns — leaving
The viscous term is \(\nu\,\pp_{y}^{2}u=\nu U^{2}f'''/\nu x=\left(U^{2}/x\right)f'''\) by Equation (A.1371). Equating,
and the factor \(U^{2}/x\) divides out of both sides — every explicit occurrence of \(x\) has cancelled, which is precisely the statement that the ansatz Equation (A.1365) is a similarity solution. What remains is \(f'''=-\tfrac{1}{2}ff''\), that is Equation (A.1366).
The boundary conditions transcribe as follows. At \(y=0\), \(\eta=0\): \(u=0\) gives \(f'(0)=0\) by Equation (A.1368), and then \(v=0\) gives \(\eta f'-f=0\) at \(\eta=0\), hence \(f(0)=0\) by Equation (A.1369). As \(y\to\infty\), \(u\to U\) gives \(f'(\infty)=1\). Three conditions for a third-order equation, which is why the problem is well posed and why \(f''(0)\) is determined rather than free.
∎The three constants
With \(\mathrm{Re}_{x}=Ux/\nu\) and \(\mathrm{Re}_{L}=UL/\nu\), the solution of Equation (A.1366) gives
the last for one side of a plate of length \(L\) and unit span, \(F\) being in \(\mathrm{N}/\mathrm{m}\) of span. Rests on Equations (A.1347), (A.1366) and (A.1371).
Derives Proposition A.801. Thickness. By Equation (A.1368), \(u/U=f'(\eta)\), so the station at which the velocity has reached \(99\,\mathrm{\%}\) of the free stream is the station at which \(\eta=5.0\), by the second member of Equation (A.1347). Inverting Equation (A.1365),
Local friction. The wall stress is \(\tau_{w}=\mu\left(\pp_{y}u\right)_{y=0}\), in \(\mathrm{Pa}\), and by Equation (A.1371)
Dividing by the dynamic pressure \(\tfrac{1}{2}\rho U^{2}\) and using \(\mu=\rho\nu\),
which with \(f''(0)=0.3321\) is \(0.6642\), the \(0.664\) of Phenomenon 36.12.
Plate drag. Integrate Equation (A.1380) along the plate:
the integral converging at the leading edge despite the singularity of Equation (A.1380) there. Then
which is \(1.3284\). The relation \(C_{D}=2c_{f}(L)\) is exact and independent of the numerical value of \(f''(0)\): it is the statement that the mean of \(x^{-1/2}\) over \((0,L)\) is twice its endpoint value.
∎The momentum thickness Equation (36.25) of the Blasius profile is
so that Equation (36.25) returns exactly the drag Equation (A.1382) computed from the wall stress. Rests on Equation (36.25), Equation (A.1366) and Proposition A.801.
Derives Proposition A.802. The first member is Equation (A.1368) substituted into the definition of \(\vartheta\) in Equation (36.25), with \(\dd y=\sqrt{\nu x/U}\,\dd\eta\). For the identity, differentiate the product \(f\left(1-f'\right)\) and use the Blasius equation Equation (A.1366) in the form \(ff''=-2f'''\):
Integrate from \(0\) to \(\infty\). On the left, \(f(0)=0\) kills the lower limit, and at the upper limit \(1-f'\) vanishes exponentially while \(f\) grows only linearly, so the product vanishes; the left side is therefore zero. On the right the first term is \(I\) and the second is \(2\left[f''\right]_{0}^{\infty}=-2f''(0)\), since \(f''\to0\) with \(1-f'\). Hence \(I=2f''(0)\).
Substituting into Equation (36.25), \(F=\rho U^{2}\vartheta(L)=2f''(0)\rho U^{2}\sqrt{\nu L/U} =2\mu Uf''(0)\sqrt{UL/\nu}\), which is Equation (A.1382). The two routes to the drag — integrating the stress the fluid exerts on the plate, and traversing the wake far downstream — agree exactly, as Proposition 36.14 requires, and neither uses the other in its derivation.
∎A plate \(1\,\mathrm{m}\) long and of unit span, held edge-on in water of kinematic viscosity \(10^{-6}\,\mathrm{m}^{2}/\mathrm{s}\) and density \(10^{3}\,\mathrm{kg}/\mathrm{m}^{3}\), at \(1\,\mathrm{m}/\mathrm{s}\). Then \(\mathrm{Re}_{L}=10^{6}\), \(\mathrm{Re}_{L}^{-1/2}=10^{-3}\), and the dynamic pressure is \(\tfrac{1}{2}\rho U^{2}=500\,\mathrm{Pa}\), so
per side. The layer is half a percent of the plate, which is the number Phenomenon 36.12 quotes, and the drag on a square metre of wetted surface is under a newton — a useful calibration of how small viscous friction is, and of how completely it nevertheless governs the flow when it separates. Rests on Proposition A.801 and Phenomenon 36.12.
What these equations cannot describe
Transition. The solution above is laminar. The plate's own layer becomes turbulent at \(\mathrm{Re}_{x}\) of order \(5\times 10^{5}\), so in Example A.803 the laminar solution is honest over roughly the first half metre and the quoted whole-plate drag is the laminar figure carried past its own domain of validity. The chapter records the same transition for the pipe at Section 36.3, and the threshold is no better determined here than it is there.
Separation. Equation (A.1355) is
parabolic in \(x\): it is marched downstream from an initial
profile, and information travels only forwards. That is exactly what
fails at separation. As the wall stress approaches zero under a rising
external pressure, the marching solution develops a square-root
singularity in \(x\) at the station where \(\left(\pp_{y}u\right)_{y=0}\)
vanishes — Goldstein's singularity, established in 1948 — and the
computation cannot be continued through it. The layer equations
therefore cannot describe the separated region whose existence is the
principal observation of Phenomenon 36.12. This is why
Proposition 36.13 is proved in the chapter by an expansion
of the full equations at the wall, and not from
Equation (36.26): the criterion
Equation (36.24) has to be obtained by an argument that
survives at the point where these equations break down. Goldstein's
paper has no key in references.bib and the attribution here is
made in words.
Leading edge. Equation (A.1380) diverges as \(x\to0\), where \(\delta\) is comparable with \(x\) and the assumption \(\pp_{x}^{2}u\ll\pp_{y}^{2}u\) used in Theorem A.798 fails. The divergence is integrable, so Equation (A.1382) is finite and the drag coefficient is unaffected at leading order; but the stress distribution near the leading edge is not given correctly by this theory.
Prandtl's Boundary-Layer Equations and the Blasius Similarity Solution discharges the derivation owed at Proposition 36.15 of Experiment: Fluid Flow and Turbulence, in the boundary-layer experiment Section 36.5. Three things done here are used there and are worth carrying back. Theorem A.798 derives the exponent that Equation (36.23) estimated, so the chapter's resolution of the d'Alembert paradox rests on a reduction rather than on an order-of-magnitude argument. Corollary A.799 identifies what the outer flow gives the layer, which is the content of the chapter's remark that the pressure is “impressed” on the layer from outside and is the hypothesis under which Proposition 36.13 is stated. And Proposition A.802 closes the loop with Proposition 36.14: the drag measured by a traverse of the wake and the drag computed from the stress at the wall are the same number, exactly, for the one flow in which both can be evaluated in closed form.