Hilbert Spaces
The vector-space machinery of Linear Algebra and Representation Theory was developed largely without regard to dimension, but the passage to infinite dimension is not innocent: limits enter, and with them the topological and metric notions of Topological and Metric Spaces. A Hilbert space is the structure in which the algebra and the topology cooperate — an inner-product space whose associated metric is complete — and it is the arena in which the postulates of quantum mechanics will be formulated in The Postulates of Quantum Mechanics.
The subject grew out of Hilbert's spectral theory of integral equations [Hilbert:1912], was axiomatised by von Neumann for exactly that quantum-mechanical purpose [vonNeumann:1932], and its standard modern exposition is Reed–Simon [Reed:1972]; those three are the references of record for this chapter.
Three things are new here, and each is a place where the finite-dimensional intuition of Linear Algebra and Representation Theory fails. A subspace need not be closed, so it need not admit an orthogonal complement; an operator need not be defined on the whole space, and if it is symmetric and everywhere defined it is forced to be bounded; and a self-adjoint operator need have no eigenvector at all, so the decomposition of a state into eigenstates has to be replaced by an integral against a projection-valued measure. The chapter is organised around repairing those three failures.
Three classical results are used below and are not derived in this treatise; each is flagged where it is used and all three are in [Reed:1972]. (i) The convergence theorems of Lebesgue integration — monotone convergence and Fatou's lemma — which enter only through the completeness of \(L^{2}\) (Theorem 12.12); the treatise develops no measure theory, and Probability and Statistics makes the same declaration. (ii) The closed-graph theorem for Banach spaces, used once, in Theorem 12.74; it rests on the Baire category theorem, which is not proved here either. (iii) The Weierstrass approximation theorem, used once, to pin down the functional calculus in Proposition 12.61. Everything else in this chapter is proved from Linear Algebra and Representation Theory, Topological and Metric Spaces, Real Analysis and Complex Analysis.
The long derivations this chapter sends to Appendix A quote two further results — the Riesz–Markov representation theorem, in The Spectral Theorem for a Bounded Self-Adjoint Operator, and the Hilbert–Schmidt embedding property of a nuclear space, in The Nuclear Spectral Theorem of Gelfand and Maurin — together with the Radon–Nikodym theorem and the dominated convergence theorem, which belong to the same measure theory as (i). Each is stated as a displayed theorem where it is used, and each of those four appendix sections closes with a remark titled What is quoted here naming its imports exactly.
Definitions
Consider a pre-Hilbert space \(\mathcal{H}\) — a vector space equipped with an inner product — of infinite dimension, together with the norm associated with that inner product and the metric associated with that norm. We say that \(\mathcal{H}\) is a Hilbert space if and only if \(\mathcal{H}\) is complete, that is, if and only if every Cauchy sequence in \(\mathcal{H}\) converges. Rests on Definition 5.17, Definition 6.27 and Equation (5.44).
Nothing in Definition 12.2 is new except the word complete. The inner product is the map \(\braket{\ }{\ }\) of Definition 5.17, antilinear in its first slot and linear in its second (Equations (5.34) and (5.36)); the norm is \(\norm{v}=\sqrt{\braket{v}{v}}\) of Equation (5.44); the metric is \(d(u,v)=\norm{u-v}\) of Equation (5.45); and completeness is Definition 6.27 read in that metric. The infinite dimension is not a hypothesis of any theorem below — a finite-dimensional inner-product space is automatically complete — but it is the case of interest, and we shall say Hilbert space for any complete inner-product space and note the finite-dimensional degeneracies where they matter.
Throughout, \(\mathcal{H}\) is a complex Hilbert space unless the real case is named; \(\mathbb{K}=\C\), and \(a^{\ast}\) is the complex conjugate of \(a\in\C\). Vectors are written \(x,y,z\) or, where the quantum-mechanical reading is intended, \(\ket{\psi}\). The dagger \(A^{\dagger}\) denotes the adjoint of Theorem 12.38 throughout the treatise, never the transpose, which is written \(A\transpose\) (Linear Algebra and Representation Theory).
Three consequences of the axioms are used on nearly every page below: the Cauchy–Schwarz inequality, the continuity of the inner product that follows from it, and the parallelogram law, which is what completeness will be levered against in Section 12.2.1.
Let \(\mathcal{H}\) be an inner-product space. Then
and the inner product is jointly continuous: if \(u_{n}\to u\) and \(v_{n}\to v\) in norm, then \(\braket{u_{n}}{v_{n}}\to\braket{u}{v}\) in \(\C\). Rests on Proposition 5.19 and Definition 12.2.
Derives Proposition 12.4. Equation (12.1) is Equation (5.43) with both sides put under a square root through Equation (5.44). That inequality is available here without any further argument: Remark 5.20 records that its proof used only the five axioms of Definition 5.17 — neither a basis, nor finite dimension, nor completeness — so it holds verbatim in the present setting.
For continuity, a convergent sequence is bounded, say \(\norm{v_{n}}\leq C\) for all \(n\). Then, adding and subtracting \(\braket{u}{v_{n}}\) and using Equations (5.35) and (5.37),
by Equation (12.1) applied to each term, and both summands tend to zero.
∎The norm is continuous, \(\abs{\norm{u}-\norm{v}}\leq\norm{u-v}\); and if \(x_{n}\to x\) with \(\braket{y}{x_{n}}=0\) for all \(n\), then \(\braket{y}{x}=0\). Orthogonality to a fixed vector is therefore a closed condition. Rests on Proposition 12.4 and Definition 5.22.
Derives Corollary 12.5. The first is the triangle inequality Equation (5.42) used twice, \(\norm{u}\leq\norm{u-v}+\norm{v}\) and symmetrically. The second is Proposition 12.4 with the constant sequence \(y\) in the first slot.
∎For all \(u,v\) in an inner-product space,
and the inner product is recovered from the norm alone by
Rests on Definition 5.17 and Equation (5.44).
Derives Proposition 12.6. Expanding by Equations (5.35) and (5.37) and the scalar rules Equations (5.34) and (5.36), for any \(\lambda\in\C\),
Taking \(\lambda=+1\) and \(\lambda=-1\) and adding cancels the two cross terms and gives Equation (12.2).
For Equation (12.3), put \(\lambda=\ii^{k}\) in Equation (12.4), multiply by \(\ii^{-k}\) and sum over \(k=0,1,2,3\). The first two terms carry \(\sum_{k}\ii^{-k}=0\) and vanish. The third carries \(\sum_{k}\ii^{-k}\ii^{k}=4\) and leaves \(4\braket{u}{v}\). The fourth carries \(\sum_{k}\ii^{-k}\ii^{-k}=\sum_{k}(-1)^{k}=0\) and vanishes.
∎Equation (12.2) is a genuine test, not a mere identity: a norm satisfies it if and only if it is the norm associated with an inner product, and Equation (12.3) then constructs that inner product from the norm. The supremum norm on continuous functions fails it — take \(u,v\) two bumps with disjoint supports and equal heights, so that all four norms in Equation (12.2) equal that height and the identity reads \(2=4\) — which is why the space of continuous functions with the supremum norm is a Banach space and not a Hilbert space. The whole of the geometry below, orthogonality and projection included, is unavailable there.
A normed space is complete if and only if every series \(\sum_{n}x_{n}\) with \(\sum_{n}\norm{x_{n}}<\infty\) converges in it. Rests on Definitions 5.18 and 6.27.
Derives Proposition 12.8. (\(\Rightarrow\)) The partial sums \(S_{N}=\sum_{n\leq N}x_{n}\) satisfy \(\norm{S_{M}-S_{N}}\leq\sum_{n=N+1}^{M}\norm{x_{n}}\) for \(M>N\), which is a tail of a convergent series of non-negative reals and so tends to zero; \((S_{N})\) is Cauchy and converges by completeness.
(\(\Leftarrow\)) Let \((y_{n})\) be Cauchy. Choose \(n_{1}<n_{2}<\cdots\) with \(\norm{y_{m}-y_{n_{k}}}<2^{-k}\) for all \(m\geq n_{k}\). Then \(x_{1}=y_{n_{1}}\) and \(x_{k+1}=y_{n_{k+1}}-y_{n_{k}}\) have \(\sum_{k}\norm{x_{k}}\leq\norm{y_{n_{1}}}+\sum_{k}2^{-k}<\infty\), so by hypothesis the telescoping series converges, that is, \(y_{n_{k}}\to y\) for some \(y\). A Cauchy sequence with a convergent subsequence converges to the same limit, by the triangle inequality.
∎Two examples carry the whole subject, and Theorem 12.33 will show that they are one example seen twice.
Let
This is a Hilbert space. The series defining \(\braket{x}{y}\) converges absolutely, because \(\abs{x_{n}^{\ast}y_{n}}\leq\tfrac{1}{2}(\abs{x_{n}}^{2} +\abs{y_{n}}^{2})\); the five axioms of Definition 5.17 are inherited termwise from \(\C\); and \(\ell^{2}\) is a vector space because the same bound applied to \(\abs{x_{n}+y_{n}}^{2}\leq 2\abs{x_{n}}^{2}+2\abs{y_{n}}^{2}\) keeps the sum finite. Rests on Definitions 5.17 and 12.2.
The space Equation (12.5) is complete. Rests on Example 12.9 and Axiom 7.1.
Derives Proposition 12.10. Let \((x^{(m)})_{m}\) be Cauchy in \(\ell^{2}\). For each fixed \(n\), \(\abs{x^{(m)}_{n}-x^{(p)}_{n}}\leq\norm{x^{(m)}-x^{(p)}}\), so \((x^{(m)}_{n})_{m}\) is Cauchy in \(\C\) and converges to some \(x_{n}\) by the completeness of \(\R\) (Theorem 7.8, applied to real and imaginary parts). Given \(\varepsilon>0\) choose \(M\) with \(\norm{x^{(m)}-x^{(p)}}<\varepsilon\) for \(m,p\geq M\). For every \(N\) and every \(m,p\geq M\),
a finite sum, in which we may let \(p\to\infty\) termwise to get \(\sum_{n=1}^{N}\abs{x^{(m)}_{n}-x_{n}}^{2}\leq\varepsilon^{2}\); since \(N\) is arbitrary, \(\norm{x^{(m)}-x}\leq\varepsilon\). In particular \(x=x^{(M)}-(x^{(M)}-x)\) lies in \(\ell^{2}\), and \(x^{(m)}\to x\).
∎Let \(\Omega\subseteq\R^{N}\) be measurable and let \(L^{2}(\Omega)\) be the set of measurable \(f:\Omega\longrightarrow\C\) with \(\int_{\Omega}\abs{f}^{2}<\infty\), two functions being identified when they agree outside a set of measure zero, with
The identification is not a convenience but a necessity: without it \(\braket{f}{f}=0\) would not force \(f=0\), and Equation (5.31) would fail. An element of \(L^{2}\) is therefore an equivalence class, and a statement about “the value \(f(x_{0})\)” is meaningless in \(L^{2}\) — a point that returns in Section 12.6.3, where the delta functional is shown not to be an element of \(L^{2}\) at all. Rests on Definitions 5.17 and 12.2.
\(L^{2}(\Omega)\) is complete, and is therefore a Hilbert space. Rests on Example 12.11 and Proposition 12.8.
Derives Theorem 12.12. By Proposition 12.8 it suffices to sum an absolutely convergent series: let \(\sum_{k}\norm{f_{k}}=S<\infty\) and put \(g_{K}=\sum_{k\leq K}\abs{f_{k}}\), an increasing sequence of non-negative measurable functions with \(\norm{g_{K}}\leq S\) by the triangle inequality. By monotone convergence (Remark 12.1) \(g_{K}\uparrow g\) pointwise with \(\int\abs{g}^{2}\leq S^{2}\), so \(g\) is finite outside a null set. Where \(g\) is finite the series \(\sum_{k}f_{k}(x)\) converges absolutely in \(\C\); call its sum \(f(x)\), defined almost everywhere, and note \(\abs{f}\leq g\), so \(f\in L^{2}\). Finally \(\abs{f-\sum_{k\leq K}f_{k}}^{2}\leq g^{2}\) pointwise and tends to zero almost everywhere, so Fatou's lemma applied to \(g^{2}-\abs{f-\sum_{k\leq K}f_{k}}^{2}\) gives \(\norm{f-\sum_{k\leq K}f_{k}}\longrightarrow0\).
∎The theorem is the joint work of Riesz [Riesz:1907] and Fischer [Fischer:1907], published within weeks of one another in 1907, and it is what made \(L^{2}\) usable: before it, “convergence in the mean” had no limit to converge to. Theorem 12.33 will show that \(\ell^{2}\) and \(L^{2}\) are not two examples but one, which is the mathematical content of the equivalence of matrix and wave mechanics.
Orthonormal bases and separability
Orthogonality and projections
In finite dimension every subspace splits off an orthogonal complement, \(\mathbb{V}=\mathbb{S}\oplus\mathbb{S}^{\perp}\) (Corollary 5.27), and the proof simply orthonormalises a basis of \(\mathbb{S}\). In infinite dimension that argument stops at the first step, because \(\mathbb{S}\) need have no finite basis, and the conclusion itself becomes false for a general subspace: Remark 5.28 flagged the gap and reserved its repair for this section. The repair rests on completeness and on the parallelogram law, and it begins with a purely metric statement about convex sets.
A subset \(C\) of a vector space is convex if \(tx+(1-t)y\in C\) whenever \(x,y\in C\) and \(t\in[0,1]\). Every subspace, and every translate of a subspace, is convex. Rests on Definition 5.6.
Let \(C\subseteq\mathcal{H}\) be non-empty, closed and convex, and let \(x\in\mathcal{H}\). Then there is exactly one \(y\in C\) with
Rests on Definition 12.13, Proposition 12.6 and Definition 12.2.
Derives Theorem 12.14. Apply Equation (12.2) to the vectors \(u=x-z_{1}\) and \(v=x-z_{2}\), for \(z_{1},z_{2}\in C\). Then \(u+v=2x-(z_{1}+z_{2})\) and \(u-v=z_{2}-z_{1}\), so
The midpoint \(\tfrac{1}{2}(z_{1}+z_{2})\) lies in \(C\) by convexity, so the last norm is at least \(d\) and
Take a minimising sequence \((z_{n})\) in \(C\) with \(\norm{x-z_{n}}\to d\): the right-hand side, with \(z_{1}=z_{n}\) and \(z_{2}=z_{m}\), tends to \(2d^{2}+2d^{2}-4d^{2}=0\) as \(n,m\to\infty\), so \((z_{n})\) is Cauchy and converges, by Definition 12.2, to some \(y\), which lies in \(C\) because \(C\) is closed. Continuity of the norm (Corollary 12.5) gives \(\norm{x-y}=d\).
Uniqueness is the same estimate: if \(y\) and \(y'\) both realise the infimum, then \(\norm{y-y'}^{2}\leq2d^{2}+2d^{2}-4d^{2}=0\).
∎All three hypotheses are needed and each fails somewhere. Without closedness the infimum need not be attained (take \(C\) an open ball). Without convexity the minimiser need not be unique (take \(C\) a sphere and \(x\) its centre). Without completeness the Cauchy sequence has nowhere to go: on the pre-Hilbert space of continuous functions in \(L^{2}[-1,1]\), the closed convex set of functions with \(\int_{-1}^{0}f=\int_{0}^{1}f\) has no closest point to a suitable \(x\). It is completeness that makes Theorem 12.14 a theorem about Hilbert spaces rather than about inner-product spaces, and everything in this section inherits that.
For any subset \(S\subseteq\mathcal{H}\),
Rests on Definition 5.22 and Equation (5.52).
For any \(S\subseteq\mathcal{H}\), \(S^{\perp}\) is a closed subspace, \(S\subseteq S^{\perp\perp}\), and \(S^{\perp}=\overline{\gen{S}}^{\perp}\), where \(\overline{\gen{S}}\) is the closure of the set of finite linear combinations of elements of \(S\). Moreover \(S\cap S^{\perp}\subseteq\set{0}\). Rests on Definition 12.16 and Corollary 12.5.
Derives Proposition 12.17. \(S^{\perp}\) is a subspace by linearity of the bracket in its second slot, Equations (5.34) and (5.35), and it is closed as an intersection over \(x\in S\) of the sets \(\set{y\mid\braket{x}{y}=0}\), each closed by Corollary 12.5. If \(y\in S^{\perp}\) then \(y\) is orthogonal to every finite combination of elements of \(S\), again by linearity, and hence to every limit of such combinations, again by Corollary 12.5; the reverse inclusion is immediate from \(S\subseteq\overline{\gen{S}}\). That \(S\subseteq S^{\perp\perp}\) is a restatement of the definition read the other way round, using hermiticity Equation (5.33). Finally \(x\in S\cap S^{\perp}\) gives \(\braket{x}{x}=0\), so \(x=0\) by Equation (5.31).
∎Let \(M\subseteq\mathcal{H}\) be a closed subspace. Then
that is, every \(x\in\mathcal{H}\) has a unique decomposition \(x=x_{\parallel}+x_{\perp}\) with \(x_{\parallel}\in M\) and \(x_{\perp}\in M^{\perp}\); moreover \(x_{\parallel}\) is the closest point of \(M\) to \(x\), and \(\norm{x}^{2}=\norm{x_{\parallel}}^{2} +\norm{x_{\perp}}^{2}\). Rests on Theorem 12.14 and Proposition 12.17.
Derives Theorem 12.18. A subspace is convex and \(M\) is closed and contains \(0\), so Theorem 12.14 supplies a unique closest point \(x_{\parallel}\in M\); put \(x_{\perp}=x-x_{\parallel}\).
We claim \(x_{\perp}\in M^{\perp}\). Let \(m\in M\) with \(m\neq0\) and let \(t\in\C\). Since \(x_{\parallel}+tm\in M\), minimality gives \(\norm{x_{\perp}-tm}^{2}\geq\norm{x_{\perp}}^{2}\), while Equation (12.4) with \(\lambda=-t\) reads
Choose \(t=\braket{m}{x_{\perp}}/\norm{m}^{2}\). Using \(\braket{x_{\perp}}{m}=\braket{m}{x_{\perp}}^{\ast}\) (Equation (5.33)), each of the last three terms of Equation (12.11) equals \(\abs{\braket{m}{x_{\perp}}}^{2}/\norm{m}^{2}\) with signs \(+,-,-\), so
Comparison with the minimality inequality forces \(\braket{m}{x_{\perp}}=0\), and \(m\) was arbitrary.
Uniqueness: if \(x=m+w=m'+w'\) with \(m,m'\in M\) and \(w,w'\in M^{\perp}\), then \(m-m'=w'-w\) lies in \(M\cap M^{\perp}=\set{0}\) by Proposition 12.17. The norm identity is Equation (12.4) with \(\lambda=1\) and the cross terms vanishing — the Pythagorean theorem.
∎For a closed subspace \(M\), \(M^{\perp\perp}=M\). For an arbitrary subset \(S\), \(S^{\perp\perp}=\overline{\gen{S}}\). In particular a subspace \(D\subseteq\mathcal{H}\) is dense if and only if \(D^{\perp}=\set{0}\). Rests on Theorem 12.18 and Proposition 12.17.
Derives Corollary 12.19. \(M\subseteq M^{\perp\perp}\) always (Proposition 12.17). Conversely let \(x\in M^{\perp\perp}\) and split \(x=x_{\parallel}+x_{\perp}\) by Theorem 12.18. Then \(x_{\perp}=x-x_{\parallel}\in M^{\perp\perp}\), since \(M^{\perp\perp}\) is a subspace containing \(M\); but also \(x_{\perp}\in M^{\perp}\), and \(M^{\perp}\cap M^{\perp\perp}=\set{0}\). Hence \(x=x_{\parallel}\in M\).
For general \(S\), apply this to the closed subspace \(M=\overline{\gen{S}}\) and use \(S^{\perp}=M^{\perp}\) from Proposition 12.17. The density criterion is the case \(S=D\): \(D\) is dense exactly when \(\overline{\gen{D}}=\mathcal{H}\), i.e.\ when \(D^{\perp\perp}=\mathcal{H}\), i.e. when \(D^{\perp}=\set{0}\).
∎The density criterion of Corollary 12.19 is used more often below than the projection theorem itself: it converts every question of the form “is this family total?” into the computation of an orthogonal complement, and it is the standard way to test whether a domain is dense (Section 12.5.1).
For \(M\) a closed subspace, the map \(P_{M}:\mathcal{H} \longrightarrow\mathcal{H}\), \(P_{M}x=x_{\parallel}\), given by Theorem 12.18, is the orthogonal projection onto \(M\). Rests on Theorem 12.18.
\(P=P_{M}\) is linear and satisfies
with \(\norm{Px}\leq\norm{x}\), image \(M\) and kernel \(M^{\perp}\). Conversely, if \(P:\mathcal{H}\longrightarrow\mathcal{H}\) is linear and satisfies Equation (12.12), then its image is a closed subspace \(M\) and \(P=P_{M}\). Rests on Definition 12.20 and Theorem 12.18.
Derives Proposition 12.21. Linearity. If \(x=m+w\) and \(x'=m'+w'\) are the decompositions of Theorem 12.18, then \(ax+bx'=(am+bm')+(aw+bw')\) is a decomposition of \(ax+bx'\) of the required kind, and by uniqueness it is the one; so \(P(ax+bx')=aPx+bPx'\).
Idempotence and norm. \(Px\in M\), and a vector of \(M\) is its own closest point, so \(P^{2}=P\). The norm identity of Theorem 12.18 gives \(\norm{Px}\leq\norm{x}\). The image is \(M\) and \(Px=0\) exactly when \(x=x_{\perp}\in M^{\perp}\).
Symmetry. Write \(x=Px+(x-Px)\) and \(y=Py+(y-Py)\) with the second term of each in \(M^{\perp}\). Then \(\braket{Px}{y}=\braket{Px}{Py}\) and \(\braket{x}{Py}=\braket{Px}{Py}\), because in each case the omitted term pairs a vector of \(M\) with a vector of \(M^{\perp}\).
Converse. Let \(M=\im P\). If \(x\in M\), say \(x=Pz\), then \(Px=P^{2}z=Pz=x\), so \(M=\set{x\mid Px=x}\). This set is closed, and the argument needs no boundedness of \(P\): if \(x_{n}\in M\) and \(x_{n}\to x\), then for every \(y\) the symmetry in Equation (12.12) gives \(\braket{x_{n}}{Py-y}=\braket{Px_{n}}{y}-\braket{x_{n}}{y}=0\), whence \(\braket{x}{Py-y}=0\) by Proposition 12.4, and reading the symmetry backwards, \(\braket{Px-x}{y}=0\) for all \(y\), so \(Px=x\). For any \(x\) and any \(m\in M\), \(\braket{m}{x-Px}=\braket{Pm}{x-Px} =\braket{m}{Px-P^{2}x}=0\), using Equation (12.12) twice; so \(x-Px\in M^{\perp}\) and \(Px\) is the \(M\)-component of \(x\).
∎Equation (12.12) is the whole algebraic content of an orthogonal projection, and it is what makes projections the carriers of propositions in quantum mechanics: \(P^{2}=P\) says that asking the same question twice adds nothing, and the symmetry says the question has real answers. Two projections commute if and only if the corresponding subspaces decompose compatibly, and non-commuting projections are the formal seat of complementarity. The second identity of Equation (12.12) is exactly self-adjointness, \(P^{\dagger}=P\), once the adjoint is available (Theorem 12.38); it is written out here because the adjoint is defined two sections later and nothing above depends on it. The use made of all this is in The Postulates of Quantum Mechanics [vonNeumann:1932].
Gram–Schmidt orthogonalization
Let \(\set{v_{1},v_{2},\dots}\) be a finite or countable linearly independent family in an inner-product space. Then there is an orthonormal family \(\set{e_{1},e_{2},\dots}\) with
and consequently \(\overline{\gen{\set{e_{k}}}} =\overline{\gen{\set{v_{k}}}}\). Rests on Proposition 5.26 and Definition 5.25.
Derives Proposition 12.23. The recursion Equation (5.51) is finite at every step, so Proposition 5.26 constructs \(e_{1},\dots,e_{k}\) for each \(k\) and the construction at step \(k\) does not disturb its predecessors; Remark 5.28 already recorded that this much survives to a countable family. Only the last clause is new. By Equation (12.13) the two increasing unions \(\bigcup_{k}\gen{\set{e_{1},\dots,e_{k}}}\) and \(\bigcup_{k}\gen{\set{v_{1},\dots,v_{k}}}\) coincide, and each is exactly the set of finite linear combinations of the corresponding family; equal sets have equal closures.
∎Let \(\mathcal{H}\) be an inner-product space containing a countable dense subset. Then there is a finite or countable orthonormal family \(\set{e_{k}}\) with \(\overline{\gen{\set{e_{k}}}}=\mathcal{H}\). Rests on Proposition 12.23 and Definition 6.27.
Derives Corollary 12.24. Let \(\set{w_{1},w_{2},\dots}\) be dense. Discard \(w_{n}\) whenever it lies in \(\gen{\set{w_{1},\dots,w_{n-1}}}\); the surviving subfamily \(\set{v_{k}}\) is linearly independent and spans the same set of finite linear combinations, so \(\overline{\gen{\set{v_{k}}}}\supseteq\overline{\set{w_{n}}} =\mathcal{H}\). Now apply Proposition 12.23.
∎Take \(\mathcal{H}=L^{2}\bigl([-1,1]\bigr)\) and apply Proposition 12.23 to the monomials \(v_{k}(x)=x^{k-1}\), which are linearly independent because a non-trivial polynomial has finitely many roots and so cannot vanish almost everywhere. The output is, up to normalisation, the sequence of Legendre polynomials; running the same recursion in \(L^{2}\) of a weighted measure \(w(x)\,\dd x\) produces the Hermite, Laguerre and Chebyshev families according to the weight. That these are the eigenfunctions of the corresponding second-order differential operators is the content of Sturm–Liouville theory (Ordinary Differential Equations and Sturm–Liouville Theory); that the resulting family is total, so that the expansion of Section 12.2.3 converges to the function expanded, is a separate fact about each weight and is not supplied by Gram–Schmidt. Rests on Proposition 12.23 and Example 12.11.
The construction is due to Gram [Gram:1883], who introduced it for least-squares approximation by series of functions, and to Schmidt [Schmidt:1907], who put it in the form used here in the course of his work on integral equations.
Fourier expansion, Bessel and Parseval
A finite or countable family \(\set{e_{n}}\) in \(\mathcal{H}\) with \(\braket{e_{n}}{e_{m}}=\delta_{nm}\) is an orthonormal system. For \(x\in\mathcal{H}\) the scalars
are the Fourier coefficients of \(x\) in that system. Rests on Definition 5.25 and Equation (5.49).
Let \(\set{e_{n}}\) be an orthonormal system and \(x\in\mathcal{H}\). For every \(N\) the partial sum \(S_{N}=\sum_{n\leq N}c_{n}e_{n}\) is the orthogonal projection of \(x\) onto \(\gen{\set{e_{1},\dots,e_{N}}}\), hence the unique closest point of that subspace to \(x\), and
Rests on Definition 12.26 and Theorem 12.18.
Derives Proposition 12.27. Expanding with Equations (5.35) and (5.37) and using orthonormality,
since each of the three quantities \(\braket{x}{S_{N}}\), \(\braket{S_{N}}{x}\) and \(\norm{S_{N}}^{2}\) equals \(\sum_{n\leq N}\abs{c_{n}}^{2}\) — for the first, \(\braket{x}{S_{N}}=\sum_{n\leq N}c_{n}\braket{x}{e_{n}} =\sum_{n\leq N}c_{n}c_{n}^{\ast}\) by Equation (5.33), and the other two are the same computation — so that the two cross terms together cancel one copy of \(\norm{S_{N}}^{2}\). The same computation gives \(\braket{e_{m}}{x-S_{N}}=0\) for \(m\leq N\), so \(x-S_{N}\) is orthogonal to the finite-dimensional — hence closed — subspace \(\gen{\set{e_{1},\dots,e_{N}}}\), and \(S_{N}\) is its orthogonal projection of \(x\) by the uniqueness in Theorem 12.18. The left side of Equation (12.16) is non-negative, so \(\sum_{n\leq N}\abs{c_{n}}^{2}\leq\norm{x}^{2}\) for every \(N\); a bounded increasing sequence of partial sums converges, by Axiom 7.1, and its limit obeys the same bound.
∎Bessel's inequality says that the coefficients of any vector are square summable. The converse question — which square-summable coefficient sequences arise — has the best possible answer, and it is completeness that supplies it.
Let \(\set{e_{n}}\) be an orthonormal system in a Hilbert space and let \(c_{n}\in\C\). The series \(\sum_{n}c_{n}e_{n}\) converges in \(\mathcal{H}\) if and only if \(\sum_{n}\abs{c_{n}}^{2}<\infty\), and then \(\norm{\sum_{n}c_{n}e_{n}}^{2}=\sum_{n}\abs{c_{n}}^{2}\) and \(\braket{e_{m}}{\sum_{n}c_{n}e_{n}}=c_{m}\). Rests on Definitions 12.2 and 12.26.
Derives Proposition 12.28. By orthonormality and the Pythagorean identity, for \(M>N\),
So the partial sums of \(\sum_{n}c_{n}e_{n}\) form a Cauchy sequence in \(\mathcal{H}\) if and only if the partial sums of \(\sum_{n}\abs{c_{n}}^{2}\) form a Cauchy sequence in \(\R\). By Definition 12.2 and Theorem 7.8 respectively, each is equivalent to the convergence of the corresponding series. The norm identity follows by letting \(N=0\) and \(M\to\infty\) in Equation (12.17) and using Corollary 12.5; the coefficient identity follows from Proposition 12.4, which allows \(\braket{e_{m}}{\cdot}\) to be taken through the limit.
∎An orthonormal system \(\set{e_{n}}\) in \(\mathcal{H}\) is an orthonormal basis, or a complete orthonormal system, if it is maximal: the only vector orthogonal to every \(e_{n}\) is \(0\). Rests on Definitions 12.16 and 12.26.
Let \(\set{e_{n}}\) be an orthonormal system in a Hilbert space \(\mathcal{H}\). The following are equivalent.
-
\(\set{e_{n}}\) is maximal, i.e. an orthonormal basis (Definition 12.29);
-
the finite linear combinations of the \(e_{n}\) are dense in \(\mathcal{H}\);
-
every \(x\in\mathcal{H}\) is the sum of its Fourier series,
\begin{equation}\tag{12.18} x=\sum_{n}\braket{e_{n}}{x}\,e_{n}\ec \end{equation}the series converging in the norm of \(\mathcal{H}\);
-
Parseval's identity holds for every \(x\in\mathcal{H}\),
\begin{equation}\tag{12.19} \norm{x}^{2}=\sum_{n}\abs{\braket{e_{n}}{x}}^{2}\ep \end{equation}
Any of these implies the polarized form
Rests on Proposition 12.27, Proposition 12.28 and Corollary 12.19.
Derives Theorem 12.30. (1) \(\Rightarrow\) (3). By Equation (12.15) the coefficients \(c_{n}=\braket{e_{n}}{x}\) are square summable, so by Proposition 12.28 the series \(\sum_{n}c_{n}e_{n}\) converges to some \(y\in\mathcal{H}\) with \(\braket{e_{m}}{y}=c_{m}=\braket{e_{m}}{x}\) for every \(m\). Then \(x-y\) is orthogonal to every \(e_{m}\), hence \(x=y\) by maximality.
(3) \(\Rightarrow\) (2). The partial sums of Equation (12.18) are finite linear combinations converging to \(x\).
(2) \(\Rightarrow\) (1). If \(x\) is orthogonal to every \(e_{n}\) then \(x\in\set{e_{n}}^{\perp}=\overline{\gen{\set{e_{n}}}}^{\perp} =\mathcal{H}^{\perp}=\set{0}\), by Proposition 12.17 and density.
(3) \(\Rightarrow\) (4). Take norms in Equation (12.18) using the norm identity of Proposition 12.28.
(4) \(\Rightarrow\) (1). If \(x\perp e_{n}\) for every \(n\), then Equation (12.19) gives \(\norm{x}^{2}=0\).
Finally Equation (12.20) follows from Equation (12.19) by polarization: apply Equation (12.3) to both sides, each of the four terms of which is an instance of Equation (12.19). Alternatively, substitute Equation (12.18) for \(y\) and take \(\braket{x}{\cdot}\) through the sum by Proposition 12.4.
∎An orthonormal basis in the sense of Definition 12.29 is not a basis in the algebraic sense of Definition 5.14: the expansion Equation (12.18) is an infinite sum, a limit in the topology, and not a finite linear combination. The two notions coincide only in finite dimension. An infinite-dimensional Hilbert space does have algebraic bases, but they are uncountable and useless; the topological notion is the one physics means, and Linear Algebra and Representation Theory said as much when it reserved this section. The identity Equation (12.20), written in Dirac notation as \(\sum_{n}\ketbra{e_{n}}{e_{n}}=\identity\), is the completeness relation used constantly from The Postulates of Quantum Mechanics onwards.
\(\mathcal{H}\) is separable if it contains a countable dense subset. Rests on Definitions 6.27 and 12.2.
Let \(\mathcal{H}\) be a separable Hilbert space of infinite dimension. Then there is a linear bijection \(U:\mathcal{H}\longrightarrow\ell^{2}\) preserving the inner product, \(\braket{Ux}{Uy}_{\ell^{2}}=\braket{x}{y}\). Any two separable infinite-dimensional Hilbert spaces are therefore isomorphic; up to isomorphism there is only one. Rests on Corollary 12.24, Theorem 12.30 and Proposition 12.10.
Derives Theorem 12.33. By Corollary 12.24 there is a countable orthonormal family \(\set{e_{n}}\) whose finite combinations are dense, which is condition (2) of Theorem 12.30; the family is infinite because \(\mathcal{H}\) is. Define \(Ux=(\braket{e_{n}}{x})_{n}\). It lands in \(\ell^{2}\) by Equation (12.15), is linear because the bracket is linear in its second slot, and is isometric by Equation (12.19); an isometric linear map is injective. It is surjective because a sequence \((c_{n})\in\ell^{2}\) is, by Proposition 12.28, the coefficient sequence of the convergent sum \(\sum_{n}c_{n}e_{n}\). Preservation of the inner product is Equation (12.20), or equivalently isometry plus Equation (12.3). Composing the map for one space with the inverse of the map for another gives the last claim.
∎Theorem 12.33 is a mathematical triviality with a large physical consequence. Heisenberg's mechanics is written on \(\ell^{2}\) — infinite matrices acting on square-summable coefficient sequences — and Schrödinger's on \(L^{2}\), and the two were presented as rival theories. They are formulations on the same space: \(L^{2}\) is separable, so Theorem 12.33 produces a bracket-preserving bijection with \(\ell^{2}\), and the map is explicitly the assignment of Fourier coefficients in any orthonormal basis, Equation (12.14). This is the observation von Neumann made the organising principle of his treatise [vonNeumann:1932], and it is why The Postulates of Quantum Mechanics may state the postulates on an abstract \(\mathcal{H}\) without choosing a representation. The name attached to Equation (12.19) is Parseval's [Parseval:1806], who wrote the identity for trigonometric series more than a century before there was a space to write it on; the general form is inseparable from the Riesz–Fischer theorem [Riesz:1907] [Fischer:1907], of which it is the abstract shadow, and the trigonometric case is worked out in Fourier Analysis and Integral Transforms.
Bounded operators and duality
Bounded operators and their adjoints
In finite dimension every linear map is continuous, so Linear Algebra and Representation Theory never had to distinguish the algebra of linear transformations from the analysis of continuous ones. Here they part company, and the first order of business is to show that for linear maps the two words mean the same thing.
A linear map \(A:\mathcal{H}_{1}\longrightarrow\mathcal{H}_{2}\) is bounded if there is \(C\geq0\) with \(\norm{Ax}\leq C\norm{x}\) for all \(x\). The least such \(C\),
is the operator norm. The bounded operators from \(\mathcal{H}\) to itself form a set written \(\mathcal{B}(\mathcal{H})\). Rests on Definitions 5.18 and 5.35.
For a linear \(A:\mathcal{H}_{1}\longrightarrow\mathcal{H}_{2}\) the following are equivalent: (i) \(A\) is continuous at \(0\); (ii) \(A\) is continuous everywhere; (iii) \(A\) is bounded. Rests on Definition 12.35 and Proposition 6.28.
Derives Proposition 12.36. (iii) \(\Rightarrow\) (ii): \(\norm{Ax-Ax_{0}}=\norm{A(x-x_{0})} \leq\norm{A}\norm{x-x_{0}}\), so \(\delta=\varepsilon/(1+\norm{A})\) serves in Equation (6.11). (ii) \(\Rightarrow\) (i) is a special case. (i) \(\Rightarrow\) (iii): continuity at \(0\) with \(\varepsilon=1\) gives a \(\delta>0\) with \(\norm{Ax}\leq1\) whenever \(\norm{x}\leq\delta\). For \(x\neq0\) apply this to \(\delta x/\norm{x}\), which has norm \(\delta\): \(\norm{Ax}\leq\norm{x}/\delta\), so \(A\) is bounded with \(\norm{A}\leq1/\delta\).
∎Equation (12.21) is a norm on \(\mathcal{B}(\mathcal{H})\); it is submultiplicative,
the identity \(\identity\) has norm \(1\), and \(\mathcal{B}(\mathcal{H})\) is complete. Rests on Definition 12.35 and Proposition 12.8.
Derives Proposition 12.37. The norm axioms Equations (5.40), (5.41) and (5.42) follow pointwise from those of \(\mathcal{H}\) and the definition of a supremum; Equation (12.22) from \(\norm{ABx}\leq\norm{A}\norm{Bx}\leq\norm{A}\norm{B}\norm{x}\).
For completeness let \((A_{n})\) be Cauchy in \(\mathcal{B}(\mathcal{H})\). For each \(x\), \(\norm{A_{n}x-A_{m}x}\leq\norm{A_{n}-A_{m}}\norm{x}\), so \((A_{n}x)\) is Cauchy in \(\mathcal{H}\) and converges; call the limit \(Ax\). The map \(A\) is linear, being a pointwise limit of linear maps (Proposition 12.4 is not even needed: addition and scalar multiplication are continuous). Given \(\varepsilon>0\) pick \(N\) with \(\norm{A_{n}-A_{m}}\leq\varepsilon\) for \(n,m\geq N\); then \(\norm{A_{n}x-A_{m}x}\leq\varepsilon\norm{x}\), and letting \(m\to\infty\) with \(x\) fixed gives \(\norm{A_{n}x-Ax}\leq \varepsilon\norm{x}\). Hence \(A-A_{N}\) is bounded, so \(A\) is, and \(\norm{A_{n}-A}\leq\varepsilon\) for \(n\geq N\).
∎For every \(A\in\mathcal{B}(\mathcal{H})\) there is exactly one \(A^{\dagger}\in\mathcal{B}(\mathcal{H})\) with
and \(\norm{A^{\dagger}}=\norm{A}\). Rests on Definition 5.39, Theorem 12.46 and Proposition 12.37.
Derives Theorem 12.38. Fix \(x\). The map \(y\longmapsto\braket{x}{Ay}\) is linear, by Equations (5.34) and (5.35) and linearity of \(A\), and bounded, \(\abs{\braket{x}{Ay}}\leq\norm{x}\norm{A}\norm{y}\) by Equation (12.1). By the Riesz representation theorem (Theorem 12.46) there is a unique vector, call it \(A^{\dagger}x\), with \(\braket{A^{\dagger}x}{y}=\braket{x}{Ay}\) for all \(y\). No circularity is involved: Theorem 12.46 is proved in Section 12.3.2 from Theorem 12.18 alone and uses nothing from the present subsection. Uniqueness of \(A^{\dagger}x\) is uniqueness in Riesz's theorem; linearity of \(x\longmapsto A^{\dagger}x\) follows because \(\braket{A^{\dagger}(ax_{1}+bx_{2})-aA^{\dagger}x_{1} -bA^{\dagger}x_{2}}{y}=0\) for every \(y\), using Equations (5.36) and (5.37), so the vector in the first slot is \(0\) by Equation (5.31).
For the norm, Equation (12.23) with \(y=A^{\dagger}x\) gives
whence \(\norm{A^{\dagger}x}\leq\norm{A}\norm{x}\) and \(\norm{A^{\dagger}}\leq\norm{A}\). Reading Equation (12.23) backwards shows \((A^{\dagger})^{\dagger}=A\), and the same estimate then gives the reverse inequality.
∎Equation (12.23) is Equation (5.69) verbatim; what has changed is the proof. In finite dimension Proposition 5.40 built \(A^{\dagger}\) from an orthonormal basis and read off the conjugate transpose. Here there is no finite basis to write a matrix in, and the existence of \(A^{\dagger}\) is a theorem of analysis, resting on completeness through the projection theorem.
For \(A,B\in\mathcal{B}(\mathcal{H})\) and \(a\in\C\):
and
Moreover \(\ker A^{\dagger}=(\im A)^{\perp}\). Rests on Theorem 12.38 and Proposition 12.37.
Derives Proposition 12.39. Each identity in Equation (12.24) is verified by checking that both sides satisfy Equation (12.23) for the operator in question and invoking uniqueness; for the product, \(\braket{B^{\dagger}A^{\dagger}x}{y}=\braket{A^{\dagger}x}{By} =\braket{x}{ABy}\).
For Equation (12.25), one inequality is Equation (12.22) together with \(\norm{A^{\dagger}}=\norm{A}\). For the other, for any \(x\) with \(\norm{x}=1\),
using Equation (12.23) and Equation (12.1); taking the supremum over such \(x\) gives \(\norm{A}^{2}\leq\norm{A^{\dagger}A}\).
Finally \(x\in\ker A^{\dagger}\) means \(\braket{A^{\dagger}x}{y}=0\) for all \(y\), that is, \(\braket{x}{Ay}=0\) for all \(y\), that is, \(x\in(\im A)^{\perp}\).
∎A Banach algebra with an involution obeying Equation (12.24) and the norm identity Equation (12.25) is called a \(C^{\ast}\)-algebra, and Propositions 12.37 and 12.39 say that \(\mathcal{B}(\mathcal{H})\) is one. The identity is far stronger than the inequality \(\norm{A^{\dagger}A}\leq\norm{A}\norm{A^{\dagger}}\) that any Banach \(\ast\)-algebra has: it ties the norm to the algebra so tightly that a \(C^{\ast}\)-algebra admits exactly one norm making it one, so the analytic structure is determined by the algebraic. That rigidity is what the algebraic formulations of quantum theory (Axiomatic Quantum Field Theory) trade on. This treatise does not develop \(C^{\ast}\)-algebras; Equation (12.25) is proved here because it is used directly — it is what makes \(\norm{A^{2}}=\norm{A}^{2}\) for self-adjoint \(A\), and hence what controls the functional calculus of Section 12.4.2.
Let \(A\in\mathcal{B}(\mathcal{H})\). It is
-
self-adjoint (or hermitian) if \(A^{\dagger}=A\);
-
unitary if \(A^{\dagger}A=AA^{\dagger}=\identity\);
-
normal if \(A^{\dagger}A=AA^{\dagger}\);
-
positive, written \(A\geq0\), if \(\braket{x}{Ax}\geq0\) for every \(x\);
-
a projection if \(A^{2}=A=A^{\dagger}\) (Proposition 12.21);
-
compact if the image under \(A\) of the unit ball \(\set{x\mid\norm{x}\leq1}\) has compact closure (Definition 6.9).
Rests on Theorem 12.38, Definition 6.9 and Proposition 12.21.
(i) Self-adjoint and unitary operators are normal, and a projection is self-adjoint and positive. (ii) A unitary \(U\) preserves the inner product, \(\braket{Ux}{Uy}=\braket{x}{y}\), and \(\norm{U}=1\) when \(\mathcal{H}\neq\set{0}\). (iii) On a complex Hilbert space, \(A\) is self-adjoint if and only if \(\braket{x}{Ax}\in\R\) for every \(x\); in particular a positive operator is self-adjoint. (iv) \(A^{\dagger}A\) is positive for every \(A\). Rests on Definition 12.41 and Proposition 12.39.
Derives Proposition 12.42. (i) and (iv) are immediate from the definitions, (iv) because \(\braket{x}{A^{\dagger}Ax}=\norm{Ax}^{2}\geq0\). (ii) \(\braket{Ux}{Uy}=\braket{U^{\dagger}Ux}{y}=\braket{x}{y}\), and \(y=x\) gives \(\norm{Ux}=\norm{x}\).
(iii) If \(A^{\dagger}=A\) then \(\braket{x}{Ax}=\braket{Ax}{x}=\braket{x}{Ax}^{\ast}\) by Equation (5.33), so the number is real. Conversely suppose \(\braket{x}{Ax}\) is real for every \(x\), and put \(B=A-A^{\dagger}\), so that \(\braket{x}{Bx}=\braket{x}{Ax} -\braket{Ax}{x}=0\) for every \(x\). Replacing \(x\) by \(x+y\) and by \(x+\ii y\) and expanding as in Equation (12.4) gives \(\braket{x}{By}+\braket{y}{Bx}=0\) and \(\braket{x}{By}-\braket{y}{Bx}=0\), whence \(\braket{x}{By}=0\) for all \(x,y\) and \(B=0\). (The complex field is essential: on a real Hilbert space a rotation by a quarter turn in a plane has \(\braket{x}{Ax}=0\) identically and is not self-adjoint.)
∎If \(A\in\mathcal{B}(\mathcal{H})\) is self-adjoint then
In particular a self-adjoint operator with \(\braket{x}{Ax}=0\) for all \(x\) is \(0\). Rests on Propositions 12.39 and 12.42.
Derives Proposition 12.43. Write \(M\) for the supremum. That \(M\leq\norm{A}\) is Equation (12.1). For the converse, note first that \(\abs{\braket{z}{Az}}\leq M\norm{z}^{2}\) for every \(z\), by homogeneity. Since \(A\) is self-adjoint, \(\braket{y}{Ax}=\braket{Ay}{x} =\braket{x}{Ay}^{\ast}\), so
Bounding each term on the left and using Equation (12.2),
Take \(\norm{x}=\norm{y}=1\), so \(\abs{\operatorname{Re}\braket{x}{Ay}}\leq M\); replacing \(x\) by \(\ee^{\ii\theta}x\) for a suitable \(\theta\) makes \(\braket{x}{Ay}\) real and non-negative, giving \(\abs{\braket{x}{Ay}}\leq M\). Now take the supremum over unit \(x\): if \(Ay\neq0\) the choice \(x=Ay/\norm{Ay}\) gives \(\abs{\braket{x}{Ay}}=\norm{Ay}\), so \(\norm{Ay}\leq M\), and if \(Ay=0\) that bound is trivial. Hence \(\norm{Ay}\leq M\) for every unit \(y\), i.e. \(\norm{A}\leq M\).
∎Let \(A\in\mathcal{B}(\mathcal{H})\) be compact and self-adjoint on a Hilbert space \(\mathcal{H}\neq\set{0}\). Then \(\mathcal{H}\) has an orthonormal basis of eigenvectors of \(A\): there are real \(\lambda_{1},\lambda_{2},\dots\) with \(\lambda_{n}\to0\) and an orthonormal system \(\set{e_{n}}\) with
the series converging in norm. Every non-zero eigenvalue has finite-dimensional eigenspace, and \(\norm{A}=\max_{n}\abs{\lambda_{n}}\). Rests on Definition 12.41, Proposition 12.43 and Theorem 12.30.
The derivation is carried out in The Hilbert–Schmidt Theorem for Compact Self-Adjoint Operators. The route is: the supremum in the numerical-radius formula Equation (12.26) is attained on a compact operator, which produces a first eigenvector; its orthogonal complement is invariant, and the argument recurses; the eigenvalues tend to zero because an orthonormal sequence of eigenvectors with eigenvalues bounded away from \(0\) has an image with no convergent subsequence. The attainment step is the one that needs compactness, and the appendix carries it out without the weak topology, which the standard treatment [Reed:1972] uses: self-adjointness alone forces a maximising sequence to satisfy \(Ax_{n}-\lambda x_{n}\to0\) in norm, after which the sequential form of compactness delivers a genuine eigenvector.
Full derivation in Appendix A.
Derives Theorem 12.44.
Equation (12.27) is the exact infinite-dimensional analogue of the finite-dimensional spectral theorem Equation (5.105), and it is the only case in which the analogy is that close — the eigenvectors really do form a basis. It is also the historical starting point: this is the theorem Hilbert proved for the symmetric integral kernels of his spectral theory [Hilbert:1912] and that Schmidt recast in the language used here [Schmidt:1907]. The kernels concerned,
define compact operators on \(L^{2}(\Omega)\), and \(K\) is self-adjoint exactly when \(K(y,x)=K(x,y)^{\ast}\). The Green functions of Partial Differential Equations are of this form, which is why a boundary-value problem has a discrete spectrum of modes while the free operators of Section 12.4.1 need not.
The Riesz representation theorem
A continuous linear functional on \(\mathcal{H}\) is a bounded linear map \(f:\mathcal{H}\longrightarrow\C\). The set of them is the topological dual \(\mathcal{H}^{\ast}\), normed by Equation (12.21). Rests on Definitions 5.44 and 12.35.
For every \(f\in\mathcal{H}^{\ast}\) there is exactly one \(z\in\mathcal{H}\) with
and \(\norm{f}=\norm{z}\). Rests on Theorem 12.18 and Definition 12.45.
Derives Theorem 12.46. If \(f=0\) take \(z=0\); uniqueness is proved below. Otherwise let \(N=\ker f\), a subspace, closed because \(f\) is continuous and \(N=f^{-1}(\set{0})\) is the preimage of a closed set (Definition 6.6), and proper because \(f\neq0\). By Theorem 12.18 applied to \(N\), and by \(N^{\perp}\neq\set{0}\) — which holds because otherwise \(\mathcal{H}=N\oplus\set{0}=N\) — we may choose \(u\in N^{\perp}\) with \(\norm{u}=1\); note \(f(u)\neq0\), since \(u\notin N\).
For arbitrary \(x\), the vector \(f(x)\,u-f(u)\,x\) is annihilated by \(f\), by linearity, so it lies in \(N\) and is orthogonal to \(u\):
Hence \(f(x)=f(u)\braket{u}{x}=\braket{f(u)^{\ast}u}{x}\) by Equation (5.36), and \(z=f(u)^{\ast}u\) represents \(f\).
Uniqueness. If \(\braket{z}{x}=\braket{z'}{x}\) for all \(x\), take \(x=z-z'\) and use Equation (5.31). Norm. \(\abs{f(x)}\leq\norm{z}\norm{x}\) by Equation (12.1), so \(\norm{f}\leq\norm{z}\); and \(f(z)=\norm{z}^{2}\) gives \(\norm{f}\geq\norm{z}\).
∎The map \(J:\mathcal{H}\longrightarrow\mathcal{H}^{\ast}\), \(Jz=\braket{z}{\cdot}\), is a bijection, is isometric, and is antilinear: \(J(az_{1}+bz_{2})=a^{\ast}Jz_{1}+b^{\ast}Jz_{2}\). Rests on Theorem 12.46 and Definition 12.45.
Derives Corollary 12.47. Surjectivity and injectivity are the existence and uniqueness in Theorem 12.46, isometry is its last clause, and antilinearity is Equation (5.36).
∎Corollary 12.47 is the whole content of Dirac's notation [Dirac:1958]. The edition cited is the fourth, and deliberately so: the bra–ket split is not in the first edition of the Principles of 1930 [Dirac:1930b], which is the source of the delta-function calculus used later in this chapter and of nothing else here. Dirac introduced the split in a note of 1939 [Dirac:1939], and brought it into the book at the third edition of 1947. Writing the vector \(z\) as \(\ket{z}\) and the functional \(Jz\) as \(\bra{z}\), the value of the functional on a vector is \(\braket{z}{x}\) — the bracket split into a bra and a ket, which is where the words come from. The antilinearity of \(J\) is the rule that scalars are conjugated on passing the vertical bar, \(\bra{az}=a^{\ast}\bra{z}\), and it is not a convention but a theorem: it is forced by Equation (5.36). In finite dimension the same identification was made in Proposition 5.59, where it needed only a basis; the point of Theorem 12.46 is that it survives without one.
The identification is special to Hilbert spaces and is exactly what fails elsewhere. The dual of a Banach space is in general a different space, and even for a pre-Hilbert space that is not complete the theorem is false — completeness enters through Theorem 12.18. It fails again, in the opposite direction, when one insists on writing \(\ket{x}\) and \(\ket{p}\) for position and momentum “eigenstates”: those are functionals on a smaller space than \(\mathcal{H}\) and correspond to no vector of \(\mathcal{H}\) at all. That is the subject of Section 12.6.3, and it is the one place in this treatise where the bra–ket notation has to be read with care.
Self-adjointness and the spectral theorem
Spectrum and resolvent
Let \(A\in\mathcal{B}(\mathcal{H})\). The resolvent set is
the map \(R(\lambda)=(A-\lambda\identity)^{-1}\) on \(\rho(A)\) is the resolvent, and the spectrum is \(\sigma(A)=\C\setminus\rho(A)\). Rests on Definitions 5.45 and 12.35.
For \(\lambda\in\sigma(A)\) write \(T=A-\lambda\identity\). Then \(\lambda\) belongs to
-
the point spectrum \(\sigma_{p}(A)\) if \(T\) is not injective; \(\lambda\) is then an eigenvalue and \(\ker T\) the eigenspace;
-
the continuous spectrum \(\sigma_{c}(A)\) if \(T\) is injective, \(\im T\) is dense, and the inverse map \(T^{-1}\), defined on \(\im T\), is unbounded;
-
the residual spectrum \(\sigma_{r}(A)\) if \(T\) is injective and \(\im T\) is not dense.
Rests on Definition 12.49 and Corollary 12.19.
\(\sigma(A)\) is the disjoint union of \(\sigma_{p}(A)\), \(\sigma_{c}(A)\) and \(\sigma_{r}(A)\). Rests on Definitions 12.2 and 12.50.
Derives Proposition 12.51. Exclusivity is clear from the statements. For exhaustiveness, suppose \(\lambda\) belongs to none of the three: then \(T\) is injective, \(\im T\) is dense, and \(T^{-1}\) is bounded on \(\im T\), say \(\norm{T^{-1}y}\leq c\norm{y}\). Then \(\im T\) is closed: if \(Tx_{n}\to y\), the sequence \((Tx_{n})\) is Cauchy, so \(\norm{x_{n}-x_{m}}\leq c\norm{Tx_{n}-Tx_{m}}\) shows \((x_{n})\) is Cauchy, and \(x_{n}\to x\) for some \(x\) by Definition 12.2, with \(Tx=y\) because \(T\) is continuous. A dense closed subspace is the whole space, so \(\im T=\mathcal{H}\) and \(T^{-1}\in\mathcal{B}(\mathcal{H})\), i.e. \(\lambda\in\rho(A)\).
∎In finite dimension injectivity and surjectivity coincide (Proposition 5.47) and only \(\sigma_{p}\) survives; Remark 5.48 already recorded that this is the identity that infinite dimension destroys.
If \(\abs{\lambda}>\norm{A}\) then \(\lambda\in\rho(A)\) and
Hence \(\sigma(A)\) is contained in the closed disc of radius \(\norm{A}\) about the origin, and \(\norm{R(\lambda)}\to0\) as \(\abs{\lambda}\to\infty\). Rests on Proposition 12.37 and Definition 12.49.
Derives Proposition 12.52. By Equation (12.22), \(\norm{A^{n}/\lambda^{n+1}}\leq\norm{A}^{n}/\abs{\lambda}^{n+1}\), a geometric series of ratio \(\norm{A}/\abs{\lambda}<1\) (Proposition 7.46), so the series converges absolutely and hence in \(\mathcal{B}(\mathcal{H})\), by Propositions 12.8 and 12.37; its sum \(S\) obeys the stated norm bound. Multiplying the partial sums by \(A-\lambda\identity\) telescopes to \(-\identity+A^{N+1}/\lambda^{N+1}\), whose second term tends to \(0\); by continuity of multiplication in a Banach algebra, \((A-\lambda\identity)S=S(A-\lambda\identity)=-\identity\). Hence \(A-\lambda\identity\) is invertible with \(R(\lambda)=(A-\lambda\identity)^{-1}=-S\), which is Equation (12.31), and \(\norm{R(\lambda)}=\norm{S}\) obeys the bound already established.
∎\(\rho(A)\) is open, and for \(x,y\in\mathcal{H}\) the scalar function \(\lambda\longmapsto\braket{x}{R(\lambda)y}\) is holomorphic on \(\rho(A)\). Rests on Proposition 12.52 and Definition 8.5.
Derives Proposition 12.53. Let \(\lambda_{0}\in\rho(A)\) and write \(A-\lambda\identity=(A-\lambda_{0}\identity) \bigl[\identity-(\lambda-\lambda_{0})R(\lambda_{0})\bigr]\). For \(\abs{\lambda-\lambda_{0}}<1/\norm{R(\lambda_{0})}\) the bracket is invertible by the same geometric-series argument as in Proposition 12.52, so \(\lambda\in\rho(A)\) and
Pairing Equation (12.32) with \(x\) and \(y\) and using Proposition 12.4 to take the bracket through the sum exhibits \(\braket{x}{R(\lambda)y}\) as a convergent power series about \(\lambda_{0}\), hence holomorphic there (Theorem 7.50, and Complex Analysis for the complex statement).
∎For \(A\in\mathcal{B}(\mathcal{H})\) with \(\mathcal{H}\neq\set{0}\), the set \(\sigma(A)\) is compact and non-empty. Rests on Proposition 12.53, Theorem 8.18 and Proposition 12.52.
Derives Theorem 12.54. \(\sigma(A)\) is closed because \(\rho(A)\) is open (Proposition 12.53) and bounded by Proposition 12.52; a closed bounded subset of \(\C\), identified with \(\R^{2}\) as a metric space, is compact by the Heine–Borel theorem in \(\R^{N}\) (Section 6.1.3, where the one-dimensional case Theorem 6.11 is proved and the \(N\)-dimensional one is recorded as owed).
Suppose \(\sigma(A)=\varnothing\). Then for every \(x,y\) the function \(g(\lambda)=\braket{x}{R(\lambda)y}\) is entire, by Proposition 12.53, and by the norm bound in Equation (12.31) together with Equation (12.1) it satisfies \(\abs{g(\lambda)}\leq\norm{x}\norm{y}/(\abs{\lambda}-\norm{A})\) for \(\abs{\lambda}>\norm{A}\), so \(g\) is bounded on \(\C\) and tends to \(0\) at infinity. By Liouville's theorem (Theorem 8.18) \(g\) is constant, and the constant is \(0\). Since \(x\) and \(y\) were arbitrary, \(R(\lambda)=0\) for every \(\lambda\) — impossible, because \(R(\lambda)(A-\lambda\identity)=\identity\neq0\) on \(\mathcal{H}\neq\set{0}\).
∎Let \(A\in\mathcal{B}(\mathcal{H})\) be self-adjoint. Then \(\sigma(A)\subseteq\R\), the residual spectrum is empty, eigenvalues are real, and eigenvectors belonging to distinct eigenvalues are orthogonal. Rests on Definition 12.41, Theorem 12.54 and Corollary 12.19.
Derives Theorem 12.55. Let \(\lambda=a+\ii b\) with \(a,b\in\R\) and \(b\neq0\), and put \(B=A-a\identity\), which is self-adjoint. Since \(\braket{x}{Bx}\) is real (Proposition 12.42), expanding as in Equation (12.4) with \(\lambda=-\ii b\) makes the two cross terms \(-\ii b\braket{Bx}{x}\) and \(+\ii b\braket{x}{Bx}\) cancel, leaving
So \(A-\lambda\identity\) is injective with \(\norm{(A-\lambda\identity)^{-1}}\leq1/\abs{b}\) on its image, and that image is dense: by Equation (12.24) the adjoint of \(A-\lambda\identity\) is \(A-\lambda^{\ast}\identity\), so
by Proposition 12.39 and by Equation (12.33) applied to \(\lambda^{\ast}\), which has the same \(\abs{b}\); density then follows from Corollary 12.19. So \(A-\lambda\identity\) is injective with dense image and bounded inverse on it, which by the argument of Proposition 12.51 puts \(\lambda\) in \(\rho(A)\). Hence \(\sigma(A)\subseteq\R\).
The same computation for real \(\lambda\) shows \(\bigl(\im(A-\lambda\identity)\bigr)^{\perp} =\ker(A-\lambda\identity)\), so if \(A-\lambda\identity\) is injective its image is dense: the residual spectrum is empty. If \(Ax=\lambda x\) with \(x\neq0\) then \(\lambda\norm{x}^{2}=\braket{x}{Ax}\) is real; and if \(Ax=\lambda x\), \(Ay=\mu y\) with \(\lambda\neq\mu\) real, then \(\lambda\braket{x}{y}=\braket{Ax}{y}=\braket{x}{Ay}=\mu\braket{x}{y}\), forcing \(\braket{x}{y}=0\).
∎On \(\mathcal{H}=L^{2}\bigl([0,1]\bigr)\) define
Then \(Q\) is bounded and self-adjoint with \(\norm{Q}=1\), and
The operator has a full interval of spectrum and not one eigenvector. Rests on Definition 12.50, Theorem 12.55 and Example 12.11.
Derives Example 12.56. Bounded and self-adjoint. \(\abs{tf(t)}\leq\abs{f(t)}\) on \([0,1]\) gives \(\norm{Qf}\leq\norm{f}\); and since \(t\) is real, \(\braket{f}{Qg}=\int_{0}^{1}f(t)^{\ast}t\,g(t)\,\dd t =\braket{Qf}{g}\), which is Equation (12.23).
No eigenvectors. If \(Qf=\lambda f\) then \((t-\lambda)f(t)=0\) for almost every \(t\), so \(f\) vanishes almost everywhere on \([0,1]\setminus\set{\lambda}\), a set whose complement in \([0,1]\) has measure zero; hence \(f=0\) as an element of \(L^{2}\) (Example 12.11). This is exactly where the identification of functions agreeing almost everywhere does its work.
\(\sigma(Q)\subseteq[0,1]\). For \(\lambda\notin[0,1]\) the function \(t\longmapsto(t-\lambda)^{-1}\) is continuous on the compact set \([0,1]\), hence bounded there by Theorem 7.24, so multiplication by it is a bounded operator and is a two-sided inverse of \(Q-\lambda\identity\).
\([0,1]\subseteq\sigma(Q)\). Fix \(\lambda\in[0,1]\) and let \(I_{n}\) be the intersection of \([0,1]\) with the interval of half-width \(1/n\) about \(\lambda\), of length \(\ell_{n}\geq1/n\). Put \(f_{n}=\ell_{n}^{-1/2}\) on \(I_{n}\) and \(0\) elsewhere, so \(\norm{f_{n}}=1\) while
If \(Q-\lambda\identity\) had a bounded inverse \(R\) we would get \(1=\norm{f_{n}}\leq\norm{R}\,\norm{(Q-\lambda)f_{n}}\leq\norm{R}/n\) for every \(n\), which is false. Hence \(\lambda\in\sigma(Q)\), and \(\norm{Q}=1\) follows from Proposition 12.52 and \(1\in\sigma(Q)\). By Theorem 12.55 the residual spectrum is empty and by the previous paragraph the point spectrum is, so all of \([0,1]\) is continuous spectrum.
∎Example 12.56 is the operator Remark 5.79 promised, and it settles the shape of the rest of the chapter. Every statement of Theorem 5.75 that mentions eigenvectors is simply false for \(Q\): there is no orthonormal basis of eigenvectors, no diagonal matrix Equation (5.105), no decomposition of \(\mathcal{H}\) into eigenspaces. What survives is the real spectrum, and what replaces the eigenbasis is an integral over that spectrum against a projection-valued measure — for \(Q\) itself, the measure that assigns to a Borel set \(\Omega\subseteq[0,1]\) the projection “multiply by the indicator of \(\Omega\)”. Physically, the difference is the difference between a bound state, which is a genuine eigenvector with a normalisable wavefunction, and a scattering state, which is not; the two coexist in the hydrogen spectrum of The Hydrogen Atom and are the whole subject of Scattering Theory. It is also why a position “eigenstate” \(\ket{x_{0}}\) cannot be a vector of \(L^{2}\) — Section 12.6.3 says what it is instead.
The spectral theorem
A projection-valued measure on \(\R\), acting on \(\mathcal{H}\), is a map \(E\) assigning to every Borel set \(\Omega\subseteq\R\) a projection \(E(\Omega)\in\mathcal{B}(\mathcal{H})\) (Definition 12.41) such that
-
\(E(\varnothing)=0\) and \(E(\R)=\identity\);
-
\(E(\Omega_{1}\cap\Omega_{2})=E(\Omega_{1})E(\Omega_{2})\);
-
for pairwise disjoint \(\Omega_{1},\Omega_{2},\dots\) and every \(x\in\mathcal{H}\),
\begin{equation}\tag{12.36} E\Bigl(\bigcup_{n}\Omega_{n}\Bigr)x =\sum_{n}E(\Omega_{n})\,x\ec \end{equation}the series converging in the norm of \(\mathcal{H}\).
For \(x,y\in\mathcal{H}\) the complex measure \(\Omega\longmapsto\braket{x}{E(\Omega)y}\) is written \(\braket{x}{\dd E(\lambda)y}\); taking \(y=x\) gives a positive measure of total mass \(\norm{x}^{2}\). Rests on Definition 12.41 and Proposition 12.21.
The last remark is the reason the definition is shaped this way: for a unit vector \(x\), \(\Omega\longmapsto\braket{x}{E(\Omega)x} =\norm{E(\Omega)x}^{2}\) is a probability measure on \(\R\). That is already the Born rule in embryo, before any physics has been assumed (The Postulates of Quantum Mechanics).
Let \(A\in\mathcal{B}(\mathcal{H})\) be self-adjoint. Then there is exactly one projection-valued measure \(E\) on \(\R\), supported on \(\sigma(A)\), such that
written \(A=\int_{\sigma(A)}\lambda\,\dd E(\lambda)\). Every bounded operator commuting with \(A\) commutes with every \(E(\Omega)\).
Equivalently, in the multiplication form: if \(\mathcal{H}\) is separable there is a measure space \((M,\mu)\) with \(\mu(M)<\infty\), a bounded real measurable function \(a\) on \(M\), and a unitary \(U:\mathcal{H}\longrightarrow L^{2}(M,\mu)\) with
Every bounded self-adjoint operator is a multiplication operator in disguise. Rests on Definition 12.58, Theorem 12.55 and Proposition 12.43.
Both halves are derived in The Spectral Theorem for a Bounded Self-Adjoint Operator. The route is the standard one: the continuous functional calculus is built first, from the fact that for a self-adjoint operator the norm of a polynomial in it equals the supremum of that polynomial on the spectrum; a complex measure is then obtained for each pair of vectors by the Riesz–Markov representation theorem; the calculus is extended to bounded Borel functions, and the projections \(E(\Omega)\) are the values it takes on indicator functions. The multiplication form follows by decomposing \(\mathcal{H}\) into cyclic subspaces, on each of which the calculus itself provides the unitary. Two classical theorems are used there and are not proved in this treatise — Riesz–Markov, which is measure theory, and Stone–Weierstrass, of which only the Weierstrass approximation theorem already quoted in Remark 12.1 is needed; the appendix states each precisely where it is used and collects both in a remark, and nothing else is assumed. The original is von Neumann's memoir [vonNeumann:1930]; the modern account followed there is [Reed:1972], chapters VII and VIII.
Full derivation in Appendix A.
Derives Theorem 12.59.
The uniqueness of the functional calculus is worth isolating from that argument, and it costs only a paragraph, so it is proved here: it is what makes “the” functions of an observable well defined — without it the expression \(f(A)\) would depend on how it was constructed.
Given \(E\) as in Theorem 12.59 and a bounded Borel function \(f:\sigma(A)\longrightarrow\C\), the operator \(f(A)\) is defined by
Rests on Theorem 12.59 and Definition 12.58.
Let \(A\in\mathcal{B}(\mathcal{H})\) be self-adjoint and let \(\Phi_{1},\Phi_{2}\) be maps from the continuous functions on \(\sigma(A)\) to \(\mathcal{B}(\mathcal{H})\) such that each is linear and multiplicative, sends the constant function \(1\) to \(\identity\) and the function \(\lambda\longmapsto\lambda\) to \(A\), and is continuous for the supremum norm on the domain and the operator norm on the target. Then \(\Phi_{1}=\Phi_{2}\). Rests on Definition 12.60 and Proposition 12.37.
Derives Proposition 12.61. Linearity and multiplicativity, together with the two normalisations, force \(\Phi_{i}(p)=p(A)\) for every polynomial \(p\), where \(p(A)\) is the algebraic expression: the monomial \(\lambda^{n}\) goes to \(A^{n}\) by multiplicativity and induction, and linear combinations follow. So \(\Phi_{1}\) and \(\Phi_{2}\) agree on polynomials. By the Weierstrass approximation theorem (Remark 12.1) the polynomials are dense, in the supremum norm, in the continuous functions on the compact set \(\sigma(A)\) (Theorem 12.54). Two continuous maps agreeing on a dense subset of a metric space agree everywhere.
∎In finite dimension Theorem 12.59 is Theorem 5.75, already proved, and the translation is worth carrying out once because it fixes the meaning of every symbol above. Let \(\dim\mathcal{H}=n<\infty\) and let \(A\) be self-adjoint, with distinct eigenvalues \(\lambda_{1},\dots,\lambda_{r}\) and \(P_{i}\) the orthogonal projection onto the eigenspace of \(\lambda_{i}\). Then Theorem 5.75 says
and the projection-valued measure is \(E(\Omega)=\sum_{\lambda_{i}\in\Omega}P_{i}\), which manifestly satisfies the three conditions of Definition 12.58; the integral Equation (12.37) is then a finite sum and reproduces Equation (12.40). The functional calculus Equation (12.39) becomes \(f(A)=\sum_{i}f(\lambda_{i})P_{i}\), which is the familiar rule that a function of a diagonalizable operator acts by that function on each eigenvalue. What the general theorem does is exactly to replace the finite sum over eigenvalues by an integral over the spectrum — nothing more, and nothing less. Example 12.56 shows why nothing less will do.
Three of the quantum postulates of The Postulates of Quantum Mechanics are statements about Theorem 12.59. That an observable is represented by a self-adjoint operator is the demand that the theorem apply at all; that the possible values of a measurement are the points of \(\sigma(A)\) is the support statement; and that the probability of finding a value in \(\Omega\) is \(\norm{E(\Omega)\psi}^{2}\) is the positive measure of Definition 12.58, whose total mass is \(1\) for a normalised state. The expectation value is \(\braket{\psi}{A\psi}\), which is Equation (12.37) read as the mean of that probability measure. Functions of observables — \(\Ham^{2}\), \(\ee^{-\ii\Ham t/\hbar}\), the projectors of a filtering measurement — are Equation (12.39), and Proposition 12.61 is what allows the treatise to speak of the function of an operator. This is the content von Neumann extracted from the formalism [vonNeumann:1930] [vonNeumann:1932], and it is the reason the postulates can be stated without approximation or interpretation of the mathematics.
Stone's theorem
A family \(\set{U(t)}_{t\in\R}\) of unitary operators on \(\mathcal{H}\) is a strongly continuous one-parameter unitary group if
and for every \(x\in\mathcal{H}\) the map \(t\longmapsto U(t)x\) is continuous from \(\R\) to \(\mathcal{H}\). The parameter \(t\) carries the SI unit \(\mathrm{s}\) throughout this section. Rests on Definitions 6.6 and 12.41.
Continuity is demanded of \(t\longmapsto U(t)x\) for each fixed \(x\) (strong continuity), not of \(t\longmapsto U(t)\) in the operator norm, which would be a much stronger and, for the groups physics needs, a false requirement — norm continuity forces the generator to be bounded, and Theorem 12.75 will show that the generators of interest are not.
Let \(H\in\mathcal{B}(\mathcal{H})\) be self-adjoint, with the SI dimension of energy, \(\mathrm{J}\). Then the series
converges in \(\mathcal{B}(\mathcal{H})\) for every \(t\) and defines a strongly continuous one-parameter unitary group with
the derivative existing in the norm of \(\mathcal{H}\). Rests on Proposition 12.37, Definition 12.64 and Proposition 12.39.
Derives Proposition 12.65. Since \(\hbar\) has the SI unit \(\mathrm{J}\,\mathrm{s}\), the combination \(Ht/\hbar\) is dimensionless and Equation (12.42) is a sum of operators, as it must be. By Equation (12.22) the \(n\)th term has norm at most \((\abs{t}\norm{H}/\hbar)^{n}/n!\), so the series converges absolutely and hence in \(\mathcal{B}(\mathcal{H})\), by Propositions 12.8 and 12.37.
The group law Equation (12.41) follows from the Cauchy product of the two absolutely convergent series (Proposition 7.49) and the binomial theorem, which is legitimate because the two exponents commute, both being polynomials in the single operator \(H\). Taking adjoints term by term — legitimate by Equation (12.24) and continuity of the adjoint, which is isometric — and using \(H^{\dagger}=H\) gives \(U(t)^{\dagger}=\exp(+\ii Ht/\hbar)=U(-t)\), which by the group law is \(U(t)^{-1}\); so each \(U(t)\) is unitary.
For Equation (12.43), the difference quotient
has norm at most \(\norm{(U(s)-\identity)/s+\ii H/\hbar}\), and the bracket is the tail of Equation (12.42) divided by \(s\), of norm at most \(\abs{s}\norm{H}^{2}\hbar^{-2}\exp(\abs{s}\norm{H}/\hbar)\), which tends to \(0\). Multiplying by \(\ii\hbar\) gives Equation (12.43), in the operator norm and hence also strongly.
∎The map \(H\longmapsto U(t)=\exp(-\ii Ht/\hbar)\) is a bijection between the self-adjoint operators on \(\mathcal{H}\) — bounded or not, in the sense of Section 12.5.2 — and the strongly continuous one-parameter unitary groups on \(\mathcal{H}\). Given the group, the operator is recovered as
on the domain \(D(H)\) of vectors for which the limit exists, and \(D(H)\) is dense with \(H\) self-adjoint on it. Rests on Definition 12.64, Proposition 12.65 and Theorem 12.59.
Half of Theorem 12.66 was proved in Proposition 12.65 for bounded \(H\), and one further piece — everything that does not require the unbounded functional calculus — is proved next; the remainder is derived in Stone's Theorem on One-Parameter Unitary Groups.
Let \(U\) be a strongly continuous one-parameter unitary group and let \(D(H)\) and \(H\) be as in Equation (12.44). Then \(D(H)\) is a subspace, \(U(t)D(H)=D(H)\) for every \(t\), and
Moreover, for \(x\in D(H)\) the curve \(t\longmapsto U(t)x\) is differentiable and
Rests on Definition 12.64 and Theorem 12.66.
Derives Proposition 12.67. \(D(H)\) is a subspace because the limit in Equation (12.44) is linear in \(x\) where it exists. If \(x\in D(H)\) then, by Equation (12.41), \(\bigl(U(t)U(s)x-U(t)x\bigr)/s=U(t)\bigl(U(s)x-x\bigr)/s\), and \(U(t)\) is isometric, so the limit as \(s\to0\) exists and equals \(-\ii\hbar^{-1}U(t)Hx\); thus \(U(t)x\in D(H)\) with \(HU(t)x=U(t)Hx\), which is Equation (12.46), and applying the same with \(-t\) gives \(U(t)D(H)=D(H)\).
For Equation (12.45), let \(x,y\in D(H)\) and let \(t\neq0\). Since \(U(t)^{\dagger}=U(-t)\), and writing \(s=-t\),
Now \(s\to0\) exactly when \(t\to0\), and by Equation (12.44) both difference quotients converge: \(\bigl(U(t)y-y\bigr)/t\to-\ii Hy/\hbar\) and \(\bigl(U(s)x-x\bigr)/s\to-\ii Hx/\hbar\). Taking the limit in Equation (12.47), with the bracket carried through the limit by Proposition 12.4 and the antilinearity Equation (5.36) applied on the right,
which is Equation (12.45).
∎The two hard halves — that the domain defined by the limit in Equation (12.44) is dense, and that the operator on it is self-adjoint rather than merely symmetric; and the converse, that every self-adjoint operator, bounded or not, exponentiates to a group — are derived in Stone's Theorem on One-Parameter Unitary Groups. The first uses a vector-valued integral to produce enough vectors in the domain by smoothing, and then an elementary differential-equation argument to kill both deficiency subspaces; the second needs the spectral theorem for unbounded self-adjoint operators, which the appendix obtains from Theorem 12.59 through the Cayley transform rather than quoting. Stone's paper [Stone:1932] and [Reed:1972], chapter VIII, carry both.
Full derivation in Appendix A.
Derives Theorem 12.66.
Equation (12.46) is the Schrödinger equation before any Hamiltonian has been chosen: written for a state vector it is
and The Postulates of Quantum Mechanics adopts it as a postulate. Stone's theorem [Stone:1932] is what makes the two halves of Equation (12.48) equivalent, and it explains the otherwise unmotivated demand that the Hamiltonian be self-adjoint and not merely symmetric: symmetry alone gives neither a unitary group nor a conserved norm, so probability would not be conserved. That is not a technicality — Section 12.5.2 exhibits a symmetric momentum operator with no self-adjoint extension at all.
The factor \(\hbar\) is carried explicitly here and everywhere in this treatise (CLAUDE.md rule 3). Its role is dimensional and is worth stating once: \(\Ham\) has the unit \(\mathrm{J}\) and \(t\) the unit \(\mathrm{s}\), so \(\Ham t\) has the unit \(\mathrm{J}\,\mathrm{s}\), which is the unit of \(\hbar\); the argument of the exponential is therefore a pure number, as the series Equation (12.42) requires. The generator of the group is thus not the Hamiltonian but \(\Ham/\hbar\), an angular frequency in \(/\mathrm{s}\) — which is the mathematical form of the Planck–Einstein relation, and the reason a stationary state of energy \(E\) oscillates at angular frequency \(E/\hbar\). The convention \(\hbar=1\), common in the literature, sets a unit of action to \(1\) and is not used here.
Unbounded operators and domains
Closed operators and domains
An operator on \(\mathcal{H}\) is a pair \((A,D(A))\) with \(D(A)\subseteq\mathcal{H}\) a subspace, the domain, and \(A:D(A)\longrightarrow\mathcal{H}\) linear. It is densely defined if \(D(A)\) is dense. Two operators are equal only if their domains agree; \(A\subseteq B\) means \(D(A)\subseteq D(B)\) and \(Bx=Ax\) on \(D(A)\), and \(B\) is then an extension of \(A\). Rests on Definition 5.35 and Corollary 12.19.
The insistence that the domain be part of the operator is not bookkeeping. Section 12.5.2 exhibits one differential expression, on one interval, that is several different operators according to the domain chosen, with different spectra and different physics; and Theorem 12.75 shows that the operators of quantum mechanics have no choice in the matter.
The graph of \((A,D(A))\) is
where \(\mathcal{H}\oplus\mathcal{H}\) carries the inner product \(\braket{(x_{1},y_{1})}{(x_{2},y_{2})} =\braket{x_{1}}{x_{2}}+\braket{y_{1}}{y_{2}}\). The operator is closed if \(\Gamma(A)\) is a closed subspace, and closable if the closure of \(\Gamma(A)\) is itself the graph of an operator, called the closure \(\overline{A}\). Rests on Definitions 6.3 and 12.69.
Unwinding Equation (12.49): \(A\) is closed if and only if, whenever \(x_{n}\in D(A)\) with \(x_{n}\to x\) and \(Ax_{n}\to y\), it follows that \(x\in D(A)\) and \(Ax=y\). This is strictly weaker than continuity, which would require the second hypothesis to follow from the first. The closure of the graph fails to be a graph exactly when it contains some \((0,y)\) with \(y\neq0\), which is the obstruction to closability.
Let \((A,D(A))\) be densely defined. Put
and \(A^{\dagger}y=z\) on it. The vector \(z\) is unique because \(D(A)\) is dense: if \(\braket{z-z'}{x}=0\) for every \(x\) in a dense set then \(z=z'\), by Corollary 12.19. Rests on Definition 12.69, Theorem 12.38 and Corollary 12.19.
A densely defined \((A,D(A))\) is symmetric if \(\braket{x}{Ay}=\braket{Ax}{y}\) for all \(x,y\in D(A)\), equivalently \(A\subseteq A^{\dagger}\); it is self-adjoint if \(A=A^{\dagger}\), that is, if in addition \(D(A^{\dagger})=D(A)\). Rests on Definitions 12.69 and 12.71.
For densely defined \(A\), the operator \(A^{\dagger}\) is closed. A self-adjoint operator is therefore closed, and a symmetric operator is closable, with \(\overline{A}\subseteq A^{\dagger}\). Rests on Definitions 12.70 and 12.71.
Derives Proposition 12.73. Let \(y_{n}\in D(A^{\dagger})\) with \(y_{n}\to y\) and \(A^{\dagger}y_{n}\to z\). For every \(x\in D(A)\), \(\braket{A^{\dagger}y_{n}}{x}=\braket{y_{n}}{Ax}\); letting \(n\to\infty\) and using Proposition 12.4 on both sides gives \(\braket{z}{x}=\braket{y}{Ax}\). By Equation (12.50) this says \(y\in D(A^{\dagger})\) with \(A^{\dagger}y=z\), which is closedness. If \(A\) is symmetric then \(A\subseteq A^{\dagger}\) with \(A^{\dagger}\) closed, so the closure of \(\Gamma(A)\) lies inside \(\Gamma(A^{\dagger})\) and is therefore a graph.
∎Let \(A\) be defined on all of \(\mathcal{H}\) and symmetric, that is \(\braket{x}{Ay}=\braket{Ax}{y}\) for all \(x,y\in\mathcal{H}\). Then \(A\) is bounded. Rests on Definition 12.72, Proposition 12.36 and Remark 12.1.
Derives Theorem 12.74. First, \(A\) is closed. Let \(x_{n}\to x\) and \(Ax_{n}\to y\). For every \(z\in\mathcal{H}\),
using symmetry twice and Proposition 12.4 for the limits. Hence \(\braket{z}{y-Ax}=0\) for all \(z\), so \(y=Ax\) by Equation (5.31): the graph is closed. Since \(A\) is everywhere defined and closed between Banach spaces, the closed-graph theorem (Remark 12.1) makes it bounded.
∎Let \(\mathcal{H}\neq\set{0}\) and let \(Q,P\in\mathcal{B}(\mathcal{H})\). Then
is impossible. Rests on Proposition 12.37 and Definition 12.35.
Derives Theorem 12.75. Assume Equation (12.51). By induction,
the step being \(QP^{n}-P^{n}Q =\left(QP-PQ\right)P^{n-1}+P\left(QP^{n-1}-P^{n-1}Q\right)\), which by the inductive hypothesis is \(\ii\hbar P^{n-1}+\ii\hbar(n-1)P^{n-1}\).
No power of \(P\) vanishes. For suppose \(m\geq1\) is least with \(P^{m}=0\). Then Equation (12.52) at \(n=m\) reads \(0=\ii\hbar\,m\,P^{m-1}\), so \(P^{m-1}=0\); minimality forces \(m-1=0\), i.e. \(P^{0}=\identity=0\), which is false on \(\mathcal{H}\neq\set{0}\).
So \(\norm{P^{n-1}}>0\) for every \(n\), and taking norms in Equation (12.52) with Equation (12.22),
whence \(\hbar\,n\leq2\norm{Q}\norm{P}\) for every \(n\in\N\) — absurd.
∎Any pair of operators satisfying Equation (12.51) on a dense domain must be unbounded, and if they are symmetric they cannot be defined on all of \(\mathcal{H}\). Rests on Theorems 12.74 and 12.75.
Derives Corollary 12.76. Unboundedness is Theorem 12.75. If such an operator were symmetric and defined on all of \(\mathcal{H}\) it would be bounded by Theorem 12.74, a contradiction.
∎Corollary 12.76 is the reason this section exists. The canonical relation Equation (12.51) is not optional — it is the content of canonical quantisation (Canonical Quantization of Fields) and the source of the uncertainty relation — and it forbids position and momentum from being tame operators. They are unbounded, hence discontinuous, hence not defined on every state; every statement about them carries a domain, and The Postulates of Quantum Mechanics must say on which vectors an observable acts. The point was von Neumann's, and it is the technical core of the axiomatisation [vonNeumann:1930] [vonNeumann:1932].
A second consequence is that a domain is a piece of physics, not a technicality to be suppressed. For a differential operator the domain encodes the boundary conditions, so choosing it is choosing the physical situation — a particle in a box, on a ring, or on a half-line — and the eigenvalue problem that results is the Sturm–Liouville problem of Ordinary Differential Equations and Sturm–Liouville Theory, whose “self-adjoint boundary conditions” are exactly the choices that make the operator self-adjoint in the sense of Definition 12.72. The next subsection works one such family out completely.
Symmetric versus self-adjoint: deficiency indices
A symmetric operator \(A\) is essentially self-adjoint if its closure \(\overline{A}\) (Proposition 12.73) is self-adjoint; equivalently, if \(A\) has exactly one self-adjoint extension. A core for a self-adjoint \(B\) is a subspace \(D\subseteq D(B)\) on which \(B\) is essentially self-adjoint. Rests on Definition 12.72 and Proposition 12.73.
The three notions — symmetric, essentially self-adjoint, self-adjoint — are genuinely distinct, and the distinction is measured by two integers.
Let \(A\) be a closed symmetric operator and let \(\mu>0\) be a fixed number carrying the same SI dimension as \(A\), so that \(A\mp\ii\mu\) is dimensionally consistent. The deficiency subspaces are
and the deficiency indices are \(n_{\pm}=\dim K_{\pm}\). They do not depend on the choice of \(\mu\). Rests on Definition 12.71, Proposition 12.39 and Corollary 12.19.
The second equality in Equation (12.53) is Proposition 12.39 carried over to the unbounded case, where it reads \(\ker A^{\dagger}=(\im A)^{\perp}\) with the same one-line proof; the independence of \(\mu\) is part of von Neumann's analysis. Writing \(\mu\) rather than \(1\) is not pedantry: for a momentum operator \(\mu\) is a momentum in \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), and \(A^{\dagger}\mp\ii\) would be the difference of a momentum and a pure number.
Let \(A\) be a closed symmetric operator with deficiency indices \((n_{+},n_{-})\).
-
\(A\) is self-adjoint if and only if \(n_{+}=n_{-}=0\);
-
\(A\) has self-adjoint extensions if and only if \(n_{+}=n_{-}\), and they are then in bijective correspondence with the unitary maps \(K_{+}\longrightarrow K_{-}\) — a family parametrised by the group \(\U(n)\) when \(n_{+}=n_{-}=n<\infty\);
-
if \(n_{+}\neq n_{-}\), \(A\) has no self-adjoint extension at all.
The criterion is proved in Von Neumann's Criterion for Self-Adjoint Extensions. The proof passes to the Cayley transform, which converts a closed symmetric operator into an isometry between the orthogonal complements of the deficiency subspaces and a self-adjoint operator into a unitary; the self-adjoint extensions of the operator then correspond to the unitary extensions of that isometry, which exist exactly when the two deficiency subspaces have the same dimension. The appendix also derives the explicit description of each extension, and the independence of the indices from \(\mu\) asserted in Definition 12.79. Von Neumann's memoir of 1930 [vonNeumann:1930] is the original.
Full derivation in Appendix A.
Derives Theorem 12.80.
The deficiency-index theory and von Neumann's extension theorem are in the second Reed–Simon volume [Reed:1975], chapter X; the first volume [Reed:1972], which is what the rest of this chapter cites, does not carry them. The two examples that follow are the ones the criterion was made for, and they differ only in the interval.
Let \(L>0\) be a length and \(\mathcal{H}=L^{2}\bigl([0,L]\bigr)\), and let
with the initial domain \(D_{0}\) of continuously differentiable functions vanishing at both endpoints, which is dense. Then \(P\) is symmetric on \(D_{0}\) but not self-adjoint; its deficiency indices are \((n_{+},n_{-})=(1,1)\); and it has a one-parameter family of self-adjoint extensions \(P_{\theta}\), \(\theta\in[0,2\pi)\), with domains
whose spectra are the discrete sets
with normalised eigenfunctions \(\psi_{n}(x)=L^{-1/2}\ee^{\ii p_{n}x/\hbar}\). Rests on Definition 12.79, Theorem 12.80 and Example 12.11.
Derives Example 12.81. Symmetry. For \(\varphi,\psi\in D_{0}\), integration by parts (Theorem 7.43) gives
which vanishes because both functions vanish at both endpoints.
The adjoint. A computation with Equation (12.50), using the du Bois-Reymond lemma to identify the weak derivative, shows that \(P^{\dagger}\) is the same differential expression on the maximal domain — absolutely continuous \(\psi\) with \(\psi'\in\mathcal{H}\) and no boundary condition [Reed:1972]. That is already the failure of self-adjointness: \(D(P^{\dagger})\) is far larger than \(D_{0}\).
Deficiency indices. Solve \(P^{\dagger}\psi=\pm\ii\mu\psi\) with \(\mu>0\) a momentum. The equation \(-\ii\hbar\psi'=\ii\mu\psi\) gives \(\psi(x)=C\ee^{-\mu x/\hbar}\), and \(-\ii\hbar\psi'=-\ii\mu\psi\) gives \(\psi(x)=C\ee^{+\mu x/\hbar}\). Both are continuous on the compact interval \([0,L]\), hence bounded, hence in \(L^{2}([0,L])\); each solution space is one-dimensional, so \(n_{+}=n_{-}=1\) and Theorem 12.80 predicts a family of extensions parametrised by \(\U(1)\) — a circle.
The extensions. On \(D_{\theta}\) the boundary term Equation (12.57) vanishes, since \(\varphi(L)^{\ast}\psi(L) =\ee^{-\ii\theta}\ee^{\ii\theta}\varphi(0)^{\ast}\psi(0) =\varphi(0)^{\ast}\psi(0)\), so each \(P_{\theta}\) is symmetric. That each is exactly self-adjoint, rather than merely symmetric, and that the \(D_{\theta}\) exhaust the possibilities — a boundary condition making Equation (12.57) vanish for all \(\varphi,\psi\) in the domain must relate \(\psi(L)\) to \(\psi(0)\) by a phase — are proved in Corollary A.274, which also computes which unitary map of Theorem 12.80 carries which value of \(\theta\).
Spectrum. \(P_{\theta}\psi=p\psi\) has the solutions \(\psi(x)=C\ee^{\ii px/\hbar}\), which lie in \(D_{\theta}\) exactly when \(\ee^{\ii pL/\hbar}=\ee^{\ii\theta}\), that is, when \(pL/\hbar=\theta+2\pi n\) for an integer \(n\) — which is Equation (12.56). These eigenfunctions are orthonormal after the stated normalisation and, being the trigonometric system of Fourier Analysis and Integral Transforms up to the phase \(\theta\), form an orthonormal basis; so the spectrum is pure point.
∎Let \(\mathcal{H}=L^{2}\bigl([0,\infty)\bigr)\) and let \(P\) be Equation (12.54) on the dense domain of continuously differentiable functions of compact support in \((0,\infty)\). Then \(P\) is symmetric, its deficiency indices are \((n_{+},n_{-})=(1,0)\), and by Theorem 12.80 it has no self-adjoint extension. There is no momentum observable for a particle confined to a half-line. Rests on Example 12.81, Theorem 12.80 and Definition 12.79.
Derives Example 12.82. Symmetry is Equation (12.57) again, with both boundary terms zero because the functions have compact support in the open half-line. The adjoint is once more the maximal operator, with no boundary condition [Reed:1972]. Solving \(P^{\dagger}\psi=\ii\mu\psi\) gives \(\psi(x)=C\ee^{-\mu x/\hbar}\), which is square integrable on \([0,\infty)\) with \(\norm{\psi}^{2}=\abs{C}^{2}\hbar/(2\mu)\), so \(n_{+}=1\). Solving \(P^{\dagger}\psi=-\ii\mu\psi\) gives \(\psi(x)=C\ee^{+\mu x/\hbar}\), whose square integral diverges unless \(C=0\), so \(n_{-}=0\). The indices differ, and case (3) of Theorem 12.80 applies.
∎The same expression Equation (12.54) behaves in three different ways according to the interval, and nothing but the domain distinguishes the cases.
-
On the whole line, \(P\) on the smooth rapidly decreasing functions is essentially self-adjoint, with purely continuous spectrum \(\R\) and no eigenvectors — the Fourier transform of Fourier Analysis and Integral Transforms turns it into multiplication by \(\hbar k\), which is Example 12.56 on an unbounded interval.
-
On a finite interval it is not self-adjoint but has a circle of self-adjoint extensions, each with a discrete spectrum Equation (12.56). The parameter \(\theta\) is physical: \(\theta=0\) is the particle on a ring, and a non-zero \(\theta\) is the ring threaded by a magnetic flux, the phase being the Aharonov–Bohm phase. The mathematics does not choose; the physical situation does.
-
On the half-line it has no self-adjoint extension. The reason is visible in the proof: the translations generated by \(P\) push probability off the end of the half-line and back on again, which no unitary group can do. By Theorem 12.66, no self-adjoint generator, no unitary group; and no unitary group, no conservation of probability.
The corresponding question for the Hamiltonian — whether \(-\hbar^{2}\nabla^{2}/(2m)+V\) is essentially self-adjoint on the smooth compactly supported functions — is the question of whether the Schrödinger equation has a unique unitary dynamics, and it is answered, potential by potential, in The Hydrogen Atom and Approximation Methods [Reed:1975] [Kato:1966].
Direct sums, tensor products and rigged spaces
Direct sums and orthogonal decompositions
Let \(\mathcal{H}_{1},\mathcal{H}_{2},\dots\) be Hilbert spaces. Their direct sum is
with componentwise operations and \(\braket{x}{y}=\sum_{n}\braket{x_{n}}{y_{n}}\). Rests on Definition 5.97 and Example 12.9.
Equation (12.58) is a Hilbert space, and the maps \(J_{n}:\mathcal{H}_{n}\longrightarrow\bigoplus_{m}\mathcal{H}_{m}\) placing a vector in the \(n\)th slot are isometric with mutually orthogonal images. Rests on Definition 12.84 and Proposition 12.10.
Derives Proposition 12.85. The inner product is well defined and the space is a vector space by the same two estimates as in Example 12.9, applied to the numbers \(\norm{x_{n}}\); the axioms of Definition 5.17 hold slotwise. Completeness is the argument of Proposition 12.10 verbatim, with the modulus \(\abs{x_{n}}\) replaced by the norm \(\norm{x_{n}}\) and the completeness of \(\C\) replaced by that of \(\mathcal{H}_{n}\): each slot converges, the finite truncations give a uniform bound, and the limit lies in the sum. The last claim is immediate, distinct slots pairing to zero. Note that \(\ell^{2}\) is the special case \(\mathcal{H}_{n}=\C\) for every \(n\).
∎Closed subspaces \(M_{1},M_{2},\dots\) of \(\mathcal{H}\) are said to decompose \(\mathcal{H}\), written \(\mathcal{H}=\bigoplus_{n}M_{n}\), if they are pairwise orthogonal and the only vector orthogonal to all of them is \(0\). Rests on Definitions 12.16 and 12.29.
If \(\mathcal{H}=\bigoplus_{n}M_{n}\) with orthogonal projections \(P_{n}=P_{M_{n}}\), then \(P_{n}P_{m}=\delta_{nm}P_{n}\) and, for every \(x\in\mathcal{H}\),
the series converging in norm. The map \(x\longmapsto(P_{1}x,P_{2}x,\dots)\) is an inner-product-preserving bijection onto \(\bigoplus_{n}M_{n}\) in the sense of Definition 12.84. Rests on Definition 12.86, Theorem 12.30 and Proposition 12.85.
Derives Proposition 12.87. Orthogonality of the subspaces gives \(\im P_{m}\subseteq M_{m} \subseteq M_{n}^{\perp}=\ker P_{n}\) for \(m\neq n\), which is \(P_{n}P_{m}=0\); and \(P_{n}^{2}=P_{n}\) is Proposition 12.21.
Choose an orthonormal basis of each \(M_{n}\) — possible in the separable case by Corollary 12.24, and in general by the same argument applied to a maximal orthonormal family — and let \(\set{e_{k}}\) be the union of all of them, which is orthonormal because the \(M_{n}\) are pairwise orthogonal. This family is maximal: a vector orthogonal to every \(e_{k}\) is orthogonal to every \(M_{n}\), hence \(0\) by Definition 12.86. So Theorem 12.30 applies, and grouping the terms of Equation (12.18) by the subspace they belong to — legitimate because the partial sums converge absolutely in the sense of Proposition 12.28, whose criterion is insensitive to order — gives Equation (12.59), the group of index \(n\) summing to \(P_{n}x\) by Proposition 12.27. The norm identity is Equation (12.19) grouped the same way, and it says precisely that the displayed map lands in Equation (12.58) and is isometric; surjectivity follows from Proposition 12.28 applied slotwise.
∎A closed subspace \(M\) reduces \(A\in\mathcal{B}(\mathcal{H})\) if \(A(M)\subseteq M\) and \(A(M^{\perp})\subseteq M^{\perp}\). Rests on Definitions 12.20 and 12.35.
A closed subspace \(M\) reduces \(A\in\mathcal{B}(\mathcal{H})\) if and only if \(P_{M}A=AP_{M}\). Rests on Definition 12.88 and Proposition 12.21.
Derives Proposition 12.89. Write \(P=P_{M}\). (\(\Leftarrow\)) If \(x\in M\) then \(Ax=APx=PAx\in M\); if \(x\in M^{\perp}\) then \(Px=0\), so \(PAx=APx=0\) and \(Ax\in M^{\perp}\) by Proposition 12.21.
(\(\Rightarrow\)) For arbitrary \(x\) split \(x=Px+(\identity-P)x\). The first term lies in \(M\), so \(APx\in M\) and \(PAPx=APx\); the second lies in \(M^{\perp}\), so \(A(\identity-P)x\in M^{\perp}\) and \(PA(\identity-P)x=0\). Adding, \(PAx=PAPx+PA(\identity-P)x=APx\).
∎The extreme case — no proper reducing subspace at all — is what representation theory calls irreducibility, and Proposition 12.89 converts it into a statement about commuting operators that can actually be used.
A set \(\mathcal{S}\subseteq\mathcal{B}(\mathcal{H})\) is self-adjoint if \(A\in\mathcal{S}\) implies \(A^{\dagger}\in\mathcal{S}\). Its commutant is
and \(\mathcal{S}\) acts irreducibly on \(\mathcal{H}\) if the only closed subspaces reducing every member of \(\mathcal{S}\) are \(\set{0}\) and \(\mathcal{H}\). Rests on Definitions 12.41 and 12.88.
Let \(\mathcal{S}\subseteq\mathcal{B}(\mathcal{H})\) be self-adjoint. Then \(\mathcal{S}\) acts irreducibly if and only if \(\mathcal{S}'=\C\,\identity\): every bounded operator commuting with every member of an irreducible self-adjoint family is a multiple of the identity, and conversely. Rests on Definition 12.90, Proposition 12.89 and Theorem 12.59.
Derives Theorem 12.91. (\(\Rightarrow\)) Let \(B\in\mathcal{S}'\). Taking adjoints in \(BA=AB\) and using Equation (12.24) gives \(A^{\dagger}B^{\dagger}=B^{\dagger}A^{\dagger}\) for every \(A\in\mathcal{S}\); since \(\mathcal{S}\) is self-adjoint this says \(B^{\dagger}\in\mathcal{S}'\), and hence the self-adjoint operators \(\tfrac{1}{2}(B+B^{\dagger})\) and \(\tfrac{1}{2\ii}(B-B^{\dagger})\) lie in \(\mathcal{S}'\) too. As \(B\) is their combination \(B_{1}+\ii B_{2}\), it is enough to treat a self-adjoint \(B\in\mathcal{S}'\).
Let \(E\) be the projection-valued measure of \(B\) (Theorem 12.59). Every \(A\in\mathcal{S}\) is a bounded operator commuting with \(B\), so by the last clause of that theorem it commutes with every \(E(\Omega)\); by Proposition 12.89 the range of \(E(\Omega)\) therefore reduces every member of \(\mathcal{S}\), and irreducibility forces \(E(\Omega)\in\set{0,\identity}\) for every Borel \(\Omega\).
The spectrum \(\sigma(B)\) is compact and non-empty (Theorem 12.54) and \(E\) is supported on it, so \(\lambda\longmapsto E((-\infty,\lambda])\) passes from \(0\) to \(\identity\) within a bounded range and \(c=\inf\set{\lambda\in\R\mid E((-\infty,\lambda])=\identity}\) is finite. The sets \((-\infty,c-1/n]\) increase to \((-\infty,c)\) and each carries \(E=0\); the sets \((c+1/n,\infty)\) increase to \((c,\infty)\) and each carries \(E=\identity-E((-\infty,c+1/n])=0\). Writing an increasing union as the disjoint union of its successive differences and applying Equation (12.36) gives \(E((-\infty,c))=0\) and \(E((c,\infty))=0\), whence \(E(\set{c})=\identity\). Then Equation (12.37) reads \(\braket{x}{By}=c\braket{x}{y}\) for all \(x,y\), that is \(B=c\,\identity\), with \(c\) real because \(B\) is self-adjoint.
(\(\Leftarrow\)) If a closed subspace \(M\) reduces every member of \(\mathcal{S}\) then \(P_{M}\in\mathcal{S}'\) by Proposition 12.89, so \(P_{M}=c\,\identity\); squaring, \(c^{2}=c\), so \(c\in\set{0,1}\) and \(M\) is \(\set{0}\) or \(\mathcal{H}\).
∎Theorem 12.91 is the Hilbert-space form of Theorem 5.155, the second lemma of Schur for finite-dimensional group representations; the hypothesis that has to be added in infinite dimension is that \(\mathcal{S}\) be closed under adjoints, without which the argument above cannot pass to the self-adjoint parts of \(B\) and the conclusion is false. The self-adjointness costs nothing in practice, because the families that arise are families of unitaries or of self-adjoint operators. That matters for the way the lemma is used on canonical pairs: position and momentum are unbounded (Corollary 12.76) and so are not members of \(\mathcal{B}(\mathcal{H})\) at all, but the unitary groups they generate are, and it is to those that Theorem 12.91 is applied in Section 12.7.
Three sources of orthogonal decompositions matter later, and all three are Proposition 12.87 with a different supply of subspaces. The spectral projections \(E(\Omega)\) of Theorem 12.59 decompose \(\mathcal{H}\) over any Borel partition of the spectrum, and by Proposition 12.89 every such piece reduces the operator — the abstract form of “diagonalise, then work block by block”. The eigenspaces of the total angular momentum decompose the state space of a rotationally invariant system, which is how Angular Momentum and Spin reduces a problem to one value of \(\ell\) at a time, and the same decomposition organises the hydrogen spectrum of The Hydrogen Atom. And a superselection rule is the statement that a decomposition is respected by every observable, so that no superposition across the summands is ever prepared or detected; Interpretations (Evidence-Anchored) takes that up.
Tensor products
The direct sum adds dimensions; the tensor product multiplies them (Remark 5.102), and it is the tensor product that describes a composite system. The algebraic construction and its universal property are those of Linear Algebra and Representation Theory; what is new here is the inner product and the completion.
Let \(\mathcal{H}_{1},\mathcal{H}_{2}\) be Hilbert spaces and let \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\) be their algebraic tensor product, the space of finite sums \(\sum_{i}u_{i}\otimes v_{i}\) characterised by the universal property that every bilinear map out of \(\mathcal{H}_{1}\times\mathcal{H}_{2}\) factors uniquely through it. On it define
extended to finite sums by sesquilinearity. The tensor product \(\mathcal{H}_{1}\otimes\mathcal{H}_{2}\) is the completion of \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\) in the associated norm. Rests on Definitions 5.17, 5.106 and 12.2.
Equation (12.61) extends to a genuine inner product on \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\), and if \(\set{e_{i}}\) and \(\set{f_{j}}\) are orthonormal bases of \(\mathcal{H}_{1}\) and \(\mathcal{H}_{2}\) then \(\set{e_{i}\otimes f_{j}}\) is an orthonormal basis of \(\mathcal{H}_{1}\otimes\mathcal{H}_{2}\). Rests on Definition 12.94, Theorem 12.30 and Proposition 12.23.
Derives Proposition 12.95. Well defined. The right-hand side of Equation (12.61) is bilinear in \((u_{2},v_{2})\) and antibilinear in \((u_{1},v_{1})\), so by the universal property it factors through the algebraic tensor product in each argument separately; the resulting form on \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\) does not depend on the representation of an element as a finite sum.
Positive definite. Let \(w=\sum_{i=1}^{r}u_{i}\otimes v_{i}\) be non-zero. Applying Proposition 12.23 to the \(v_{i}\) and re-expanding, we may assume the \(v_{i}\) orthonormal, in which case Equation (12.61) gives \(\braket{w}{w}=\sum_{i}\norm{u_{i}}^{2}\geq0\), with equality only if every \(u_{i}=0\), i.e. only if \(w=0\). The remaining axioms of Definition 5.17 are inherited from the two factors.
Basis. The family \(\set{e_{i}\otimes f_{j}}\) is orthonormal by Equation (12.61). It is maximal: its closed span contains every \(u\otimes v\), because expanding \(u\) and \(v\) by Equation (12.18) and using continuity of Equation (12.61) in each slot gives \(u\otimes v=\sum_{i,j}\braket{e_{i}}{u}\braket{f_{j}}{v}\, e_{i}\otimes f_{j}\); hence it contains \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\), which is dense in the completion. Now apply criterion (2) of Theorem 12.30.
∎For \(A_{1}\in\mathcal{B}(\mathcal{H}_{1})\) and \(A_{2}\in\mathcal{B}(\mathcal{H}_{2})\) there is a unique \(A_{1}\otimes A_{2}\in \mathcal{B}(\mathcal{H}_{1}\otimes\mathcal{H}_{2})\) with \((A_{1}\otimes A_{2})(u\otimes v)=A_{1}u\otimes A_{2}v\), and
Derives Proposition 12.96. Uniqueness is density of \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\) and continuity. For existence and the norm, factor \(A_{1}\otimes A_{2}=(A_{1}\otimes\identity)(\identity\otimes A_{2})\) and bound each factor separately. Expanding \(w=\sum_{j}u_{j}\otimes f_{j}\) in the orthonormal basis \(\set{f_{j}}\) of \(\mathcal{H}_{2}\) — possible by Proposition 12.95 — gives \(\norm{w}^{2}=\sum_{j}\norm{u_{j}}^{2}\) and
so \(\norm{A_{1}\otimes\identity}\leq\norm{A_{1}}\), and symmetrically for the other factor; hence \(\norm{A_{1}\otimes A_{2}}\leq\norm{A_{1}}\norm{A_{2}}\), which also shows the map is bounded on the dense subspace and so extends. The reverse inequality follows by testing on product vectors, \(\norm{A_{1}u\otimes A_{2}v}=\norm{A_{1}u}\norm{A_{2}v}\), and taking suprema. The adjoint identity is Equation (12.61) checked on product vectors, together with uniqueness in Theorem 12.38.
∎It is worth stating what Definition 12.94 does and does not claim. The universal property — every bilinear map factors uniquely — belongs to the algebraic tensor product \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\), and it is what makes Equation (12.61) well defined. The completion \(\mathcal{H}_{1}\otimes\mathcal{H}_{2}\) does not inherit it for arbitrary bounded bilinear maps: a bilinear \(B\) with \(\norm{B(u,v)}\leq C\norm{u}\norm{v}\) need not extend continuously to the completion, because the Hilbert norm on \(\mathcal{H}_{1}\odot\mathcal{H}_{2}\) is not the largest such norm. The Hilbert tensor product is one of several completions of the same algebraic object, singled out by Equation (12.61) and by nothing else, and it is singled out for a physical reason: it is the one whose orthonormal basis is the set of products of basis vectors (Proposition 12.95), which is the statement that a joint state is specified by an amplitude for each pair of outcomes.
For a single particle in ordinary space the state space is \(L^{2}(\R^{3})\), and for two distinguishable particles it is
the isomorphism carrying \(f\otimes g\) to the function \((x_{1},x_{2})\longmapsto f(x_{1})\,g(x_{2})\). It is an isomorphism because the products \(f(x_{1})g(x_{2})\) span a dense subspace of \(L^{2}(\R^{6})\) and because Equation (12.61) matches the Fubini factorisation of the integral over \(\R^{6}\) [Reed:1972]. Note that the configuration space of the pair is \(\R^{6}\), not \(\R^{3}\): a two-particle wavefunction is a function of six coordinates and is not a field on space. That is a statement about the formalism with observable consequences, and it is where the non-classical correlations begin. Rests on Definition 12.94 and Example 12.11.
A vector \(w\in\mathcal{H}_{1}\otimes\mathcal{H}_{2}\) is a product vector if \(w=u\otimes v\) for some \(u,v\), and is entangled otherwise. Rests on Definition 12.94.
Let \(\mathcal{H}_{1}=\mathcal{H}_{2}=\C^{2}\) with orthonormal bases \(\set{e_{0},e_{1}}\) and \(\set{f_{0},f_{1}}\). The unit vector
is entangled. Rests on Definition 12.99 and Proposition 12.95.
Derives Example 12.100. Suppose \(w=(a e_{0}+b e_{1})\otimes(c f_{0}+d f_{1})\). Expanding in the basis \(\set{e_{i}\otimes f_{j}}\), which is a basis by Proposition 12.95, and comparing coefficients with Equation (12.64) gives
From \(ad=0\) either \(a=0\), contradicting \(ac=1/\sqrt{2}\), or \(d=0\), contradicting \(bd=1/\sqrt{2}\).
∎Definition 12.94 is the mathematics behind the composition postulate: the state space of a system assembled from two others is the tensor product of their state spaces, an assumption already explicit in von Neumann [vonNeumann:1932] and stated as a postulate in The Postulates of Quantum Mechanics. Example 12.100 is then not a curiosity but a forced consequence: the tensor product contains vectors that are not products, so a composite system has states in which neither part has a state of its own. Whether nature actually realises them is a question for experiment, and it has been settled affirmatively; the argument that turns Equation (12.64) into a testable inequality is Entanglement and Bell Tests, and the measurements are in Experiment: Bell Tests. For identical rather than merely distinguishable particles the state space is not the whole tensor product but a symmetric or antisymmetric subspace of it, which is a reducing subspace in the sense of Definition 12.88; Identical Particles takes that up.
Rigged Hilbert spaces
Two objects used constantly in quantum mechanics are not elements of \(\mathcal{H}\): the plane wave, which is not square integrable, and the delta “function”, which is not a function. Neither is an approximation to be apologised for; both are legitimate objects, and the structure they live in is a triple of spaces rather than a single one.
(i) The function \(x\longmapsto\ee^{\ii px/\hbar}\) is not in \(L^{2}(\R)\) for any real \(p\). (ii) There is no \(g\in L^{2}(\R)\) with \(\braket{g}{\varphi}=\varphi(0)\) for every continuous, compactly supported \(\varphi\). Rests on Example 12.11 and Proposition 12.4.
Derives Proposition 12.102. (i) The integrand \(\abs{\ee^{\ii px/\hbar}}^{2}=1\), so the integral over \(\R\) diverges.
(ii) Suppose such a \(g\) existed. Let \(\varphi_{n}\) be the continuous function equal to \(1\) at \(0\), vanishing outside \([-1/n,1/n]\), and linear in between, so \(\varphi_{n}(0)=1\) while \(\norm{\varphi_{n}}^{2}=\int\abs{\varphi_{n}}^{2}\leq2/n\). Then Equation (12.1) gives \(1=\abs{\braket{g}{\varphi_{n}}}\leq\norm{g}\sqrt{2/n}\) for every \(n\), which is false for large \(n\).
∎A Gelfand triple, or rigged Hilbert space, is a chain
in which \(\Phi\) is a dense subspace of \(\mathcal{H}\) carrying its own, strictly finer, topology under which it is a nuclear space, and \(\Phi'\) is the space of continuous linear functionals on \(\Phi\) for that topology. The middle inclusion is the Riesz identification of Corollary 12.47: each \(x\in\mathcal{H}\) acts on \(\Phi\) by \(\varphi\longmapsto\braket{x}{\varphi}\), which is linear in \(\varphi\) by Equation (5.34) and therefore is an element of \(\Phi'\), and that action determines \(x\) because \(\Phi\) is dense. The identification \(x\longmapsto\braket{x}{\cdot}\) is itself antilinear in \(x\), exactly as Corollary 12.47 says; it is the functionals, not the embedding, that are linear. Rests on Corollary 12.47, Definition 12.45 and Corollary 12.19.
The standard triple of quantum mechanics on the line is
where \(\mathcal{S}(\R)\) is the Schwartz space of smooth functions all of whose derivatives decay faster than any power, and \(\mathcal{S}'(\R)\) is the space of tempered distributions [Schwartz:1950]. Both objects of Proposition 12.102 live in the third space: the delta is the functional \(\varphi\longmapsto\varphi(0)\), which is continuous for the topology of \(\mathcal{S}\) though not for the \(L^{2}\) norm, and the plane wave is the functional
which is the Fourier transform of Fourier Analysis and Integral Transforms evaluated at the wavenumber \(p/\hbar\). Rests on Definition 12.103, Proposition 12.102 and Example 12.11.
Let \(A\) map \(\Phi\) into \(\Phi\) continuously. A functional \(F\in\Phi'\) is a generalized eigenvector of \(A\) with eigenvalue \(\lambda\) if
On the triple Equation (12.66), with \(P=-\ii\hbar\,\dd/\dd x\) mapping \(\mathcal{S}(\R)\) into itself, the functional Equation (12.67) satisfies \(F_{p}(P\varphi)=p\,F_{p}(\varphi)\) for every \(\varphi\in\mathcal{S}(\R)\), and \(p\) ranges over all of \(\R\). Rests on Definition 12.105, Example 12.104 and Equation (12.54).
Derives Proposition 12.106. Integrate by parts in Equation (12.67); the boundary terms vanish because \(\varphi\) and all its derivatives decay faster than any power. Writing \(C=(2\pi\hbar)^{-1/2}\),
and \(\ii\hbar\,\dd(\ee^{-\ii px/\hbar})/\dd x =\ii\hbar\left(-\ii p/\hbar\right)\ee^{-\ii px/\hbar} =p\,\ee^{-\ii px/\hbar}\), which gives \(p\,F_{p}(\varphi)\).
∎Let \(\Phi\subseteq\mathcal{H}\subseteq\Phi'\) be a Gelfand triple with \(\Phi\) nuclear, and let \(A\) be a self-adjoint operator on \(\mathcal{H}\) mapping \(\Phi\) continuously into itself. Then \(A\) possesses a complete set of generalized eigenvectors: there is a measure \(\mu\) on \(\sigma(A)\) and a family \(\set{F_{\lambda}}\) in \(\Phi'\), with \(F_{\lambda}\) a generalized eigenvector of eigenvalue \(\lambda\) for \(\mu\)-almost every \(\lambda\), such that
Rests on Definition 12.105, Theorem 12.59 and Definition 12.103.
The proof is in The Nuclear Spectral Theorem of Gelfand and Maurin: it builds the direct-integral decomposition supplied by the spectral theorem Theorem 12.59, and then uses nuclearity of the test space to realise the fibre maps as continuous functionals, so that almost every fibre furnishes a generalized eigenvector. This is the one long derivation of the chapter that rests on a theory this treatise does not develop — the theory of nuclear spaces, for which Gelfand and Vilenkin, volume 4, is the reference of record [Gelfand:1964]. The appendix isolates exactly what is assumed, in one displayed statement: that a nuclear test space embeds in \(\mathcal{H}\) through some Hilbert–Schmidt map. Everything else is derived, and two things the statement above suppresses are made explicit there: \(\mathcal{H}\) and \(\Phi\) are assumed separable, and when the spectrum of \(A\) has multiplicity greater than one the integrand of Equation (12.69) carries a sum over a multiplicity index — the display as written is the simple-spectrum case, which is the case of Proposition 12.106 and of every use made of the theorem in this treatise.
Full derivation in Appendix A.
Derives Theorem 12.107.
Equation (12.69) is the completeness relation that physics writes as \(\int\ketbra{\lambda}{\lambda}\,\dd\mu(\lambda)=\identity\), and Theorem 12.107 is what licenses it. The rule this treatise follows is therefore precise, and it is worth stating once for all the later chapters that use it.
-
A ket \(\ket{\lambda}\) belonging to a point of continuous spectrum denotes an element of \(\Phi'\), not of \(\mathcal{H}\). It is not a state, it cannot be normalised, and \(\braket{\lambda}{\lambda}\) is meaningless. Proposition 12.102 is the proof that no other reading is available.
-
An expression \(\braket{\lambda}{\varphi}\) is legitimate whenever \(\varphi\in\Phi\) — it is the value \(F_{\lambda}(\varphi)\) — and the “normalisation” \(\braket{\lambda}{\lambda'} =\delta(\lambda-\lambda')\) is shorthand for Equation (12.69), an identity between functionals and never between numbers.
-
A physical state is always a vector of \(\mathcal{H}\). A wave packet built by superposing generalized eigenvectors against a square-integrable amplitude is such a vector; a single plane wave is not, which is why every scattering calculation in Scattering Theory that is done with plane waves must be read as a statement about packets.
The triple also explains, retrospectively, why Dirac's formalism worked before it had a foundation [Dirac:1930b]: he was computing in \(\Phi'\) with functionals that the nuclear spectral theorem later showed to exist. The rigorous setting is due to Gelfand and his collaborators [Gelfand:1964], built on Schwartz's theory of distributions [Schwartz:1950]. It is a completion of the formalism, not a correction to it, and no result of this chapter is disturbed by it: the spectral theorem Theorem 12.59 remains the statement that carries the physics, and Equation (12.69) is a convenient way of writing its conclusion in the notation physicists actually use.
The canonical commutation relations in Weyl form
Theorem 12.75 shows that no pair of bounded operators satisfies \(\comm{\hat{q}}{\hat{p}}=\ii\hbar\identity\), so any realisation of that relation involves unbounded operators, each with a domain that is not the whole space (Corollary 12.76). This is not a technicality to be waved through. Two pairs can satisfy the relation on every vector where both sides are defined and still be inequivalent, because the relation says nothing about what happens off the common domain: momentum on a half-line has no self-adjoint realisation at all (Example 12.82), and momentum on an interval has a whole circle of them (Example 12.81), and in neither case does the formal commutator notice. The repair is to exponentiate, replacing the unbounded generators by the unitary groups they generate — which by Theorem 12.66 carry exactly the same information — and to impose the commutation relation on those.
A Weyl system of one degree of freedom on \(\mathcal{H}\) is a pair of strongly continuous one-parameter unitary groups (Definition 12.64)
with \(Q\) and \(P\) the self-adjoint generators supplied by Theorem 12.66, satisfying the Weyl relation
Here \(Q\) carries the SI unit \(\mathrm{m}\) and \(P\) the unit \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), so that the parameters carry \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\) and \(\mathrm{m}\) respectively and both \(\alpha Q/\hbar\) and \(\alpha\beta/\hbar\) are pure numbers. For \(f\) degrees of freedom the parameters become vectors \(\vect{\alpha},\vect{\beta}\in\R^{f}\), the exponents become \(\vect{\alpha}\cdot\vect{Q}\) and \(\vect{\beta}\cdot\vect{P}\), the phase in Equation (12.71) becomes \(\ee^{\ii\vect{\alpha}\cdot\vect{\beta}/\hbar}\), and the groups belonging to different degrees of freedom are required to commute. Rests on Definition 12.64, Theorem 12.66 and Theorem 12.75.
On \(\mathcal{H}=L^{2}(\R)\) put
Both are unitary, both are groups, and both are strongly continuous — the first by dominated convergence, the second because translation is continuous in the mean. Their generators are multiplication by \(x\) and \(-\ii\hbar\,\dd/\dd x\), on the domains where the corresponding limits Equation (12.44) exist. Applying the two definitions in the two orders, \((U(\alpha)V(\beta)\psi)(x)=\ee^{\ii\alpha x/\hbar}\psi(x-\beta)\) while \((V(\beta)U(\alpha)\psi)(x) =\ee^{\ii\alpha(x-\beta)/\hbar}\psi(x-\beta)\), so Equation (12.71) holds. This system is also irreducible in the sense of Definition 12.90 — a fact taken here from [Reed:1972], theorem VIII.14 and its preamble, not proved: the step that does the work is that a bounded operator commuting with multiplication by every bounded measurable function is itself such a multiplication, which is the maximal abelian property of \(L^{\infty}(\R)\) acting on \(L^{2}(\R)\), and only then does commuting with every translation force the multiplier to be constant. Rests on Definition 12.109 and Example 12.11.
In any Weyl system,
the domains being \(V(-\beta)D(Q)\) and \(U(-\alpha)D(P)\) respectively. Conversely either identity in Equation (12.73) implies Equation (12.71). Rests on Definition 12.109, Theorem 12.66 and Proposition 12.67.
Derives Proposition 12.111. Multiply Equation (12.71) on the left by \(V(\beta)^{\dagger}=V(-\beta)\):
For fixed \(\beta\) the left-hand side is, as a function of \(\alpha\), a strongly continuous one-parameter unitary group: the group law and the continuity are inherited from \(U\) because conjugation by a fixed unitary is multiplicative and isometric. Its generator is \(V(\beta)^{\dagger}QV(\beta)\) on the domain \(V(-\beta)D(Q)\), because
and \(V(\beta)^{\dagger}\) is isometric, so the limit on the left exists exactly when the limit Equation (12.44) defining \(Q\) exists at the vector \(V(\beta)x\), and is then its image under \(V(\beta)^{\dagger}\). The right-hand side of Equation (12.74) is likewise such a group, and its generator is \(Q+\beta\identity\) on \(D(Q)\): the summand \(\beta\identity\) is bounded and self-adjoint, and by Proposition 12.65 it contributes exactly the constant phase \(\ee^{\ii\alpha\beta/\hbar}\). Two groups that are equal have equal generators, by the bijection of Theorem 12.66; that is the first identity of Equation (12.73), and the second follows in the same way from Equation (12.71) rewritten as \(U(\alpha)^{\dagger}V(\beta)U(\alpha) =\ee^{-\ii\alpha\beta/\hbar}V(\beta)\). Conversely, exponentiating either identity of Equation (12.73) through the bijection of Theorem 12.66 returns Equation (12.74), hence Equation (12.71).
∎Let \((U,V)\) be a Weyl system and let
and set \(g(Q)=\int_{\R}m(\alpha)\,U(\alpha)\,\dd\alpha\), a bounded operator defined by the convergent vector-valued integral. Then for every \(x\in D(P)\) one has \(g(Q)x\in D(P)\) and
where \(g'(Q)=\int_{\R}(\ii\alpha/\hbar)\,m(\alpha)\,U(\alpha)\, \dd\alpha\) is the operator built the same way from \(g'\). Rests on Proposition 12.111, Theorem 12.66 and Definition 12.109.
Derives Corollary 12.112. By Equation (12.74) and linearity of the integral,
which is the operator built the same way from the translate \(\lambda\longmapsto g(\lambda+\beta)\). Subtracting \(g(Q)\), dividing by \(\beta\) and comparing with \(g'(Q)\), the operator norm of the difference is at most
because every \(U(\alpha)\) has norm \(1\). The integrand tends to zero pointwise as \(\beta\to0\) and is bounded by \(2\abs{\alpha}\abs{m(\alpha)}/\hbar\), using \(\abs{\ee^{\ii\theta}-1}\le\abs{\theta}\); that bound is integrable by the hypothesis in Equation (12.76), so dominated convergence (Remark 12.1) sends Equation (12.79) to zero. The right-hand side of Equation (12.78) is therefore differentiable in \(\beta\) at \(\beta=0\) in the operator norm, with derivative \(g'(Q)\).
Now fix \(x\in D(P)\) and write, for \(\beta\neq0\),
The left-hand side converges to \(g'(Q)x\) by the previous paragraph. In the first term on the right, \((V(\beta)x-x)/\beta\to-\ii Px/\hbar\) because \(x\in D(P)\) and \(V\) has generator \(P\) (Equation (12.44)), \(g(Q)\) is bounded, and \(V(-\beta)\to\identity\) strongly, so that term converges to \(-\ii g(Q)Px/\hbar\). The second term therefore converges as well, and by Equation (12.44) applied to the group \(V(-\cdot)\) its convergence is the statement that \(g(Q)x\in D(P)\), the limit being \(+\ii P\,g(Q)x/\hbar\). Collecting,
which is Equation (12.77).
∎Equation (12.77) is the identity every derivation of Ehrenfest's theorem uses, and the reason for proving it in the class Equation (12.76) rather than for a general function is that this class needs nothing but the Weyl relation: the operator \(g(Q)\) is built from the group itself, so no functional calculus for the unbounded \(Q\) has to be invoked and no question of domains arises beyond the one settled in the proof. The class contains every bounded potential whose Fourier transform decays fast enough — a Gaussian well, a smoothly screened Coulomb potential, any \(C^{2}\) function of compact support — and for a polynomial the identity is elementary and is obtained directly from \(\comm{P}{Q}=-\ii\hbar\) by induction. For a general bounded Borel \(g\) the identity still holds, with \(g(Q)\) read through the functional calculus of Definition 12.60 extended to unbounded operators in Stone's Theorem on One-Parameter Unitary Groups; that extension is not carried out here, and where a chapter needs it the fact is flagged.
Let \(f\) be finite and let \((U,V)\) be a Weyl system of \(f\) degrees of freedom, in the sense of Definition 12.109, on a separable Hilbert space \(\mathcal{H}\), acting irreducibly (Definition 12.90). Then there is a unitary \(T:\mathcal{H}\longrightarrow L^{2}(\R^{f})\) carrying it to the Schrödinger system Equation (12.72), and \(T\) is unique up to a phase. Without the irreducibility assumption, every such Weyl system on a separable space is a direct sum (Definition 12.84) of copies of the Schrödinger system. Rests on Definition 12.109, Definition 12.90 and Theorem 12.91.
Theorem 12.114 is quoted, not proved. The argument builds from the Weyl operators a weighted average \(\Pi=\iint w(\alpha,\beta)\,U(\alpha)V(\beta)\,\dd\alpha\,\dd\beta\) with a Gaussian weight \(w\) chosen so that the phase Equation (12.71) reproduces itself on multiplication, shows by that relation alone that \(\Pi^{2}=\Pi=\Pi^{\dagger}\) and that \(\Pi\) is not zero, deduces from irreducibility and Theorem 12.91 that its range is one-dimensional, and then shows that the Weyl operators applied to a unit vector of that range generate \(\mathcal{H}\) and reproduce the Schrödinger action — whence the unitary. It is functional analysis about the Weyl algebra, and the treatise takes the statement from [Reed:1972], theorem VIII.14, where the proof is given in full; the uniqueness question was von Neumann's, raised by Weyl's formulation of the commutation relations in exponentiated form. What is proved here is everything Equation (12.71) yields directly — the covariance Equation (12.73), the commutator Equation (12.77), and the Schur step Theorem 12.91 on which the quoted argument turns — so exactly one thing is owed, and it is named.
The hypothesis that \(f\) be finite is essential and its failure is physics rather than pathology. A field has infinitely many degrees of freedom, the theorem then fails, and inequivalent irreducible representations of the commutation relations exist in profusion; the choice among them becomes a physical question, and it is the origin of superselection sectors, of inequivalent vacua, and of the theorems of Axiomatic Quantum Field Theory. Within its hypotheses the theorem is what makes “the” Schrödinger representation a legitimate phrase: the quantum kinematics of a system with finitely many canonical pairs is fixed, up to unitary equivalence, by the commutation relations alone, which is the statement The Poisson Algebra and the Canonical Bridge to Quantum Mechanics needs when it asks what a quantization can and cannot be.