Lie Groups, Lie Algebras, and Fibre Bundles

Contents
  1. Lie groups
  2. Representations of Lie groups
  3. The Lie algebra associated with a Lie group
  4. Central extensions
  5. Contraction of Lie algebras
  6. Fibre bundles
  7. Invariant polynomials
  8. The Chern–Weil theorem and transgression forms

A continuous symmetry of a physical system is described by a group whose elements depend smoothly on a set of real parameters: a Lie group. This chapter introduces Lie groups and their generators, builds the universal enveloping algebra and the Casimir operators that label an irreducible representation, and works out the rotation group \(\SO(2,\R)\) as the fundamental example. It then constructs the three compact groups on which the Standard Model rests — \(\SO(3,\R)\), its double cover \(\SU(2)\), and \(\SU(3)\) — together with their irreducible representations, and the Lie algebras \(\mathfrak{t}(D)\) and \(\mathfrak{so}(p,q)\) with the algebras built from them: Lorentz, Poincaré, (anti-)de Sitter and conformal, in general dimension \(D = p + q\) and signature \((p,q)\). The two Casimir invariants of the Poincaré algebra, which are mass and spin, are obtained in the observed \(3+1\) dimensions. The closing sections develop invariant polynomials on a principal fibre bundle and prove the Chern–Weil theorem, from which the transgression and Chern–Simons forms follow.

Notation 14.1.

Throughout this chapter \(D = p + q\) denotes the dimension of the underlying pseudo-Euclidean space; capital Latin indices \(A, B, C, \ldots\) run from \(1\) to \(D\); and \(\eta_{AB} = \diag(\underbrace{+1,\ldots,+1}_{p},\underbrace{-1,\ldots,-1}_{q})\) is the flat metric of signature \((p,q)\), used to raise and lower indices, \(x_A = \eta_{AB}\,x^{B}\). The upright symbol \(\mathrm{D}\) denotes the covariant exterior derivative and must not be confused with the italic dimension \(D\).

Lie groups

Definition and generators

Lie group

Definition 14.2 (Lie group).

Consider an \(n\)-dimensional differentiable manifold \(G\). To denote the points of the manifold we introduce the differentiable operator \(g\), with \(g(\theta^{a}) \in G\), where the \(\theta^{a}\), \(a = 1, \ldots, M\), are real parameters. We say that \(G\) is a Lie group if and only if

  1. the algebraic structure \((G,\circ)\) is a group;

  2. the operator \(g\) is of class \(C^{\infty}\) in the parameters \(\theta^{a}\).

Thus, in order to characterise a continuous-group transformation, the transformation must be a function of the old coordinates and of the parameters \(\theta^{a}\); that is, continuous-group transformations take the form

\begin{equation}\tag{14.1} \bar{x}^{i} = f^{i}(\theta^{a}, x^{i})\ep \end{equation}

Generator of a Lie group

Consider an infinitesimal continuous-group transformation in a Lie group \(G\),

\begin{equation*} \bar{x}^{i} = f^{i}(\delta\theta^{a}, x^{i})\ep \end{equation*}

Since with no variation of the parameters the coordinates do not change, we may write

\begin{equation*} x^{i} = f^{i}(0, x^{i})\ec \end{equation*}

so the infinitesimal variation of the coordinates is given by

\begin{align} \delta\bar{x}^{i} &= \bar{x}^{i} - x^{i}\notag\\ &= f^{i}(\delta\theta^{a}, x^{i}) - f^{i}(0, x^{i})\notag\\ &\approx f^{i}(0, x^{i}) + \left.\pdv{f^{i}}{\theta^{a}}\right|_{\theta^{a}=0}\delta\theta^{a} - f^{i}(0, x^{i})\notag\\ &= \left.\pdv{f^{i}}{\theta^{a}}\right|_{\theta^{a}=0} \delta\theta^{a}\ep\tag{14.2} \end{align}

Consider now the variation of a field \(F = F(x)\) defined on \(G\) between two infinitesimally close points \(P\) and \(Q\):

\begin{align*} \delta F &= F(Q) - F(P)\\ &= F(x + \delta x) - F(x)\\ &\approx F(x) + \pdv{F}{x^{i}}\,\delta x^{i} - F(x)\\ &= \pdv{F}{x^{i}}\,\delta x^{i}\\ \text{[Equation (14.2)]}\quad &= \pdv{F}{x^{i}} \left.\pdv{f^{i}}{\theta^{a}}\right|_{\theta^{a}=0}\delta\theta^{a}\\ &= \left(\left.\pdv{f^{i}}{\theta^{a}}\right|_{\theta^{a}=0}\pp_{i}\right) F\,\delta\theta^{a}\ep \end{align*}

We define the generators of the Lie group as the operators

\begin{equation}\tag{14.3} J_{a} = \left.\pdv{f^{i}}{\theta^{a}}\right|_{\theta^{a}=0}\pp_{i}\ec \end{equation}

so that

\begin{equation}\tag{14.4} \delta F = J_{a} F\,\delta\theta^{a}\ep \end{equation}

From Equations (14.2) and (14.3) (insert the infinitesimal coordinate variation into the variation of the field and read off the operator acting on it).

The generators of a Lie group owe their name to the fact that they generate the variation of a field, as Equation (14.4) shows. They are mathematical objects that can always be assigned to a Lie group, and they are of fundamental importance in physics for the study of the symmetries of a system, as the physical parts of this treatise will show at length.

Geometric interpretation of the Lie bracket

In the preceding setting, let us examine what happens when we apply the variations along different paths leading from \(P\) to \(Q\). Let \(\theta^{a}_{1}\) and \(\theta^{a}_{2}\) be the two sets of parameters used to follow the two distinct paths. From Equation (14.4) we have

\begin{equation*} \delta_{1}F = J_{a}F\,\delta\theta^{a}_{1}\ec\qquad \delta_{2}F = J_{b}F\,\delta\theta^{b}_{2}\ep \end{equation*}

Then

\begin{equation*} \delta_{2}\delta_{1}F = J_{b}J_{a}F\,\delta\theta^{a}_{1}\delta\theta^{b}_{2}\ec\qquad \delta_{1}\delta_{2}F = J_{a}J_{b}F\,\delta\theta^{b}_{2}\delta\theta^{a}_{1}\ec \end{equation*}

so that the variation of the variations becomes

\begin{align*} \delta_{1}\delta_{2}F - \delta_{2}\delta_{1}F &= J_{a}J_{b}F\,\delta\theta^{b}_{2}\delta\theta^{a}_{1} - J_{b}J_{a}F\,\delta\theta^{a}_{1}\delta\theta^{b}_{2}\\ &= \comm{J_{a}}{J_{b}}F\,\delta\theta^{a}_{1}\delta\theta^{b}_{2}\ep \end{align*}

Hence, for the quantity \(\delta_{1}\delta_{2}F - \delta_{2}\delta_{1}F\) to be, indeed, a variation of \(F\) leading from \(P\) to \(Q\) — that is, of the form Equation (14.4) — the operator \(\comm{J_{a}}{J_{b}}\) must be a linear combination of the generators:

\begin{equation}\tag{14.5} \comm{J_{a}}{J_{b}} = C^{c}{}_{ab}\,J_{c}\ep \end{equation}

From Equation (14.4) (demanding that the commutator of two variations be itself a variation of the same form).

Exponentiation of the generators

A finite transformation is obtained by exponentiating the generators:

\begin{equation}\tag{14.6} g(\theta^{a}) = \ee^{J_{a}\theta^{a}}\ep \end{equation}

The source asserts Equation (14.6) without derivation. Two statements are hidden in it and they are of different kinds. That a one-parameter family of transformations closing under composition is an exponential is a theorem, proved next. That a many-parameter family can be written as the single exponential Equation (14.6) is a statement about the choice of parameters, and is discussed in Remark 14.4.

Proposition 14.3 (A one-parameter subgroup is an exponential).

Let \(\mathbb{V}\) be a finite-dimensional real or complex vector space and let \(t \mapsto g(t)\) be a differentiable curve of linear operators on \(\mathbb{V}\) satisfying

\begin{equation}\tag{14.7} g(s+t) = g(s)\,g(t)\ec\qquad g(0) = \identity\ep \end{equation}

Then, with \(J = \left.\dv{g}{t}\right|_{t=0}\),

\begin{equation}\tag{14.8} \boxed{g(t) = \ee^{tJ} = \sum_{n=0}^{\infty}\frac{t^{n}}{n!}\,J^{n}}\ec \end{equation}

the series converging absolutely for every \(t\). In particular \(g(t)\) is invertible, with \(g(t)^{-1} = g(-t)\). Rests on Definition 14.2 and Equation (14.3).

Proof.

Derives Proposition 14.3. Fix any submultiplicative norm on the operators of \(\mathbb{V}\), so that \(\norm{J^{n}} \le \norm{J}^{n}\). Then \(\sum_{n}\abs{t}^{n}\norm{J}^{n}/n! = \ee^{\abs{t}\norm{J}} < \infty\), so the series Equation (14.8) converges absolutely in the finite-dimensional — hence complete — space of operators, term by term differentiation is legitimate, and \(\dv{}{t}\ee^{tJ} = J\ee^{tJ} = \ee^{tJ}J\).

Differentiate Equation (14.7) with respect to \(s\) at \(s = 0\), holding \(t\) fixed. The left-hand side gives \(\dv{g}{t}(t)\) and the right-hand side gives \(g(t)\left.\dv{g}{s}\right|_{s=0} = g(t)J\), so

\begin{equation}\tag{14.9} \dv{g}{t} = g(t)\,J\ec\qquad g(0) = \identity\ep \end{equation}

Put \(h(t) = g(t)\ee^{-tJ}\). Since \(J\) commutes with \(\ee^{-tJ}\),

\begin{equation*} \dv{h}{t} = g(t)J\ee^{-tJ} - g(t)J\ee^{-tJ} = 0\ec \end{equation*}

so \(h\) is constant and equal to \(h(0) = \identity\). Multiplying on the right by \(\ee^{tJ}\) gives Equation (14.8). Finally \(g(t)g(-t) = g(0) = \identity\) by Equation (14.7).

Remark 14.4 (What the exponential formula asserts, and what it does not).

For several parameters, \(\ee^{\theta^{a}J_{a}}\) is a well-defined operator for every choice of the \(\theta^{a}\), and for fixed \(\theta^{a}\) the curve \(t \mapsto \ee^{t\theta^{a}J_{a}}\) is a one-parameter subgroup by Proposition 14.3. The map \(\theta^{a} \mapsto \ee^{\theta^{a}J_{a}}\) is smooth, and its differential at \(\theta^{a} = 0\) sends \(\theta^{a}\) to \(\theta^{a}J_{a}\), which is injective because the generators are linearly independent; by the inverse function theorem the map is therefore a diffeomorphism of a neighbourhood of the origin in parameter space onto a neighbourhood of the identity in \(G\). Equation (14.6) is the statement that the parameters \(\theta^{a}\) of Definition 14.2 have been chosen to be those coordinates — the canonical coordinates of the first kind. Three things it does not say. It does not say that parameters compose additively: unless the generators commute, \(\ee^{X}\ee^{Y} \neq \ee^{X+Y}\), the leading correction to the exponent being \(\tfrac{1}{2}\comm{X}{Y}\). It does not say that every element of a connected Lie group is a single exponential; what is true, and is all that Corollary 14.18 uses, is that the exponentials cover a neighbourhood of the identity and that a connected group is generated by any such neighbourhood. And when the generators are realized as the differential operators Equation (14.3) rather than as matrices, \(\ee^{\theta^{a}J_{a}}\) is to be read as the flow of the vector field \(\theta^{a}J_{a}\), which exists locally by the existence and uniqueness theorem for ordinary differential equations and obeys Equation (14.7) for that reason.

Alternative expression for the generators

From Equation (14.6) we see that, differentiating with respect to \(\theta^{a}\),

\begin{align*} \dv{g}{\theta^{a}} &= \ee^{J_{b}\theta^{b}} \dv{}{\theta^{a}}\left(J_{c}\theta^{c}\right)\\ &= \ee^{J_{b}\theta^{b}} J_{c}\,\delta^{c}{}_{a}\\ &= \ee^{J_{b}\theta^{b}} J_{a}\ec \end{align*}

and evaluating at \(\theta^{a} = 0\) we obtain

\begin{equation}\tag{14.10} J_{a} = \left.\dv{g}{\theta^{a}}\right|_{\theta^{a}=0}\ep \end{equation}

From Equation (14.6) (differentiate the exponential with respect to a parameter and evaluate at the identity).

The universal enveloping algebra

The bracket Equation (14.5) is not an associative product. The juxtaposition \(J_{a}J_{b}\) has no meaning in the Lie algebra itself; it acquires one only inside a representation, where the generators are operators and juxtaposition is composition. Yet the quantities that label the states of a physical system — the squared angular momentum, the squared four-momentum — are polynomials of exactly that kind. Before they can be discussed once and for all, rather than representation by representation, they must be given a home.

Definition 14.5 (Universal enveloping algebra).

Let \(\mathfrak{g}\) be a Lie algebra over \(\R\) with basis \(\set{T_{a}}\), \(a = 1,\ldots,\dim\mathfrak{g}\), and structure constants fixed as in Equation (14.5),

\begin{equation}\tag{14.11} \comm{T_{a}}{T_{b}} = C^{c}{}_{ab}\,T_{c}\ep \end{equation}

Its universal enveloping algebra \(U(\mathfrak{g})\) is the associative algebra with unit generated by the symbols \(T_{a}\) subject to the relations

\begin{equation}\tag{14.12} \boxed{T_{a}T_{b} - T_{b}T_{a} = C^{c}{}_{ab}\,T_{c}}\ec \end{equation}

equivalently the quotient of the tensor algebra of the vector space \(\mathfrak{g}\) by the two-sided ideal generated by all elements \(X\otimes Y - Y\otimes X - \comm{X}{Y}\) with \(X, Y \in \mathfrak{g}\). Rests on Equation (14.5).

Remark 14.6 (What $U(\mathfrak{g})$ is for).

Every linear representation \(D\) of \(\mathfrak{g}\) on a vector space \(\mathbb{V}\) — that is, every linear map with \(D\left(\comm{X}{Y}\right) = D(X)D(Y) - D(Y)D(X)\) — extends uniquely to a homomorphism of associative algebras \(U(\mathfrak{g}) \rightarrow \operatorname{End}(\mathbb{V})\), because Equation (14.12) is precisely the relation the operators \(D(T_{a})\) already satisfy. This is the sense in which \(U(\mathfrak{g})\) is universal, and it is what makes a statement proved once in \(U(\mathfrak{g})\) a statement about every representation at once.

Theorem 14.7 (Poincaré–Birkhoff–Witt).

The ordered monomials \(T_{1}^{n_{1}}T_{2}^{n_{2}}\cdots T_{M}^{n_{M}}\), with \(M = \dim\mathfrak{g}\) and \(n_{a} \in \N\cup\set{0}\), form a basis of \(U(\mathfrak{g})\) as a vector space. In particular \(\mathfrak{g}\) embeds in \(U(\mathfrak{g})\), and no relation holds among the generators beyond those forced by Equation (14.12). Rests on Definition 14.5.

That the ordered monomials span \(U(\mathfrak{g})\) is immediate from the defining relations, which allow any product to be reordered at the cost of terms of lower degree; the content of the theorem is their linear independence, and it is what guarantees that the Casimir element built below is not accidentally the zero element of the enveloping algebra.

Derives Theorem 14.7.

Casimir operators

Definition 14.8 (Casimir element).

A Casimir element, or Casimir operator, of a Lie algebra \(\mathfrak{g}\) is an element \(C \in U(\mathfrak{g})\) commuting with every generator,

\begin{equation}\tag{14.13} \boxed{\comm{C}{T_{a}} = 0\ec\qquad a = 1,\ldots,\dim\mathfrak{g}}\ec \end{equation}

equivalently an element of the centre \(Z\left(U(\mathfrak{g})\right)\). Rests on Definition 14.5.

A Casimir element is in general not an element of \(\mathfrak{g}\): for a semisimple algebra the centre of \(\mathfrak{g}\) itself is zero, and the invariants appear only at quadratic and higher degree in the generators. This is the whole reason for Definition 14.5.

Notation 14.9 (Casimirs against structure constants).

Casimir elements are written \(C_{2}, C_{3}, \ldots\), the single numerical subscript recording the degree of the polynomial in the generators. The structure constants always carry three algebra indices, \(C^{c}{}_{ab}\), so no ambiguity arises. Note also that Section 14.7 writes the same structure constants with the index pair first, \(\comm{\vect{T}_{A}}{\vect{T}_{B}} = C_{AB}{}^{C}\vect{T}_{C}\); only the index placement differs, and \(C_{ab}{}^{c} = C^{c}{}_{ab}\).

The Killing form and the quadratic Casimir

Every Lie algebra carries a canonical symmetric bilinear form, built from the structure constants alone, and it is this form that produces the Casimir element every physicist meets first.

Definition 14.10 (Killing form).

The Killing form of \(\mathfrak{g}\) is

\begin{equation}\tag{14.14} \boxed{\kappa(X,Y) = \tr\left(\ad_{X}\ad_{Y}\right)}\ec \qquad \ad_{X}Y = \comm{X}{Y}\ec \end{equation}

a symmetric bilinear form on \(\mathfrak{g}\). Since \(\left(\ad_{T_{a}}\right)^{c}{}_{b} = C^{c}{}_{ab}\) is the matrix of \(\ad_{T_{a}}\) in the basis \(\set{T_{a}}\), its components are

\begin{equation}\tag{14.15} \kappa_{ab} = \kappa(T_{a},T_{b}) = C^{c}{}_{ad}\,C^{d}{}_{bc}\ep \end{equation}

Rests on Equation (14.5).

Lemma 14.11 (Invariance of the Killing form).

For all \(X, Y, Z \in \mathfrak{g}\),

\begin{equation}\tag{14.16} \kappa\left(\comm{X}{Y},Z\right) + \kappa\left(Y,\comm{X}{Z}\right) = 0\ep \end{equation}

Rests on Definition 14.10.

Proof.

Derives Lemma 14.11. The adjoint map is itself a representation, \(\ad_{\comm{X}{Y}} = \comm{\ad_{X}}{\ad_{Y}}\), which is the Jacobi identity rewritten. Hence

\begin{align*} \kappa\left(\comm{X}{Y},Z\right) &= \tr\left(\comm{\ad_{X}}{\ad_{Y}}\ad_{Z}\right) = \tr\left(\ad_{X}\ad_{Y}\ad_{Z}\right) - \tr\left(\ad_{Y}\ad_{X}\ad_{Z}\right)\ec\\ \kappa\left(Y,\comm{X}{Z}\right) &= \tr\left(\ad_{Y}\comm{\ad_{X}}{\ad_{Z}}\right) = \tr\left(\ad_{Y}\ad_{X}\ad_{Z}\right) - \tr\left(\ad_{Y}\ad_{Z}\ad_{X}\right)\ep \end{align*}

Adding the two lines, the terms \(\mp\tr(\ad_{Y}\ad_{X}\ad_{Z})\) cancel against each other, and the survivors \(\tr(\ad_{X}\ad_{Y}\ad_{Z})\) and \(-\tr(\ad_{Y}\ad_{Z}\ad_{X})\) cancel by cyclicity of the trace.

Corollary 14.12 (Total antisymmetry of the structure constants).

Lower the upper index of the structure constants with the Killing form,

\begin{equation}\tag{14.17} C_{abc} := C^{d}{}_{ab}\,\kappa_{dc} = \kappa\left(\comm{T_{a}}{T_{b}},T_{c}\right)\ep \end{equation}

Then \(C_{abc}\) is antisymmetric under the exchange of any two of its three indices. Rests on Lemma 14.11 and Definition 14.10.

Proof.

Derives Corollary 14.12. Antisymmetry in the first two indices is the antisymmetry of the bracket. For the last two, put \(X = T_{a}\), \(Y = T_{b}\), \(Z = T_{c}\) in Equation (14.16) and use the symmetry of \(\kappa\):

\begin{equation*} C_{abc} = \kappa\left(\comm{T_{a}}{T_{b}},T_{c}\right) = -\kappa\left(T_{b},\comm{T_{a}}{T_{c}}\right) = -\kappa\left(\comm{T_{a}}{T_{c}},T_{b}\right) = -C_{acb}\ep \end{equation*}

The transpositions \((ab)\) and \((bc)\) generate the whole symmetric group on three letters, so \(C_{abc}\) changes sign under every transposition.

Theorem 14.13 (Cartan's criterion).

A finite-dimensional Lie algebra over a field of characteristic zero is semisimple — it contains no nonzero abelian ideal — if and only if its Killing form is nondegenerate. Rests on Definition 14.10.

One direction — that a nonzero abelian ideal lies in the radical of the Killing form — is a short index computation; the converse, that a degenerate Killing form produces a solvable ideal, goes through Cartan's criterion for solvability and the Jordan decomposition of an endomorphism, and is what makes the theorem long. It is used here to know when the Killing form may be inverted, and again in Whitehead's lemmas in the section on central extensions.

Derives Theorem 14.13.

When \(\kappa_{ab}\) is nondegenerate we write \(\kappa^{ab}\) for its inverse, \(\kappa^{ab}\kappa_{bc} = \delta^{a}{}_{c}\), and use it freely to raise algebra indices.

Definition 14.14 (Invariant symmetric tensor).

A totally symmetric array \(k^{a_{1}\cdots a_{r}}\) is invariant if, for every \(c\),

\begin{equation}\tag{14.18} \boxed{\sum_{l=1}^{r} C^{a_{l}}{}_{bc}\, k^{a_{1}\cdots a_{l-1}\,b\,a_{l+1}\cdots a_{r}} = 0}\ep \end{equation}

Rests on Equation (14.5).

Remark 14.15 (This is the invariance already proved).

Condition Equation (14.18) is the infinitesimal invariance of Definition 14.109 with the indices raised: it is Equation (14.148), established as part 2 of Theorem 14.110, transported to the contravariant position by \(\kappa^{ab}\). On a semisimple algebra, where the Killing form may be inverted, the two notions therefore coincide, and every \(G\)-invariant polynomial of Section 14.7 supplies one of these tensors. Nothing new has to be proved to use them.

Theorem 14.16 (Invariant tensors give Casimir operators).

Let \(k^{a_{1}\cdots a_{r}}\) be a totally symmetric invariant tensor in the sense of Definition 14.14. Then

\begin{equation}\tag{14.19} \boxed{C_{r} = k^{a_{1}\cdots a_{r}}\, T_{a_{1}}T_{a_{2}}\cdots T_{a_{r}} \in U(\mathfrak{g})} \end{equation}

is a Casimir element. Rests on Definitions 14.5, 14.8 and 14.14.

Proof.

Derives Theorem 14.16. In the associative algebra \(U(\mathfrak{g})\) the commutator with a fixed element obeys the Leibniz rule, \(\comm{XY}{Z} = X\comm{Y}{Z} + \comm{X}{Z}Y\), so

\begin{align*} \comm{C_{r}}{T_{c}} &= k^{a_{1}\cdots a_{r}}\sum_{l=1}^{r} T_{a_{1}}\cdots T_{a_{l-1}}\comm{T_{a_{l}}}{T_{c}} T_{a_{l+1}}\cdots T_{a_{r}}\\ &= \sum_{l=1}^{r} k^{a_{1}\cdots a_{r}}\,C^{b}{}_{a_{l}c}\, T_{a_{1}}\cdots T_{a_{l-1}}\,T_{b}\,T_{a_{l+1}}\cdots T_{a_{r}}\ep \end{align*}

In the \(l\)-th term rename the two summation indices \(a_{l}\) and \(b\) into each other. This is a relabelling of dummy indices and moves no operator past another, so the ordered string becomes literally \(T_{a_{1}}T_{a_{2}}\cdots T_{a_{r}}\), the same string for every value of \(l\), while the coefficient becomes \(C^{a_{l}}{}_{bc}\,k^{a_{1}\cdots a_{l-1}\,b\,a_{l+1}\cdots a_{r}}\). The sum over \(l\) may therefore be taken inside:

\begin{equation*} \comm{C_{r}}{T_{c}} = \left(\sum_{l=1}^{r} C^{a_{l}}{}_{bc}\, k^{a_{1}\cdots a_{l-1}\,b\,a_{l+1}\cdots a_{r}}\right) T_{a_{1}}T_{a_{2}}\cdots T_{a_{r}}\ec \end{equation*}

and the parenthesis vanishes by Equation (14.18).

Corollary 14.17 (The quadratic Casimir).

Let \(\mathfrak{g}\) be semisimple, so that \(\kappa_{ab}\) is invertible by Theorem 14.13. Then

\begin{equation}\tag{14.20} \boxed{C_{2} = \kappa^{ab}\,T_{a}T_{b}} \end{equation}

commutes with every generator of \(\mathfrak{g}\). Rests on Theorem 14.16, Corollary 14.12 and Theorem 14.13.

Proof.

Derives Corollary 14.17. The statement is Theorem 14.16 for \(r=2\) and \(k^{ab} = \kappa^{ab}\), whose invariance is Corollary 14.12 with both free indices raised. It is worth seeing the two-index case done by hand, because it is the whole mechanism in three lines. By the Leibniz rule,

\begin{equation*} \comm{C_{2}}{T_{c}} = \kappa^{ab}\left(T_{a}\comm{T_{b}}{T_{c}} + \comm{T_{a}}{T_{c}}T_{b}\right) = \kappa^{ab}C^{d}{}_{bc}\,T_{a}T_{d} + \kappa^{ab}C^{d}{}_{ac}\,T_{d}T_{b}\ep \end{equation*}

Fix \(c\) and set

\begin{equation*} M^{ad} := \kappa^{ab}\,C^{d}{}_{bc} = \kappa^{ab}\kappa^{de}\,C_{bce}\ec \end{equation*}

using Equation (14.17). Exchanging the two free indices and then renaming the dummies \(b \leftrightarrow e\),

\begin{equation*} M^{da} = \kappa^{db}\kappa^{ae}\,C_{bce} = \kappa^{de}\kappa^{ab}\,C_{ecb} = -\kappa^{ab}\kappa^{de}\,C_{bce} = -M^{ad}\ec \end{equation*}

the last step by Corollary 14.12, since \(C_{ecb}\) differs from \(C_{bce}\) by the transposition of the first and third indices. The first term above is \(M^{ad}T_{a}T_{d}\); the second is \(\kappa^{ab}C^{d}{}_{ac}T_{d}T_{b} = M^{bd}T_{d}T_{b}\) by the symmetry of \(\kappa\), which is \(M^{ad}T_{d}T_{a}\) after renaming \(b \rightarrow a\). Hence

\begin{equation*} \comm{C_{2}}{T_{c}} = M^{ad}\left(T_{a}T_{d} + T_{d}T_{a}\right) = 0\ec \end{equation*}

an antisymmetric array contracted with a symmetric one.

Corollary 14.18 (A Casimir acts as a number on an irreducible representation).

Let \(G\) be a connected Lie group with Lie algebra \(\mathfrak{g}\), let \(U\) be a finite-dimensional irreducible representation of \(G\) on a complex vector space \(\mathbb{V}\), let \(D\) be the induced representation of \(\mathfrak{g}\) as in Equation (14.6), and let \(C\) be a Casimir element. Then

\begin{equation}\tag{14.21} \boxed{D(C) = c\,\identity\ec\qquad c \in \C}\ep \end{equation}

Rests on Definition 14.8, Remark 14.6, Equation (14.6) and Theorem 5.153.

Proof.

Derives Corollary 14.18. By Remark 14.6 the representation \(D\) extends to \(U(\mathfrak{g})\), so \(D(C)\) is a well-defined operator on \(\mathbb{V}\) and \(\comm{D(C)}{D(T_{a})} = D\left(\comm{C}{T_{a}}\right) = 0\) for every \(a\) by Equation (14.13). An operator commuting with \(D(T_{a})\) commutes with every power of \(D(T_{a})\), hence with the exponential \(\ee^{\theta^{a}D(T_{a})} = U\left(g(\theta^{a})\right)\) of Equation (14.6), that is, with \(U(g)\) for every \(g\) in a neighbourhood of the identity. A connected topological group is generated by any neighbourhood of its identity, so \(D(C)\) commutes with \(U(g)\) for every \(g \in G\). Schur's first lemma (Theorem 5.153), applied to the irreducible representation \(U\), gives \(D(C) = c\,\identity\).

Remark 14.19 (A Casimir eigenvalue labels; one Casimir may not separate).

Equation (14.21) is what makes a Casimir useful. The number \(c\) takes the same value on every vector of the representation and is unchanged by any change of basis, so it is an invariant label of the equivalence class of the irreducible representation, and two representations with different values of \(c\) cannot be equivalent. The converse does not follow, and it is worth being precise about it rather than repeating the usual slogan. A second imprecision has to be removed first: there is no such thing as the quadratic Casimir, because any nonzero multiple of a Casimir is again one and its eigenvalue scales with it, so a numerical eigenvalue means something only once the normalisation is stated. For \(\mathfrak{su}(2)\), in the normalisation \(\vect{J}^{2} = J_{1}^{2}+J_{2}^{2}+J_{3}^{2}\) of Equation (14.49) — which by Lemma 14.41 is twice the Killing-form Casimir \(C_{2} = \kappa^{ab}J_{a}J_{b}\) of Equation (14.20) — the quadratic Casimir does separate the classes, because Theorem 14.43 gives its eigenvalue as \(j(j+1)\) and \(j \mapsto j(j+1)\) is injective for \(j \ge 0\). For \(\mathfrak{su}(3)\) it does not: the fundamental and the antifundamental representations of Proposition 14.55 share the value of \(C_{2}\) and are told apart only by a second, cubic invariant. How many invariants are needed in general is the subject of what follows.

Rank, and how many Casimir operators there are

Definition 14.20 (Cartan subalgebra and rank).

A Cartan subalgebra \(\mathfrak{h} \subset \mathfrak{g}\) of a semisimple Lie algebra is a maximal abelian subalgebra whose elements all act diagonalizably in the adjoint representation. Its dimension is the same for every such subalgebra and is called the rank \(\ell\) of \(\mathfrak{g}\).

Theorem 14.21 (Racah).

For a finite-dimensional semisimple Lie algebra of rank \(\ell\), the centre \(Z\left(U(\mathfrak{g})\right)\) is a polynomial algebra in \(\ell\) algebraically independent Casimir elements. A complete set of invariant labels for its finite-dimensional irreducible representations therefore consists of \(\ell\) numbers, and of no fewer. Rests on Definitions 14.8 and 14.20.

The derivation identifies the centre of the enveloping algebra with the adjoint-invariant polynomials on the algebra, by the symmetrization map, and then those with the Weyl-invariant polynomials on a Cartan subalgebra; it reduces the theorem to two results of the structure theory of semisimple Lie algebras — the Chevalley restriction theorem and Chevalley's theorem on finite reflection groups — which this treatise quotes rather than proves, and the appendix says exactly where. The systematic use of such invariants as representation labels in spectroscopy is Racah's.

Derives Theorem 14.21.

The rank is read off the algebra without difficulty in the cases this treatise needs. For \(\mathfrak{so}(3) \cong \mathfrak{su}(2)\) any single generator spans a maximal abelian subalgebra, so \(\ell = 1\) and there is exactly one Casimir: it is the \(\vect{J}^{2}\) constructed in Definition 14.40. For \(\mathfrak{su}(3)\) the diagonal traceless matrices give \(\ell = 2\), hence two Casimirs, one quadratic and one cubic (Proposition 14.54). The remaining family this treatise needs is \(\mathfrak{so}(p,q)\), and because the count of Lorentz invariants rests on it the statement is proved rather than quoted.

Proposition 14.22 (The rank of $\mathfrak{so}(p,q)$).

Let \(D = p+q \ge 3\). Whatever the signature,

\begin{equation}\tag{14.22} \boxed{\ell\left(\mathfrak{so}(p,q)\right) = \left\lfloor\frac{D}{2}\right\rfloor}\ep \end{equation}

In particular the Lorentz algebra \(\mathfrak{so}(3,1)\) of the observed spacetime has rank \(2\), and therefore, by Theorem 14.21, exactly two Casimir operators. Rests on Definition 14.20, Equation (14.90) and Theorem 14.21.

Proof.

Derives Proposition 14.22. Write \(r = \lfloor D/2 \rfloor\), work in the defining representation on \(\C^{D}\), and let \(\eta\) be diagonal as in Notation 14.1, so that \(X \in \mathfrak{so}(p,q)\) means

\begin{equation}\tag{14.23} \eta(Xu,v) + \eta(u,Xv) = 0\qquad\text{for all }u,v\ep \end{equation}

Diagonalizability in the defining representation is what is used throughout; that it agrees with the diagonalizability in the adjoint representation asked for by Definition 14.20 is standard structure theory, quoted here as Theorem 14.13 is. One direction is immediate — in a basis diagonalizing \(X\) the map \(Y \mapsto \comm{X}{Y}\) is diagonal on the matrix units, with eigenvalue the difference of two eigenvalues of \(X\) — and the other is that the semisimple and nilpotent parts of an element of a semisimple linear Lie algebra again lie in it, so that an element with diagonalizable adjoint action has central, hence zero, nilpotent part.

A commuting set of size \(r\). Take

\begin{equation}\tag{14.24} H_{k} = J_{(2k-1)(2k)}\ec\qquad k = 1,\ldots,r\ec \end{equation}

one generator for each of the \(r\) disjoint index pairs \(\set{1,2}\), \(\set{3,4}\), and so on. No two of them share an index, so every \(\eta\) on the right of Equation (14.90) carries two distinct indices and vanishes, \(\eta\) being diagonal; hence \(\comm{H_{k}}{H_{l}} = 0\). In the defining representation \(J_{AB}\) is the matrix \(\left(J_{AB}\right)^{C}{}_{D} = \delta^{C}_{A}\eta_{BD} - \delta^{C}_{B}\eta_{AD}\), the identification already made for \(\mathfrak{so}(3)\) at Equation (14.32): it acts inside the plane of its own two coordinates and annihilates the others, and on that plane it is either the rotation generator Equation (14.27), with eigenvalues \(\pm\ii\), or a boost generator, with eigenvalues \(\pm1\); in both cases it is diagonalizable, and so is any linear combination of the \(H_{k}\), the blocks being independent. The span of the \(H_{k}\) is therefore an abelian subalgebra of diagonalizable elements, of dimension \(r\).

No larger one exists. Let \(\mathfrak{h}\) be any abelian subalgebra of \(\mathfrak{so}(p,q)\) all of whose elements are diagonalizable, and let \(\mathfrak{h}_{\C}\) be its complex span inside the complex \(D\times D\) matrices. Since the elements of \(\mathfrak{h}\) are real matrices, a real basis of \(\mathfrak{h}\) stays independent over \(\C\) — the real and imaginary parts of a vanishing complex combination vanish separately — so \(\dim_{\C}\mathfrak{h}_{\C} = \dim_{\R}\mathfrak{h}\). Commuting diagonalizable operators are simultaneously diagonalizable, and every complex combination of operators diagonal in one basis is diagonal in that basis, so there are a basis \(\set{v_{1},\ldots,v_{D}}\) of \(\C^{D}\) and complex-linear functionals \(\lambda_{i}\) on \(\mathfrak{h}_{\C}\) with \(Xv_{i} = \lambda_{i}(X)v_{i}\). Condition Equation (14.23) is complex-bilinear and therefore holds on \(\mathfrak{h}_{\C}\); applied to \(u = v_{i}\), \(v = v_{j}\) it gives

\begin{equation*} \left(\lambda_{i}(X)+\lambda_{j}(X)\right)\eta(v_{i},v_{j}) = 0 \qquad\text{for all }X \in \mathfrak{h}_{\C}\ep \end{equation*}

Collect the basis vectors by their functional: for each \(\alpha\) occurring among the \(\lambda_{i}\), let \(V_{\alpha}\) be the span of those \(v_{i}\) with \(\lambda_{i} = \alpha\). The display says \(\eta(V_{\alpha},V_{\beta}) = 0\) unless \(\beta = -\alpha\), so nondegeneracy of \(\eta\) forces \(-\alpha\) to occur whenever \(\alpha\) does, and forces the pairing of \(V_{\alpha}\) with \(V_{-\alpha}\) to be nondegenerate, whence \(\dim V_{\alpha} = \dim V_{-\alpha}\). The occurring functionals thus consist of \(k\) pairs \(\pm\alpha_{1},\ldots,\pm\alpha_{k}\) with \(\alpha_{i} \neq 0\), possibly together with \(0\); each pair accounts for at least two basis vectors, so \(2k \le D\) and \(k \le r\). Finally the occurring functionals span the dual of \(\mathfrak{h}_{\C}\), since an \(X\) annihilated by all of them acts as zero on every \(v_{i}\) and is therefore zero, and they span the same space as \(\alpha_{1},\ldots,\alpha_{k}\). Hence \(\dim_{\R}\mathfrak{h} = \dim_{\C}\mathfrak{h}_{\C} \le k \le r\).

The subalgebra Equation (14.24) is therefore maximal, and every maximal one has the same dimension \(r\).

Remark 14.23 (Racah's theorem does not apply to every algebra that matters).

Theorem 14.21 is a statement about semisimple algebras, and the Poincaré algebra Equation (14.99) is not semisimple: it is a semidirect sum with an abelian ideal, so its Killing form is degenerate and Theorem 14.13 excludes it. Its invariants must be found by hand, which is done in Theorem 14.65, and the count that comes out — two, in the observed \(3+1\) dimensions — is not predicted by Racah's theorem but merely happens to agree with the rank of its Lorentz subalgebra. Conflating the two is an easy error and would give the wrong answer for the Galilei algebra of Example 14.80, whose invariants include the central mass.

Remark 14.24 (Casimir operators and SI units).

Everything in this subsection is dimensionless. The generators of Equation (14.3) are differential operators built from coordinates and derivatives with respect to them, the structure constants are pure numbers, and a Casimir element is a polynomial in the generators, so its eigenvalue on an irreducible representation is a pure number too. A physical observable is obtained by multiplying a generator by the constant that carries its dimension, and for the rotation generators that constant is the reduced Planck constant \(\hbar = h/2\pi\), with \(h = 6.62607015\times 10^{-34}\,\mathrm{J}\,\mathrm{s}\) exactly, by the definition of the kilogram [BIPM:2019]. The observable angular momentum is \(\hbar\vect{J}\), of SI unit \(\mathrm{J}\,\mathrm{s}\), and the observable Casimir is \(\hbar^{2}\vect{J}^{2}\), of SI unit \(\mathrm{J}^{2}\,\mathrm{s}^{2}\). It is worth saying this once and plainly, because the literature almost never does: the familiar statement “the eigenvalue is \(j(j+1)\)” is a statement about a dimensionless generator, and the measured quantity is \(\hbar^{2}j(j+1)\). The same distinction returns for the Poincaré algebra in Proposition 14.60, where the first invariant is \(m^{2}c^{2}\) and not \(m^{2}\).

Representations of Lie groups

The group $\SO(2,\R)$

Discussion of the definition

Consider the vector space \(\R^{2} = \operatorname{span}\set{\ket{e_{x}}, \ket{e_{y}}}\) and let \(R(\phi)\) be the operator that takes the basis vector \(\ket{e_{i}}\) to the rotated vector \(R(\phi)\ket{e_{i}}\) [Goldstein:2002].

Elementary trigonometry on the rotated basis gives

\begin{align*} R(\phi)\ket{e_{x}} &= \cos\phi\,\ket{e_{x}} + \sin\phi\,\ket{e_{y}}\ec\\ R(\phi)\ket{e_{y}} &= -\sin\phi\,\ket{e_{x}} + \cos\phi\,\ket{e_{y}}\ec \end{align*}

or, compactly,

\begin{equation*} R(\phi)\ket{e_{i}} = \left(D(\phi)\right)^{j}{}_{i}\ket{e_{j}}\ec \end{equation*}

with

\begin{equation}\tag{14.25} D(\phi) = \begin{pmatrix} \cos\phi & -\sin\phi\\ \sin\phi & \cos\phi \end{pmatrix}\ep \end{equation}

Now consider an arbitrary vector \(\ket{x} \in \R^{2}\). In the basis above, \(\ket{x} = x^{i}\ket{e_{i}}\). As we have seen, the operator \(R(\phi)\) rotates a vector when acting upon it:

\begin{align*} \ket{\bar{x}} &= R(\phi)\ket{x}\ec\\ \bar{x}^{j}\ket{e_{j}} &= R(\phi)\,x^{i}\ket{e_{i}} = x^{i}\left(D(\phi)\right)^{j}{}_{i}\ket{e_{j}}\ec \end{align*}

so that

\begin{equation*} \bar{x}^{j} = \left(D(\phi)\right)^{j}{}_{i}\,x^{i}\ep \end{equation*}

But does the rotation of the vector preserve its norm? This seems obvious; let us see what happens if we impose it. We require

\begin{align*} \norm{\ket{x}} &\stackrel{!}{=} \norm{\ket{\bar{x}}}\ec\\ x^{i}x^{i} &= \bar{x}^{j}\bar{x}^{j}\\ &= \left(D(\phi)\right)^{j}{}_{k}\left(D(\phi)\right)^{j}{}_{l}\,x^{k}x^{l}\\ &= \left(D(\phi)\transpose\right)^{k}{}_{j}\left(D(\phi)\right)^{j}{}_{l}\, x^{k}x^{l}\ec \end{align*}

whence

\begin{equation*} \left(D(\phi)\transpose\right)^{k}{}_{j}\left(D(\phi)\right)^{j}{}_{l} = \delta^{k}{}_{l}\ec \end{equation*}

or, in matrix notation,

\begin{equation}\tag{14.26} D(\phi)\transpose D(\phi) = \identity\ep \end{equation}

In conclusion, we have found the condition for the operator \(R(\phi)\) to be an automorphism belonging to the group \(\Ogrp(2,\R)\). Moreover,

\begin{equation*} \det D(\phi) = \cos^{2}\phi + \sin^{2}\phi = 1\ec \end{equation*}

and therefore \(R(\phi) \in \SO(2,\R)\).

With this we can conclude the following fact: the set of all two-dimensional rotations forms a group, which is isomorphic to the group \(\SO(2,\R)\). Furthermore, one can see geometrically that rotating a vector by an angle \(\phi_{1}\) and then by an angle \(\phi_{2}\) is exactly the same as doing it the other way around — first by \(\phi_{2}\) and then by \(\phi_{1}\) — which means that \(\SO(2,\R)\) is an abelian group.

Generator

We have

\begin{align*} D(\phi) &= \begin{pmatrix} \cos\phi & -\sin\phi\\ \sin\phi & \cos\phi \end{pmatrix}\\ &= \begin{pmatrix} \cos\phi & 0\\ 0 & \cos\phi \end{pmatrix} + \begin{pmatrix} 0 & -\sin\phi\\ \sin\phi & 0 \end{pmatrix}\\ &= \cos\phi \begin{pmatrix} 1 & 0\\ 0 & 1 \end{pmatrix} + \sin\phi \begin{pmatrix} 0 & -1\\ 1 & 0 \end{pmatrix}\ep \end{align*}

Defining the matrix

\begin{equation}\tag{14.27} J = \begin{pmatrix} 0 & -1\\ 1 & 0 \end{pmatrix}\ec \end{equation}

we may write

\begin{equation*} D(\phi) = \identity\cos\phi + J\sin\phi\ep \end{equation*}

Note that

\begin{equation*} J^{2} = \begin{pmatrix} 0 & -1\\ 1 & 0 \end{pmatrix} \begin{pmatrix} 0 & -1\\ 1 & 0 \end{pmatrix} = \begin{pmatrix} -1 & 0\\ 0 & -1 \end{pmatrix} = -\identity\ec \end{equation*}

and thus

\begin{equation*} J^{3} = -J\ec\qquad J^{4} = \identity\ec\qquad J^{5} = J\ec \end{equation*}

so that, for \(n \in \N\),

\begin{equation*} J^{2n} = (-1)^{n}\,\identity\ec\qquad J^{2n+1} = (-1)^{n}\,J\ep \end{equation*}

On the other hand, recalling the Taylor series of the cosine and the sine,

\begin{equation*} \cos\phi = \sum^{\infty}_{n=0}\frac{(-1)^{n}}{(2n)!}\,\phi^{2n}\ec\qquad \sin\phi = \sum^{\infty}_{n=0}\frac{(-1)^{n}}{(2n+1)!}\,\phi^{2n+1}\ec \end{equation*}

we obtain

\begin{align*} D(\phi) &= \sum^{\infty}_{n=0}\frac{(-1)^{n}\,\identity}{(2n)!}\,\phi^{2n} + \sum^{\infty}_{n=0}\frac{(-1)^{n}\,J}{(2n+1)!}\,\phi^{2n+1}\\ &= \sum^{\infty}_{n=0}\frac{J^{2n}}{(2n)!}\,\phi^{2n} + \sum^{\infty}_{n=0}\frac{J^{2n+1}}{(2n+1)!}\,\phi^{2n+1}\\ &= \sum^{\infty}_{n=0}\frac{J^{n}}{n!}\,\phi^{n}\ec \end{align*}

whence

\begin{equation}\tag{14.28} D(\phi) = \ee^{J\phi}\ec \end{equation}

so \(J\) is the generator of the group \(\SO(2,\R)\).

Irreducible representations

We know that the group \(\SO(2,\R)\) is abelian, and we have seen that the irreducible representations of an abelian group are necessarily one-dimensional — Proposition 5.156, a consequence of the first Schur lemma (Theorem 5.153). Which, then, are the possible irreducible representations? Let \(\ket{\lambda}\) be an eigenvector of \(J\),

\begin{equation*} J\ket{\lambda} = \lambda\ket{\lambda}\ec \end{equation*}

so that, by Equation (14.28),

\begin{align*} R(\phi)\ket{\lambda} &= \ee^{J\phi}\ket{\lambda}\\ &= \ee^{\lambda\phi}\ket{\lambda}\ep \end{align*}

But since one must have

\begin{equation*} R(\phi)\ket{x} = R(\phi + 2\pi)\ket{x}\ec \end{equation*}

it follows that

\begin{align*} \ee^{\lambda\phi}\ket{\lambda} &= \ee^{\lambda(\phi+2\pi)}\ket{\lambda}\\ &= \ee^{\lambda\phi}\,\ee^{2\pi\lambda}\ket{\lambda} \implies \ee^{2\pi\lambda} = 1\ec \end{align*}

so that

\begin{equation*} \lambda = \ii m\ec\qquad m \in \Z\ep \end{equation*}

Therefore, the irreducible representations of \(\SO(2,\R)\) are

\begin{equation}\tag{14.29} R^{m}(\phi) = \ee^{\ii m\phi}\ec\qquad m \in \Z\ep \end{equation}

From Equation (14.28) (diagonalize the generator and impose \(2\pi\)-periodicity of the rotation).

The groups $\SO(3)$, $\SU(2)$ and $\SU(3)$

Three compact groups carry almost all of the group theory the physical parts of this treatise need. The rotation group \(\SO(3,\R)\) is the symmetry of space itself. The special unitary group \(\SU(2)\) is its double cover, and the existence of that cover — not any postulate about particles — is why half-integer angular momentum exists. The group \(\SU(3)\) is the first case here in which the rank exceeds one, so that two Casimir operators are required by Theorem 14.21 and the weight diagram becomes two-dimensional. The gauge group of the Standard Model is \(\SU(3)\times\SU(2)\times\U(1)\), and this subsection supplies the mathematics that Quantum Chromodynamics and Electroweak Unification and the Higgs Boson will use; the physics, and every claim about what the groups act on, is left to them. The quantum-mechanical treatment of angular momentum, which starts from the algebra derived here, is Angular Momentum and Spin [Sakurai:2017] [CohenTannoudji:1977].

Remark 14.25 ($\SU(3)$ is an internal symmetry).

This treatise instantiates every spacetime construction in the observed \(3+1\) dimensions. The \(3\) of \(\SU(3)\) is not a spacetime dimension: it is the dimension of an internal complex vector space attached to each point, on which the group acts without touching the coordinates at all. Nothing in this subsection is a statement about a three-dimensional spacetime.

The rotation group $\SO(3,\R)$

Definition 14.26 (The rotation group).
\begin{equation}\tag{14.30} \SO(3,\R) = \set{R \in \GL(3,\R) \mid R\transpose R = \identity,\ \det R = 1}\ep \end{equation}

This is the case \(p = 3\), \(q = 0\) of Section 14.3.2: the defining condition \(R\transpose\eta R = \eta\) of the pseudo-orthogonal group reduces to \(R\transpose R = \identity\) when \(\eta = \identity\), and \(\det R = 1\) selects the component containing the identity. Accordingly Equation (14.89) gives

\begin{equation}\tag{14.31} \dim\mathfrak{so}(3) = \frac{3 \cdot 2}{2} = 3\ep \end{equation}

From Equation (14.89) (the case \(p=3\), \(q=0\), so that \(D=3\)).

Proposition 14.27 (Generators of $\mathfrak{so}(3)$).

The Lie algebra \(\mathfrak{so}(3)\) is the space of real antisymmetric \(3\times3\) matrices. A basis of it is

\begin{equation}\tag{14.32} \left(L_{k}\right)_{ij} = -\epsilon_{kij}\ec \end{equation}

that is

\begin{equation}\tag{14.33} L_{1} = \begin{pmatrix} 0 & 0 & 0\\ 0 & 0 & -1\\ 0 & 1 & 0\end{pmatrix} \ec\quad L_{2} = \begin{pmatrix} 0 & 0 & 1\\ 0 & 0 & 0\\ -1 & 0 & 0\end{pmatrix} \ec\quad L_{3} = \begin{pmatrix} 0 & -1 & 0\\ 1 & 0 & 0\\ 0 & 0 & 0\end{pmatrix} \ec \end{equation}

and it satisfies

\begin{equation}\tag{14.34} \comm{L_{i}}{L_{j}} = \epsilon_{ijk}\,L_{k}\ep \end{equation}

Rests on Definition 14.26.

Proof.

Derives Proposition 14.27. Write \(R = \identity + \varepsilon X + O(\varepsilon^{2})\) in Equation (14.30). To first order \(R\transpose R = \identity\) gives \(X\transpose + X = 0\), and \(\det R = 1\) gives \(\tr X = 0\), which is implied by antisymmetry; conversely \(\exp\) of an antisymmetric matrix is orthogonal with unit determinant. The real antisymmetric \(3\times3\) matrices form a three-dimensional space, matching Equation (14.31), and the \(L_{k}\) of Equation (14.32) are three independent members of it. Note that \(L_{3}\) is the generator \(J\) of \(\SO(2,\R)\) found in Equation (14.27), embedded in the \(1\)–\(2\) block.

For the bracket, compute the \((m,n)\) entry using \(\left(L_{i}\right)_{mk} = -\epsilon_{imk}\):

\begin{align*} \left(\comm{L_{i}}{L_{j}}\right)_{mn} &= \epsilon_{imk}\epsilon_{jkn} - \epsilon_{jmk}\epsilon_{ikn}\\ &= \left(\delta_{in}\delta_{mj} - \delta_{ij}\delta_{mn}\right) - \left(\delta_{jn}\delta_{mi} - \delta_{ij}\delta_{mn}\right)\\ &= \delta_{in}\delta_{jm} - \delta_{im}\delta_{jn}\ec \end{align*}

where the contraction \(\epsilon_{kim}\epsilon_{knj} = \delta_{in}\delta_{mj} - \delta_{ij}\delta_{mn}\) was used twice after a cyclic relabelling. On the other hand

\begin{equation*} \left(\epsilon_{ijk}L_{k}\right)_{mn} = -\epsilon_{ijk}\epsilon_{kmn} = -\left(\delta_{im}\delta_{jn} - \delta_{in}\delta_{jm}\right)\ec \end{equation*}

which is the same expression.

For representation theory it is convenient to use Hermitian generators, because Hermitian operators are the ones that can represent observables and because the resulting brackets are those of the quantum theory.

Definition 14.28 (Hermitian rotation generators).
\begin{equation}\tag{14.35} \boxed{J_{k} = \ii L_{k}\ec\qquad \comm{J_{i}}{J_{j}} = \ii\,\epsilon_{ijk}J_{k}\ec\qquad R(\hat{n},\phi) = \exp\left(-\ii\,\phi\,\hat{n}\cdot\vect{J}\right)}\ep \end{equation}

Each \(J_{k}\) is Hermitian: \(L_{k}\) is real and antisymmetric, so \(\left(\ii L_{k}\right)^{\dagger} = -\ii L_{k}\transpose = \ii L_{k}\). The bracket follows from Equation (14.34) by \(\comm{J_{i}}{J_{j}} = -\comm{L_{i}}{L_{j}} = -\epsilon_{ijk}L_{k} = \ii\epsilon_{ijk}J_{k}\). Rests on Proposition 14.27.

Remark 14.29 (Where $\hbar$ enters).

Equation (14.35) is a statement about dimensionless matrices. The physical angular momentum operator is \(\hat{\vect{\mathcal{J}}} = \hbar\vect{J}\), of SI unit \(\mathrm{J}\,\mathrm{s}\), and multiplying Equation (14.35) by \(\hbar^{2}\) turns it into \(\comm{\hat{\mathcal{J}}_{i}}{\hat{\mathcal{J}}_{j}} = \ii\hbar\,\epsilon_{ijk}\hat{\mathcal{J}}_{k}\), which is the commutator measured in the laboratory and the one used throughout Angular Momentum and Spin and Particles as Poincaré Representations (cf. Equation (94.8)). The rotation operator is correspondingly \(\exp\left(-\ii\,\phi\,\hat{n}\cdot\hat{\vect{\mathcal{J}}}/\hbar\right)\), the ratio in the exponent being dimensionless as it must be. See Remark 14.24.

Proposition 14.30 (Rodrigues formula; the exponential map is onto).

Write \(N = \hat{n}\cdot\vect{L}\) for a unit vector \(\hat{n}\). Then \(N\vect{x} = \hat{n}\times\vect{x}\), \(N^{3} = -N\), and

\begin{equation}\tag{14.36} \exp(\phi N) = \identity + \sin\phi\;N + (1-\cos\phi)\,N^{2}\ec \end{equation}

so that

\begin{equation}\tag{14.37} R(\hat{n},\phi)\,\vect{x} = \vect{x}\cos\phi + \hat{n}\left(\hat{n}\cdot\vect{x}\right)\left(1-\cos\phi\right) + \left(\hat{n}\times\vect{x}\right)\sin\phi\ep \end{equation}

Every element of \(\SO(3,\R)\) is of this form with \(\phi \in [0,\pi]\), so the exponential map \(\mathfrak{so}(3) \rightarrow \SO(3,\R)\) is surjective. Rests on Proposition 14.27, Equation (14.28) and Equation (14.25).

Proof.

Derives Proposition 14.30. From Equation (14.32), \(\left(N\vect{x}\right)_{i} = -\epsilon_{kij}n_{k}x_{j} = \epsilon_{ikj}n_{k}x_{j} = \left(\hat{n}\times\vect{x}\right)_{i}\). Hence \(N^{2}\vect{x} = \hat{n}\times(\hat{n}\times\vect{x}) = \hat{n}(\hat{n}\cdot\vect{x}) - \vect{x}\) and \(N^{3}\vect{x} = \hat{n}\times\left(\hat{n}(\hat{n}\cdot\vect{x}) - \vect{x}\right) = -\hat{n}\times\vect{x} = -N\vect{x}\). Consequently \(N^{2k+1} = (-1)^{k}N\) and \(N^{2k+2} = (-1)^{k}N^{2}\) for \(k \in \N\cup\set{0}\), and splitting the exponential series into its odd and even parts, exactly as was done for \(\SO(2,\R)\) at Equation (14.28),

\begin{equation*} \exp(\phi N) = \identity + \left(\sum_{k=0}^{\infty}\frac{(-1)^{k}\phi^{2k+1}}{(2k+1)!}\right)N + \left(\sum_{k=0}^{\infty}\frac{(-1)^{k}\phi^{2k+2}}{(2k+2)!}\right)N^{2} \ec \end{equation*}

whose two coefficients are \(\sin\phi\) and \(1 - \cos\phi\). Substituting \(N^{2}\vect{x}\) gives Equation (14.37).

Surjectivity is Euler's theorem on rotations: \(R \in \SO(3,\R)\) is a real \(3\times3\) matrix, so its characteristic polynomial is a real cubic and has a real root; the roots have modulus \(1\) because \(R\) is orthogonal, and their product is \(\det R = 1\), so the real root is \(+1\) and there is a unit vector \(\hat{n}\) with \(R\hat{n} = \hat{n}\). In an orthonormal basis adapted to \(\hat{n}\), \(R\) acts as the identity on \(\hat{n}\) and as an element of \(\SO(2,\R)\) on the orthogonal plane, which is Equation (14.25) for some angle \(\phi\); replacing \(\hat{n}\) by \(-\hat{n}\) if necessary brings \(\phi\) into \([0,\pi]\).

Proposition 14.31 ($\SO(3,\R)$ is not simply connected).

The map sending \(\vect{\phi} = \phi\hat{n}\) with \(\phi \in [0,\pi]\) to \(R(\hat{n},\phi)\) identifies \(\SO(3,\R)\) with the closed ball of radius \(\pi\) in \(\R^{3}\) with antipodal boundary points identified, that is with the real projective space \(\R P^{3}\). It is compact, connected, and not simply connected in the sense of Definition 6.17: the loop traced by \(R(\hat{n},\phi)\) as \(\phi\) runs from \(0\) to \(2\pi\) cannot be contracted to a point, while the same loop traversed twice can. Rests on Proposition 14.30, Definition 6.17 and Corollary 14.38.

Derivation. Derives Proposition 14.31. That the map is onto is Proposition 14.30, and it is injective on the open ball because \(R(\hat{n},\phi)\) determines its fixed axis and its rotation angle; on the boundary sphere \(\phi = \pi\) the only failure of injectivity is \(R(\hat{n},\pi) = R(-\hat{n},\pi)\), immediate from Equation (14.37) since \(\cos\pi = -1\) and \(\sin\pi = 0\) leave only terms even in \(\hat{n}\). A continuous bijection from a compact space to a Hausdorff space is a homeomorphism, so \(\SO(3,\R) \cong \R P^{3}\); connectedness and compactness follow. The statement about loops is Corollary 14.38 below, where it is read off the double cover.

The group $\SU(2)$

Definition 14.32 (The special unitary group in two dimensions).
\begin{equation}\tag{14.38} \SU(2) = \set{U \in \GL(2,\C) \mid U^{\dagger}U = \identity,\ \det U = 1}\ep \end{equation}
Proposition 14.33 ($\SU(2)$ is the three-sphere).

Every \(U \in \SU(2)\) has the form

\begin{equation}\tag{14.39} U = \begin{pmatrix} \alpha & \beta\\ -\bar{\beta} & \bar{\alpha}\end{pmatrix}\ec\qquad \abs{\alpha}^{2} + \abs{\beta}^{2} = 1\ec \end{equation}

and \((\alpha,\beta) \mapsto U\) is a homeomorphism from the unit sphere \(S^{3} \subset \R^{4} \cong \C^{2}\) onto \(\SU(2)\). Hence \(\SU(2)\) is compact, connected and simply connected. Rests on Definition 14.32.

Proof.

Derives Proposition 14.33. Write

\begin{equation*} U = \begin{pmatrix} a & b\\ c & d\end{pmatrix}\ep \end{equation*}

With \(\det U = 1\) the inverse is \(U^{-1} = \begin{pmatrix} d & -b\\ -c & a\end{pmatrix}\), while unitarity says \(U^{-1} = U^{\dagger} = \begin{pmatrix} \bar{a} & \bar{c}\\ \bar{b} & \bar{d}\end{pmatrix}\). Comparing entries, \(d = \bar{a}\) and \(c = -\bar{b}\), which is Equation (14.39) with \(\alpha = a\), \(\beta = b\); the determinant condition then reads \(\abs{\alpha}^{2} + \abs{\beta}^{2} = 1\). Conversely every such matrix is unitary with unit determinant. The correspondence and its inverse are given by polynomials in the real and imaginary parts, so both are continuous. Finally \(S^{3}\) is connected and simply connected by Lemma 14.34.

The last step is the only one that is not immediate, and since the distinction between \(\SU(2)\) and \(\SO(3,\R)\) turns entirely on it — and with that distinction, the existence of half-integer spin — it is proved rather than quoted. Its natural home is the topology chapter, whose Definition 6.17 it uses.

Lemma 14.34 (The $n$-sphere is simply connected for $n \ge 2$).

For \(n \ge 2\) the unit sphere \(S^{n} = \set{y \in \R^{n+1} \mid \norm{y} = 1}\) is path-connected and simply connected. Rests on Definition 6.17, Example 6.18 and Theorem 7.25.

Proof.

Derives Lemma 14.34. For \(p, q \in S^{n}\) with \(q \neq -p\) define the normalized chord

\begin{equation}\tag{14.40} \sigma_{p,q}(u) = \frac{(1-u)p + uq}{\norm{(1-u)p + uq}}\ec \qquad 0 \le u \le 1\ep \end{equation}

It is well defined, because \(\norm{(1-u)p+uq}^{2} = (1-u)^{2} + u^{2} + 2u(1-u)\,p\cdot q\) can vanish only if \(u = 1/2\) and \(p\cdot q = -1\), that is only if \(q = -p\); it is jointly continuous in \(p\), \(q\) and \(u\); it runs from \(p\) to \(q\); and its image lies in the plane spanned by \(p\) and \(q\). Path-connectedness follows at once: two points are joined by Equation (14.40) unless they are antipodal, and an antipodal pair is joined through any third point.

Every loop is homotopic to a piecewise-circular one. Let \(\gamma : [0,1] \longrightarrow S^{n}\) be continuous with \(\gamma(0) = \gamma(1)\). Each of the \(n+1\) components of \(\gamma\) is uniformly continuous by Theorem 7.25, whose proof applies verbatim to a map into \(\R^{n+1}\), so there is an \(N\) with \(\norm{\gamma(s) - \gamma(s')} < 1\) whenever \(\abs{s-s'} \le 1/N\). Put \(p_{k} = \gamma(k/N)\). Consecutive \(p_{k}\) are not antipodal, their distance being less than \(1 < 2\), so

\begin{equation*} \tilde{\gamma}(s) = \sigma_{p_{k},p_{k+1}}\left(Ns-k\right)\ec \qquad \frac{k}{N} \le s \le \frac{k+1}{N}\ec \end{equation*}

is a well-defined continuous loop with \(\tilde{\gamma}(0) = p_{0} = \gamma(0)\) and \(\tilde{\gamma}(1) = p_{N} = \gamma(1)\). For \(s\) in the \(k\)-th subinterval both \(\norm{\gamma(s)-p_{k}} < 1\) and \(\norm{\gamma(s)-p_{k+1}} < 1\), and for unit vectors \(\norm{a-b}^{2} = 2 - 2\,a\cdot b\), so \(\gamma(s)\cdot p_{k} > 1/2\) and \(\gamma(s)\cdot p_{k+1} > 1/2\). Since \(\tilde{\gamma}(s)\) is a combination of \(p_{k}\) and \(p_{k+1}\) with nonnegative coefficients divided by a norm at most \(1\),

\begin{equation*} \gamma(s)\cdot\tilde{\gamma}(s) = \frac{(1-u)\,\gamma(s)\cdot p_{k} + u\,\gamma(s)\cdot p_{k+1}} {\norm{(1-u)p_{k}+up_{k+1}}} \ge \tfrac{1}{2}\ec\qquad u = Ns-k\ep \end{equation*}

Hence \((1-t)\gamma(s) + t\tilde{\gamma}(s)\) never vanishes, its squared norm being at least \((1-t)^{2}+t^{2}\), and

\begin{equation*} H(s,t) = \frac{(1-t)\gamma(s) + t\,\tilde{\gamma}(s)} {\norm{(1-t)\gamma(s) + t\,\tilde{\gamma}(s)}} \end{equation*}

is a homotopy from \(\gamma\) to \(\tilde{\gamma}\) fixing the base point.

A piecewise-circular loop misses a point. The image of \(\tilde{\gamma}\) lies in the union of the \(N\) sets \(C_{k} = S^{n}\cap\Pi_{k}\), where \(\Pi_{k}\) is a plane of dimension at most two containing \(p_{k}\) and \(p_{k+1}\); enlarging it if necessary, take \(\dim\Pi_{k} = 2\). Because \(n+1 \ge 3\), the planes \(\Pi' = \operatorname{span}\set{e_{1},\cos\alpha\,e_{2} + \sin\alpha\,e_{3}}\) are pairwise distinct for \(\alpha \in [0,\pi)\) and infinite in number, so some \(\Pi'\) differs from all \(N\) of the \(\Pi_{k}\). Two distinct two-dimensional planes meet in a subspace of dimension at most one, so \(S^{n}\cap\Pi'\cap\Pi_{k}\) has at most two points for each \(k\); the circle \(S^{n}\cap\Pi'\) is infinite, and therefore contains a point \(x\) lying on no \(C_{k}\).

The complement of a point is \(\R^{n}\). Stereographic projection from \(x\),

\begin{equation}\tag{14.41} \pi(y) = \frac{y-(y\cdot x)\,x}{1-y\cdot x}\ec \end{equation}

maps \(S^{n}\setminus\set{x}\) into the \(n\)-dimensional space \(x^{\perp} \cong \R^{n}\) and is continuous there. So is \(w \mapsto \left(2w + (\norm{w}^{2}-1)x\right)/(\norm{w}^{2}+1)\), which lands on \(S^{n}\) — its squared norm is \(\left(4\norm{w}^{2}+(\norm{w}^{2}-1)^{2}\right)/(\norm{w}^{2}+1)^{2} = 1\) — and a direct substitution shows the two maps are mutually inverse. Hence \(S^{n}\setminus\set{x}\) is homeomorphic to \(\R^{n}\), which is simply connected by Example 6.18. The loop \(\tilde{\gamma}\) therefore contracts to its base point inside \(S^{n}\setminus\set{x}\), hence inside \(S^{n}\); composing that homotopy with \(H\) contracts \(\gamma\).

Proposition 14.35 (Generators of $\mathfrak{su}(2)$).

The Lie algebra \(\mathfrak{su}(2)\) is the space of traceless anti-Hermitian \(2\times2\) matrices, of real dimension \(3\). Writing its elements as \(-\ii\theta^{k}T_{k}\) with \(\theta^{k}\) real, a basis of Hermitian generators is

\begin{equation}\tag{14.42} \boxed{T_{k} = \tfrac{1}{2}\sigma_{k}}\ec\qquad \sigma_{1} = \begin{pmatrix} 0 & 1\\ 1 & 0\end{pmatrix}\ec\quad \sigma_{2} = \begin{pmatrix} 0 & -\ii\\ \ii & 0\end{pmatrix}\ec\quad \sigma_{3} = \begin{pmatrix} 1 & 0\\ 0 & -1\end{pmatrix}\ec \end{equation}

the Pauli matrices [Pauli:1927b]. They obey

\begin{equation}\tag{14.43} \sigma_{i}\sigma_{j} = \delta_{ij}\identity + \ii\,\epsilon_{ijk}\sigma_{k}\ec \end{equation}

and therefore

\begin{equation}\tag{14.44} \boxed{\comm{T_{i}}{T_{j}} = \ii\,\epsilon_{ijk}T_{k}}\ec\qquad \tr\left(T_{i}T_{j}\right) = \tfrac{1}{2}\delta_{ij}\ep \end{equation}

Rests on Definition 14.32.

Proof.

Derives Proposition 14.35. Setting \(U = \identity + \varepsilon X + O(\varepsilon^{2})\) in Equation (14.38) gives \(X^{\dagger} + X = 0\) and \(\tr X = 0\). A traceless anti-Hermitian \(2\times2\) matrix is \(\begin{pmatrix} \ii u & v + \ii w\\ -v + \ii w & -\ii u\end{pmatrix}\) with \(u,v,w\) real, a three-dimensional real space, and the three matrices \(-\ii\sigma_{k}/2\) span it. Equation (14.43) is checked directly: \(\sigma_{i}^{2} = \identity\) for each \(i\) by inspection, and the three products \(\sigma_{1}\sigma_{2} = \ii\sigma_{3}\), \(\sigma_{2}\sigma_{3} = \ii\sigma_{1}\), \(\sigma_{3}\sigma_{1} = \ii\sigma_{2}\) are likewise immediate, the remaining cases following from \(\sigma_{j}\sigma_{i} = 2\delta_{ij}\identity - \sigma_{i}\sigma_{j}\). Antisymmetrizing Equation (14.43) gives \(\comm{\sigma_{i}}{\sigma_{j}} = 2\ii\epsilon_{ijk}\sigma_{k}\), that is Equation (14.44); taking the trace of Equation (14.43) and using \(\tr\sigma_{k} = 0\) gives \(\tr(\sigma_{i}\sigma_{j}) = 2\delta_{ij}\).

Equation (14.44) is the same algebra as Equation (14.35). The two groups therefore have isomorphic Lie algebras, \(\mathfrak{su}(2) \cong \mathfrak{so}(3)\), and cannot be distinguished by any local construction. They are nevertheless different groups, and the next result says exactly how.

Proposition 14.36 (Closed form of the exponential).

For a unit vector \(\hat{n}\),

\begin{equation}\tag{14.45} U(\hat{n},\psi) = \exp\left(-\tfrac{\ii}{2}\psi\,\hat{n}\cdot\vect{\sigma}\right) = \identity\cos\frac{\psi}{2} - \ii\left(\hat{n}\cdot\vect{\sigma}\right)\sin\frac{\psi}{2}\ec \end{equation}

and in particular \(U(\hat{n},2\pi) = -\identity\) and \(U(\hat{n},4\pi) = \identity\). Rests on Proposition 14.35 and Equation (14.28).

Proof.

Derives Proposition 14.36. By Equation (14.43), \(\left(\hat{n}\cdot\vect{\sigma}\right)^{2} = n_{i}n_{j}\sigma_{i}\sigma_{j} = n_{i}n_{j}\delta_{ij}\identity = \identity\), the \(\epsilon\) term dropping because \(n_{i}n_{j}\) is symmetric. The even and odd powers of \(\hat{n}\cdot\vect{\sigma}\) therefore alternate exactly as \(J\) did in Equation (14.27), and the same resummation that produced Equation (14.28) produces Equation (14.45). Setting \(\psi = 2\pi\) gives \(\cos\pi = -1\), \(\sin\pi = 0\).

The double cover $\SU(2) \rightarrow \SO(3,\R)$

Theorem 14.37 ($\SU(2)$ is a two-to-one cover of $\SO(3,\R)$).

Identify \(\vect{x} \in \R^{3}\) with the traceless Hermitian matrix \(X = \vect{x}\cdot\vect{\sigma}\), and define

\begin{equation}\tag{14.46} \boxed{\Phi(U):\ \vect{x}\cdot\vect{\sigma} \longmapsto U\left(\vect{x}\cdot\vect{\sigma}\right)U^{\dagger}}\ep \end{equation}

Then \(\Phi\) is a surjective group homomorphism \(\SU(2) \rightarrow \SO(3,\R)\) with

\begin{equation}\tag{14.47} \ker\Phi = \set{\identity,-\identity} \cong \Z_{2}\ec \end{equation}

and explicitly \(\Phi\left(U(\hat{n},\phi)\right) = R(\hat{n},\phi)\). Rests on Propositions 14.30, 14.33, 14.35 and 14.36.

Proof.

Derives Theorem 14.37. We check the four claims in turn.

\(\Phi(U)\) is a well-defined real linear map on \(\R^{3}\). The matrices \(\vect{x}\cdot\vect{\sigma}\) are exactly the traceless Hermitian \(2\times2\) matrices, since \(\set{\sigma_{1},\sigma_{2},\sigma_{3}}\) is a real basis of that space. If \(X\) is traceless Hermitian then so is \(UXU^{\dagger}\): Hermiticity because \(\left(UXU^{\dagger}\right)^{\dagger} = UX^{\dagger}U^{\dagger} = UXU^{\dagger}\), and tracelessness by cyclicity together with \(U^{\dagger}U = \identity\). So \(\Phi(U)\) maps \(\R^{3}\) to \(\R^{3}\) linearly.

\(\Phi(U)\) preserves the Euclidean norm. From Equation (14.43), \(\det\left(\vect{x}\cdot\vect{\sigma}\right) = -\vect{x}\cdot\vect{x}\), as one sees by writing out \(\begin{pmatrix} x_{3} & x_{1}-\ii x_{2}\\ x_{1}+\ii x_{2} & -x_{3}\end{pmatrix}\). Since \(\det U = 1\), \(\det\left(UXU^{\dagger}\right) = \det X\), so the norm of \(\vect{x}\) is unchanged and \(\Phi(U) \in \Ogrp(3,\R)\).

The image lies in \(\SO(3,\R)\), and \(\Phi\) is a homomorphism. \(\Phi(U_{1}U_{2}) = \Phi(U_{1})\Phi(U_{2})\) is immediate from Equation (14.46). The map \(U \mapsto \det\Phi(U) \in \set{+1,-1}\) is continuous on the connected space \(\SU(2)\) (Proposition 14.33) and equals \(+1\) at \(U = \identity\), hence equals \(+1\) everywhere.

The kernel. Suppose \(UXU^{\dagger} = X\) for every traceless Hermitian \(X\), i.e. \(U\) commutes with \(\sigma_{1},\sigma_{2},\sigma_{3}\). Writing \(U = \begin{pmatrix} a & b\\ c & d\end{pmatrix}\), commuting with \(\sigma_{3}\) forces \(b = c = 0\), and commuting with \(\sigma_{1}\) then forces \(a = d\). So \(U = a\identity\), and \(\det U = a^{2} = 1\) gives \(a = \pm1\). Both values do lie in the kernel, giving Equation (14.47).

Surjectivity, and the explicit form. It suffices to compute \(\Phi\left(U(\hat{n},\phi)\right)\) and find the rotation Equation (14.37), since those exhaust \(\SO(3,\R)\) by Proposition 14.30. Abbreviate \(\mathrm{c} = \cos(\phi/2)\), \(\mathrm{s} = \sin(\phi/2)\) and \(n = \hat{n}\cdot\vect{\sigma}\), \(X = \vect{x}\cdot\vect{\sigma}\). Using Equation (14.45),

\begin{align*} UXU^{\dagger} &= \left(\mathrm{c}\,\identity - \ii\mathrm{s}\,n\right)X \left(\mathrm{c}\,\identity + \ii\mathrm{s}\,n\right)\\ &= \mathrm{c}^{2}X + \ii\,\mathrm{c}\mathrm{s}\comm{X}{n} + \mathrm{s}^{2}\,nXn\ep \end{align*}

Now Equation (14.43) gives \(\comm{X}{n} = x_{i}n_{j}\comm{\sigma_{i}}{\sigma_{j}} = 2\ii\,\epsilon_{ijk}x_{i}n_{j}\sigma_{k} = -2\ii\left(\hat{n}\times\vect{x}\right)\cdot\vect{\sigma}\), and, using \(nX + Xn = 2\left(\hat{n}\cdot\vect{x}\right)\identity\) from the symmetric part of the same identity together with \(n^{2} = \identity\),

\begin{equation*} nXn = \left(2\left(\hat{n}\cdot\vect{x}\right)\identity - Xn\right)n = 2\left(\hat{n}\cdot\vect{x}\right)n - X\ep \end{equation*}

Collecting the three pieces and using \(\mathrm{c}^{2}-\mathrm{s}^{2} = \cos\phi\), \(2\mathrm{c}\mathrm{s} = \sin\phi\), \(2\mathrm{s}^{2} = 1 - \cos\phi\),

\begin{equation*} UXU^{\dagger} = \left[\vect{x}\cos\phi + \hat{n}\left(\hat{n}\cdot\vect{x}\right)\left(1-\cos\phi\right) + \left(\hat{n}\times\vect{x}\right)\sin\phi\right]\cdot\vect{\sigma}\ec \end{equation*}

which is exactly Equation (14.37).

Corollary 14.38 ($\SU(2)$ is the universal cover).

\(\SO(3,\R) \cong \SU(2)/\set{\identity,-\identity}\), and since \(\SU(2)\) is simply connected (Proposition 14.33) it is the universal covering group of \(\SO(3,\R)\), with

\begin{equation}\tag{14.48} \pi_{1}\left(\SO(3,\R)\right) \cong \Z_{2}\ep \end{equation}

Concretely: the path \(\phi \mapsto U(\hat{n},\phi)\), \(0 \le \phi \le 2\pi\), joins \(\identity\) to \(-\identity\) in \(\SU(2)\) and projects to a closed loop in \(\SO(3,\R)\) that is not contractible, while the path with \(0 \le \phi \le 4\pi\) is a closed loop upstairs and projects to a contractible one. Rests on Theorem 14.37 and Proposition 14.33.

Everything in the corollary except Equation (14.48) — that the quotient is \(\SO(3,\R)\), that the kernel has two elements, that the two paths are as described — is proved above from the material of this chapter. The remaining step, from a two-sheeted covering by a simply connected space to the fundamental group of the base, is covering-space theory: the homotopy lifting property, unique path lifting, and the resulting isomorphism between the fundamental group of the base and the deck-transformation group, here the kernel of the covering homomorphism.

Derives Corollary 14.38.

Remark 14.39 (This is why half-integer spin exists).

A quantum state is a ray, so a symmetry need only be represented up to a phase: what a connected symmetry group must furnish is a projective unitary representation, and by Bargmann's analysis these are the ordinary unitary representations of the universal covering group [Bargmann:1954]. For rotations that covering group is \(\SU(2)\), not \(\SO(3,\R)\). Theorem 14.43 below will show that an irreducible representation of \(\mathfrak{su}(2)\) carries a label \(j\) with \(2j \in \N\cup\set{0}\) and Proposition 14.44 that every such \(j\) occurs, and Proposition 14.45 will show that exactly the integer ones descend to \(\SO(3,\R)\) while the half-integer ones do not. The half-integer representations are therefore not an extra postulate: they are what the double cover makes available, and the electron uses them [Uhlenbeck:1926] [Pauli:1927b]. The same argument in the relativistic setting replaces \(\SU(2)\) by \(\SL(2,\C)\) and is carried out in Proposition 94.10.

The irreducible representations of $\mathfrak{su}(2)$

The construction that follows is the ladder-operator argument, and it is worth noting how little goes into it: only the brackets Equation (14.44), the existence of the Casimir guaranteed by Corollary 14.17, and finite dimensionality [Wigner:1931].

Definition 14.40 (Ladder operators and the Casimir).

In the complexified algebra put

\begin{equation}\tag{14.49} J_{\pm} = J_{1} \pm \ii J_{2}\ec\qquad \vect{J}^{2} = J_{1}^{2} + J_{2}^{2} + J_{3}^{2}\ep \end{equation}

Rests on Equation (14.44).

Lemma 14.41 (Ladder algebra).
\begin{equation}\tag{14.50} \comm{J_{3}}{J_{\pm}} = \pm J_{\pm}\ec\qquad \comm{J_{+}}{J_{-}} = 2J_{3}\ec \end{equation}

and

\begin{equation}\tag{14.51} \vect{J}^{2} = J_{-}J_{+} + J_{3}\left(J_{3}+1\right) = J_{+}J_{-} + J_{3}\left(J_{3}-1\right)\ep \end{equation}

Moreover \(\vect{J}^{2}\) is the quadratic Casimir of Corollary 14.17 up to normalisation: the Killing form of \(\mathfrak{su}(2)\) in the Hermitian basis is \(\kappa_{ab} = 2\delta_{ab}\), so \(C_{2} = \kappa^{ab}J_{a}J_{b} = \tfrac{1}{2}\vect{J}^{2}\). Rests on Definition 14.40, Equation (14.44), Equation (14.15) and Corollary 14.17.

Proof.

Derives Lemma 14.41. From Equation (14.44), \(\comm{J_{3}}{J_{1}\pm\ii J_{2}} = \ii J_{2} \pm \ii(-\ii J_{1}) = \pm\left(J_{1} \pm \ii J_{2}\right)\), and \(\comm{J_{+}}{J_{-}} = -2\ii\comm{J_{1}}{J_{2}} = 2J_{3}\). Next

\begin{equation*} J_{-}J_{+} = \left(J_{1}-\ii J_{2}\right)\left(J_{1}+\ii J_{2}\right) = J_{1}^{2} + J_{2}^{2} + \ii\comm{J_{1}}{J_{2}} = \vect{J}^{2} - J_{3}^{2} - J_{3}\ec \end{equation*}

which rearranges to the first form of Equation (14.51); the second follows by exchanging the roles of \(J_{+}\) and \(J_{-}\), which flips the sign of the \(\comm{J_{1}}{J_{2}}\) term. For the Killing form use Equation (14.15) with \(C^{c}{}_{ab} = \ii\epsilon_{abc}\) from Equation (14.44): \(\kappa_{ab} = \left(\ii\epsilon_{adc}\right)\left(\ii\epsilon_{bcd}\right) = -\epsilon_{adc}\epsilon_{bcd} = \epsilon_{adc}\epsilon_{bdc} = 2\delta_{ab}\).

Remark 14.42 (The compact real form).

The value \(\kappa_{ab} = +2\delta_{ab}\) is computed in the basis of Hermitian generators \(J_{a}\), which spans the complexification. On the real algebra \(\mathfrak{su}(2)\) itself, whose elements are the anti-Hermitian \(-\ii\theta^{a}J_{a} = \theta^{a}L_{a}\), the same computation gives \(\kappa_{ab} = -2\delta_{ab}\): the Killing form of a compact real form is negative definite. Nondegeneracy, which is all that Corollary 14.17 requires, holds either way.

Theorem 14.43 (The irreducible representations of $\mathfrak{su}(2)$).

Let \(D\) be a finite-dimensional irreducible representation of \(\mathfrak{su}(2)\) on a complex inner-product space, unitary in the sense that the \(J_{a}\) are represented by Hermitian operators. Then there is a number \(j\) with

\begin{equation}\tag{14.52} \boxed{2j \in \N \cup \set{0}} \end{equation}

such that the representation has dimension \(2j+1\), an orthonormal basis \(\set{\ket{j,m}}\) with \(m = -j,-j+1,\ldots,j-1,j\), and

\begin{equation}\tag{14.53} \boxed{J_{3}\ket{j,m} = m\ket{j,m}\ec\qquad \vect{J}^{2}\ket{j,m} = j(j+1)\ket{j,m}}\ec \end{equation}
\begin{equation}\tag{14.54} J_{\pm}\ket{j,m} = \sqrt{j(j+1) - m(m\pm1)}\;\ket{j,m\pm1}\ep \end{equation}

Rests on Lemma 14.41 and Equation (14.44).

Proof.

Derives Theorem 14.43. \(J_{3}\) is Hermitian on a finite-dimensional space, so it has an eigenvector; let \(\mu\) be an eigenvalue of largest real part, with eigenvector \(\ket{\mu}\). From Equation (14.50), \(J_{3}\left(J_{\pm}\ket{\mu}\right) = \left(\mu\pm1\right)J_{\pm}\ket{\mu}\), so \(J_{\pm}\) raises and lowers the \(J_{3}\) eigenvalue by one. Maximality of \(\mu\) forces \(J_{+}\ket{\mu} = 0\). Write \(j := \mu\) and \(\ket{j,j} := \ket{\mu}\) normalised. Then by Equation (14.51),

\begin{equation*} \vect{J}^{2}\ket{j,j} = \left(J_{-}J_{+} + J_{3}(J_{3}+1)\right)\ket{j,j} = j(j+1)\ket{j,j}\ec \end{equation*}

and since \(\vect{J}^{2}\) commutes with \(J_{3}\) and with \(J_{\pm}\) its eigenspaces are invariant under the whole algebra, so by irreducibility the value \(j(j+1)\) is taken on the entire representation. (This is Corollary 14.18 again, argued here directly so that nothing about the group is needed.)

Define \(\ket{j,m}\) recursively by lowering. Because \(J_{+}\) and \(J_{-}\) are adjoint to each other when the \(J_{a}\) are Hermitian,

\begin{equation*} \norm{J_{-}\ket{j,m}}^{2} = \bra{j,m}J_{+}J_{-}\ket{j,m} = j(j+1) - m(m-1)\ec \end{equation*}

using the second form of Equation (14.51); the same computation with \(J_{+}\) and the first form gives \(\norm{J_{+}\ket{j,m}}^{2} = j(j+1)-m(m+1)\), which is Equation (14.54) once phases are fixed by choosing the square roots positive. Since the space is finite-dimensional the chain of eigenvalues \(j, j-1, j-2, \ldots\) must terminate, which by the norm formula happens exactly at the value \(m_{\min}\) with \(j(j+1) - m_{\min}(m_{\min}-1) = 0\), that is \(m_{\min} \in \set{-j,\, j+1}\); the second is excluded because \(m_{\min} \le j\) and would need \(j+1 \le j\). Hence \(m_{\min} = -j\), and the chain runs from \(j\) down to \(-j\) in unit steps, so \(j - (-j) = 2j\) is a non-negative integer. This is the whole reason half-integer angular momentum is possible: nothing in the algebra forbids the chain from having an even number of steps. The span of the \(\ket{j,m}\) is invariant under \(J_{3}\) and \(J_{\pm}\), hence under the whole algebra, and is nonzero, so by irreducibility it is the whole space, of dimension \(2j+1\).

Theorem 14.43 is a statement of necessity: an irreducible representation, if there is one, looks like that. It does not produce a single representation, and everything below — the descent to \(\SO(3,\R)\), the Clebsch–Gordan series, the spin label of a Poincaré representation — needs to know that the list is not empty. The converse is the shorter half, because the theorem has already written down the matrix elements a representation would have to have; it remains to check that they define one.

Proposition 14.44 (Existence: every allowed $j$ is realized).

For each \(j\) with \(2j \in \N\cup\set{0}\) there exists a \((2j+1)\)-dimensional irreducible representation of \(\mathfrak{su}(2)\) by Hermitian operators, and its matrix elements are Equations (14.53) and (14.54). Together with Theorem 14.43 this puts the equivalence classes of finite-dimensional irreducible representations of \(\mathfrak{su}(2)\) in bijection with the numbers \(j = 0,\tfrac{1}{2},1,\tfrac{3}{2},\ldots\) Rests on Theorem 14.43 and Lemma 14.41.

Proof.

Derives Proposition 14.44. Let \(\mathbb{V}_{j}\) be the complex inner-product space with orthonormal basis \(\set{\ket{j,m}}\), \(m = -j,-j+1,\ldots,j\), of dimension \(2j+1\), and define operators on it by

\begin{equation}\tag{14.55} J_{3}\ket{j,m} = m\ket{j,m}\ec\qquad J_{\pm}\ket{j,m} = c_{\pm}(m)\ket{j,m\pm1}\ec \end{equation}
\begin{equation}\tag{14.56} c_{\pm}(m) = \sqrt{j(j+1)-m(m\pm1)}\ec \end{equation}

with the convention that \(\ket{j,m}\) denotes the zero vector for \(\abs{m} > j\). The convention is consistent because \(j(j+1)-m(m+1) = (j-m)(j+m+1)\) and \(j(j+1)-m(m-1) = (j+m)(j-m+1)\), which are non-negative for \(-j \le m \le j\) and vanish exactly at \(c_{+}(j) = 0\) and \(c_{-}(-j) = 0\): the operators do not attempt to leave the basis. Put \(J_{1} = \tfrac{1}{2}(J_{+}+J_{-})\) and \(J_{2} = \tfrac{1}{2\ii}(J_{+}-J_{-})\).

The brackets hold. \(J_{3}J_{\pm}\ket{j,m} = (m\pm1)c_{\pm}(m)\ket{j,m\pm1}\) while \(J_{\pm}J_{3}\ket{j,m} = m\,c_{\pm}(m)\ket{j,m\pm1}\), so \(\comm{J_{3}}{J_{\pm}} = \pm J_{\pm}\). Since \(c_{+}(m-1) = c_{-}(m)\) by Equation (14.56), \(J_{+}J_{-}\ket{j,m} = c_{-}(m)^{2}\ket{j,m}\) and \(J_{-}J_{+}\ket{j,m} = c_{+}(m)^{2}\ket{j,m}\), whence

\begin{equation*} \comm{J_{+}}{J_{-}}\ket{j,m} = \left(\left[j(j+1)-m(m-1)\right] - \left[j(j+1)-m(m+1)\right]\right)\ket{j,m} = 2m\ket{j,m}\ec \end{equation*}

which is \(2J_{3}\ket{j,m}\). That is Equation (14.50), equivalent to Equation (14.44) by the computation in Lemma 14.41 run backwards.

The generators are Hermitian. The basis is orthonormal and the \(c_{\pm}\) are real, so the only nonzero matrix elements of the ladder operators are \(\bra{j,m+1}J_{+}\ket{j,m} = c_{+}(m)\) and \(\bra{j,m}J_{-}\ket{j,m+1} = c_{-}(m+1) = c_{+}(m)\). Hence \(J_{-} = J_{+}^{\dagger}\), and \(J_{1}\), \(J_{2}\), \(J_{3}\) are Hermitian.

It is irreducible. Let \(\mathbb{W} \neq 0\) be an invariant subspace. It is invariant under the Hermitian \(J_{3}\), whose eigenvalues on \(\mathbb{V}_{j}\) are the \(2j+1\) distinct numbers \(m\), so \(\mathbb{W}\) is spanned by those \(\ket{j,m}\) it contains and contains at least one of them. From any one of them the ladder operators reach every other, their coefficients \(c_{\pm}(m)\) vanishing only at the two ends of the chain. So \(\mathbb{W} = \mathbb{V}_{j}\).

Finally Equation (14.51) gives \(\vect{J}^{2}\ket{j,m} = \left(c_{+}(m)^{2}+m(m+1)\right)\ket{j,m} = j(j+1)\ket{j,m}\), so Equations (14.55) and (14.56) are Equations (14.53) and (14.54).

Proposition 14.45 (Which representations descend to $\SO(3,\R)$).

Let \(D^{(j)}\) be the representation of \(\SU(2)\) obtained by exponentiating the representation of \(\mathfrak{su}(2)\) constructed in Proposition 14.44. Then

\begin{equation}\tag{14.57} D^{(j)}(-\identity) = (-1)^{2j}\,\identity\ec \end{equation}

so \(D^{(j)}\) descends to a genuine representation of \(\SO(3,\R) = \SU(2)/\set{\pm\identity}\) if and only if \(j\) is an integer. For half-integer \(j\) the map \(R \mapsto D^{(j)}(U)\) is defined only up to the sign ambiguity in the choice of \(U\) above \(R\): a projective representation, in the sense of Remark 14.39. Rests on Proposition 14.44, Proposition 14.36 and Corollary 14.38.

Proof.

Derives Proposition 14.45. By Proposition 14.36, \(-\identity = U(\hat{e}_{3},2\pi) = \exp\left(-2\pi\ii\,T_{3}\right)\), so \(D^{(j)}(-\identity) = \exp\left(-2\pi\ii J_{3}\right)\) in the representation, and on \(\ket{j,m}\) this is \(\ee^{-2\pi\ii m}\) by Equation (14.53). The values of \(m\) are \(j, j-1,\ldots,-j\): all integers when \(j\) is an integer, all in \(\Z+\tfrac{1}{2}\) when \(j\) is half-odd-integer. In the first case \(\ee^{-2\pi\ii m} = 1\) for every \(m\), in the second \(\ee^{-2\pi\ii m} = -1\) for every \(m\); either way the operator is \((-1)^{2j}\identity\). A representation of \(\SU(2)\) factors through the quotient by \(\set{\pm\identity}\) precisely when it is trivial on that subgroup, which by Equation (14.57) happens exactly for integer \(j\).

Remark 14.46 (The integer case is what one already knew).

For \(j=1\) the representation of Theorem 14.43 is three-dimensional and is equivalent to the defining representation of \(\SO(3,\R)\) on \(\R^{3}\) complexified — both have \(\vect{J}^{2} = 2\), as one checks directly from Equation (14.35), and both are irreducible of dimension three. For general integer \(j\) the representation is the one carried by the spherical harmonics of degree \(j\), of dimension \(2j+1\). In SI terms (Remark 14.29) the observable statement is

\begin{equation}\tag{14.58} \boxed{\hat{\vect{\mathcal{J}}}^{2} = \hbar^{2}\,j(j+1)}\ec\qquad \hat{\mathcal{J}}_{3} = \hbar m\ec \end{equation}

with \(\hbar^{2}j(j+1)\) carrying the SI unit \(\mathrm{J}^{2}\,\mathrm{s}^{2}\) and \(\hbar m\) the unit \(\mathrm{J}\,\mathrm{s}\).

The Clebsch–Gordan decomposition

Theorem 14.47 (Clebsch–Gordan series).

Let \(V_{j_{1}}\) and \(V_{j_{2}}\) be the irreducible representations of \(\mathfrak{su}(2)\) of Theorem 14.43. Under the total generators \(\vect{J} = \vect{J}^{(1)}\otimes\identity + \identity\otimes\vect{J}^{(2)}\),

\begin{equation}\tag{14.59} \boxed{V_{j_{1}} \otimes V_{j_{2}} = \bigoplus_{J = \abs{j_{1}-j_{2}}}^{j_{1}+j_{2}} V_{J}}\ec \end{equation}

each \(V_{J}\) occurring exactly once, and the dimensions check:

\begin{equation}\tag{14.60} \left(2j_{1}+1\right)\left(2j_{2}+1\right) = \sum_{J=\abs{j_{1}-j_{2}}}^{j_{1}+j_{2}}\left(2J+1\right)\ep \end{equation}

Rests on Theorem 14.43 and Equation (14.44).

Proof.

Derives Theorem 14.47. That \(\vect{J}\) satisfies Equation (14.44) is immediate, the two factors commuting. Take \(j_{1} \ge j_{2}\) without loss of generality. In the product basis \(\ket{j_{1}m_{1}}\otimes\ket{j_{2}m_{2}}\) the operator \(J_{3}\) is diagonal with eigenvalue \(m = m_{1}+m_{2}\), so the multiplicity \(n(m)\) of the eigenvalue \(m\) is the number of pairs \((m_{1},m_{2})\) with \(m_{1}+m_{2} = m\), \(\abs{m_{i}} \le j_{i}\). Counting them,

\begin{equation}\tag{14.61} n(m) = \begin{cases} j_{1}+j_{2}-\abs{m}+1\ec & j_{1}-j_{2} \le \abs{m} \le j_{1}+j_{2}\ec\\ 2j_{2}+1\ec & \abs{m} \le j_{1}-j_{2}\ec\\ 0 \ec & \abs{m} > j_{1}+j_{2}\ep \end{cases} \end{equation}

Now the tensor product is a finite-dimensional representation, hence a direct sum of irreducibles \(V_{J}\) with multiplicities \(a_{J} \ge 0\); by Equation (14.53) each \(V_{J}\) contributes exactly \(1\) to \(n(m)\) for every \(\abs{m}\le J\) and nothing otherwise, so \(n(m) = \sum_{J \ge \abs{m}} a_{J}\). Subtracting consecutive values,

\begin{equation*} a_{J} = n(J) - n(J+1)\ec \end{equation*}

and Equation (14.61) makes this \(1\) for \(\abs{j_{1}-j_{2}} \le J \le j_{1}+j_{2}\) and \(0\) otherwise, which is Equation (14.59). For Equation (14.60) write \(J = j_{1}-j_{2}+k\) with \(k = 0,\ldots,2j_{2}\):

\begin{align*} \sum_{k=0}^{2j_{2}}\left(2(j_{1}-j_{2}+k)+1\right) &= \left(2j_{2}+1\right)\left(2j_{1}-2j_{2}+1\right) + 2\,\frac{2j_{2}\left(2j_{2}+1\right)}{2}\\ &= \left(2j_{2}+1\right)\left(2j_{1}+1\right)\ep \end{align*}
Remark 14.48 (What is and is not proved here).

Theorem 14.47 settles which irreducible representations occur and with what multiplicity, and that is all the group theory of the following parts requires of it. The unitary matrix relating the product basis to the coupled basis — the Clebsch–Gordan coefficients themselves, together with the \(3j\), \(6j\) and \(9j\) symbols built from them — is a separate construction and belongs where it is used. This treatise reserves it for Sections 79.4.2 and 79.4.4 of Angular Momentum and Spin, which at the time of writing carry their headings and their editorial specification but not yet the construction; nothing in the present chapter depends on it. The sources are Wigner's book for the coefficients themselves, of which the \(3j\) symbol is his symmetrized form [Wigner:1931]; Racah's second paper for the closed form of the general coefficient and for the recoupling coefficient now written \(6j\) [Racah:1942]; and his third for the recoupling of four angular momenta, the modern \(9j\) symbol [Racah:1943]. The character-theoretic route to the same decomposition, using the product rule Equation (5.195), is available from Linear Algebra and Representation Theory; the weight count used above was preferred because it needs nothing beyond Equation (14.53).

The group $\SU(3)$

Definition 14.49 (The special unitary group in three dimensions).
\begin{equation}\tag{14.62} \SU(3) = \set{U \in \GL(3,\C) \mid U^{\dagger}U = \identity,\ \det U = 1}\ec \end{equation}

whose Lie algebra \(\mathfrak{su}(3)\) is the space of traceless anti-Hermitian \(3\times3\) matrices.

Proposition 14.50 (Dimension and rank).

\(\dim\mathfrak{su}(3) = 8\) and the rank of \(\mathfrak{su}(3)\) in the sense of Definition 14.20 is \(\ell = 2\). Rests on Definitions 14.20 and 14.49.

Proof.

Derives Proposition 14.50. Linearizing Equation (14.62) as in Proposition 14.35 gives \(X^{\dagger} + X = 0\) and \(\tr X = 0\). The anti-Hermitian \(3\times3\) matrices form a real vector space of dimension \(9\) — three real diagonal entries times \(\ii\), and three complex entries above the diagonal — and the trace condition removes one, leaving \(8\). For the rank, a maximal abelian subalgebra consisting of diagonalizable elements may be taken to be the traceless diagonal anti-Hermitian matrices, of which there are two independent, and no larger commuting set of diagonalizable elements exists because a set of commuting Hermitian matrices is simultaneously diagonalizable and the diagonal traceless ones already exhaust that possibility.

Definition 14.51 (Gell-Mann basis).

The standard basis of Hermitian generators is \(T_{a} = \tfrac{1}{2}\lambda_{a}\), \(a = 1,\ldots,8\), with [GellMann:1961] [Navas:2024]

\begin{equation}\tag{14.63} \lambda_{1} = \begin{pmatrix} 0&1&0\\ 1&0&0\\ 0&0&0\end{pmatrix}\ec\quad \lambda_{2} = \begin{pmatrix} 0&-\ii&0\\ \ii&0&0\\ 0&0&0\end{pmatrix} \ec\quad \lambda_{3} = \begin{pmatrix} 1&0&0\\ 0&-1&0\\ 0&0&0\end{pmatrix}\ec \end{equation}
\begin{equation}\tag{14.64} \lambda_{4} = \begin{pmatrix} 0&0&1\\ 0&0&0\\ 1&0&0\end{pmatrix}\ec\quad \lambda_{5} = \begin{pmatrix} 0&0&-\ii\\ 0&0&0\\ \ii&0&0\end{pmatrix} \ec\quad \lambda_{6} = \begin{pmatrix} 0&0&0\\ 0&0&1\\ 0&1&0\end{pmatrix}\ec \end{equation}
\begin{equation}\tag{14.65} \lambda_{7} = \begin{pmatrix} 0&0&0\\ 0&0&-\ii\\ 0&\ii&0\end{pmatrix} \ec\quad \lambda_{8} = \frac{1}{\sqrt{3}} \begin{pmatrix} 1&0&0\\ 0&1&0\\ 0&0&-2\end{pmatrix}\ep \end{equation}

They are Hermitian and traceless and satisfy the normalisation

\begin{equation}\tag{14.66} \tr\left(\lambda_{a}\lambda_{b}\right) = 2\delta_{ab}\ec \qquad\text{i.e.}\qquad \tr\left(T_{a}T_{b}\right) = \tfrac{1}{2}\delta_{ab}\ec \end{equation}

the same normalisation as Equation (14.44). Rests on Definition 14.49.

Proposition 14.52 (Structure constants and the symmetric tensor).

There are real arrays \(f_{abc}\), totally antisymmetric, and \(d_{abc}\), totally symmetric, with

\begin{equation}\tag{14.67} \boxed{\comm{T_{a}}{T_{b}} = \ii f_{abc}T_{c}\ec\qquad \acomm{T_{a}}{T_{b}} = \tfrac{1}{3}\delta_{ab}\identity + d_{abc}T_{c}}\ec \end{equation}

recovered from the generators by

\begin{equation}\tag{14.68} f_{abc} = -2\ii\,\tr\left(\comm{T_{a}}{T_{b}}T_{c}\right)\ec\qquad d_{abc} = 2\,\tr\left(\acomm{T_{a}}{T_{b}}T_{c}\right)\ep \end{equation}

Their independent nonvanishing values are

\begin{equation}\tag{14.69} f_{123} = 1\ec\quad f_{147} = f_{246} = f_{257} = f_{345} = \tfrac{1}{2}\ec\quad f_{156} = f_{367} = -\tfrac{1}{2}\ec\quad f_{458} = f_{678} = \tfrac{\sqrt{3}}{2}\ec \end{equation}
\begin{equation}\tag{14.70} \begin{aligned} d_{118} &= d_{228} = d_{338} = \tfrac{1}{\sqrt{3}}\ec & d_{888} &= -\tfrac{1}{\sqrt{3}}\ec\\ d_{448} &= d_{558} = d_{668} = d_{778} = -\tfrac{1}{2\sqrt{3}}\ec & d_{146} &= d_{157} = d_{256} = d_{344} = d_{355} = \tfrac{1}{2}\ec\\ d_{247} &= d_{366} = d_{377} = -\tfrac{1}{2}\ep & & \end{aligned} \end{equation}

Rests on Definition 14.51.

Proof.

Derives Proposition 14.52. The product \(T_{a}T_{b}\) is a \(3\times3\) matrix, and \(\set{\identity,T_{1},\ldots,T_{8}}\) is a basis of the complex \(3\times3\) matrices, since the \(T_{c}\) span the traceless ones. Expanding \(T_{a}T_{b}\) in it and separating the antisymmetric and symmetric parts in \((a,b)\) gives Equation (14.67), the coefficient of \(\identity\) in the symmetric part being \(\tfrac{1}{3}\tr\acomm{T_{a}}{T_{b}} = \tfrac{1}{3}\delta_{ab}\) by Equation (14.66). Multiplying either relation by \(T_{c}\) and taking the trace with Equation (14.66), and using \(\tr T_{c} = 0\), gives Equation (14.68). Total symmetry of \(d\) and total antisymmetry of \(f\) then follow from Equation (14.68) and the cyclicity of the trace: \(f\) is antisymmetric in \((a,b)\) by construction and

\begin{equation*} \tr\left(\comm{T_{a}}{T_{b}}T_{c}\right) = \tr\left(T_{a}T_{b}T_{c}\right) - \tr\left(T_{b}T_{a}T_{c}\right) = \tr\left(T_{a}\comm{T_{b}}{T_{c}}\right) \end{equation*}

is invariant under cyclic permutation of \((a,b,c)\), which together with antisymmetry in the first pair gives total antisymmetry; the same argument with the anticommutator gives total symmetry of \(d\). The tables Equations (14.69) and (14.70) are then a finite computation from Equations (14.63), (14.64) and (14.65); two entries are done here as specimens. For \(f_{458}\): a direct multiplication gives \(\comm{\lambda_{4}}{\lambda_{5}} = \diag(2\ii,0,-2\ii)\), so

\begin{equation*} f_{458} = -\frac{\ii}{4}\tr\left(\comm{\lambda_{4}}{\lambda_{5}} \lambda_{8}\right) = -\frac{\ii}{4}\cdot\frac{1}{\sqrt{3}}\left(2\ii + 0 + 4\ii\right) = \frac{\sqrt{3}}{2}\ep \end{equation*}

For \(d_{247}\): \(\acomm{\lambda_{2}}{\lambda_{4}} = -\lambda_{7}\) by direct multiplication, so \(d_{247} = \tfrac{1}{4}\tr\left(-\lambda_{7}\lambda_{7}\right) = -\tfrac{1}{4}\cdot2 = -\tfrac{1}{2}\).

The remaining entries are three-matrix traces of the same kind: nine independent nonvanishing components of \(f\) and sixteen of \(d\), twenty-five in all, every other nonzero component being fixed by the total antisymmetry of \(f\) and the total symmetry of \(d\). They are computed and tabulated in full in the appendix.

Derives Proposition 14.52.

Proposition 14.53 (Killing form of $\mathfrak{su}(3)$).

In the Hermitian basis of Definition 14.51,

\begin{equation}\tag{14.71} \kappa_{ab} = f_{acd}f_{bcd} = 3\,\delta_{ab}\ec \end{equation}

which is nondegenerate — as Theorem 14.13 demands of \(\mathfrak{su}(3)\), which is simple and hence semisimple — so that \(\kappa^{ab} = \tfrac{1}{3}\delta^{ab}\). Rests on Proposition 14.52, Definition 14.51, Equation (14.15) and Theorem 14.13.

Proof.

Derives Proposition 14.53. With \(C^{c}{}_{ab} = \ii f_{abc}\) from Equation (14.67), Equation (14.15) gives \(\kappa_{ab} = \left(\ii f_{adc}\right)\left(\ii f_{bcd}\right) = -f_{adc}f_{bcd} = f_{adc}f_{bdc} = f_{acd}f_{bcd}\) after renaming.

The array is evaluated here directly. The shorter route — both \(\kappa_{ab}\) and \(\tr(T_{a}T_{b})\) are symmetric bilinear forms invariant under the adjoint action, and on a simple algebra such a form is unique up to a scale, so it is enough to evaluate one entry — would have to obtain that uniqueness from Schur's lemma applied to the adjoint representation, and Theorem 5.153 is stated for a representation on a complex space, whereas the adjoint representation of \(\mathfrak{su}(3)\) acts on the eight-dimensional real algebra. Over \(\R\) the conclusion of that lemma is false in general, so its hypotheses would first have to be recovered by complexifying and by quoting the irreducibility of the complexified adjoint — which is no cheaper than the computation that follows and rests on the same quoted simplicity.

Because \(\comm{T_{a}}{T_{c}} = \ii f_{acd}T_{d}\) and \(\tr(T_{d}T_{e}) = \tfrac{1}{2}\delta_{de}\) by Equation (14.66),

\begin{equation}\tag{14.72} \sum_{c}\tr\left(\comm{T_{a}}{T_{c}}\comm{T_{b}}{T_{c}}\right) = -f_{acd}f_{bce}\,\tr\left(T_{d}T_{e}\right) = -\tfrac{1}{2}f_{acd}f_{bcd} = -\tfrac{1}{2}\kappa_{ab}\ep \end{equation}

Expanding the two commutators gives four traces. Summed over \(c\) they collapse to two, by cyclicity of the trace together with \(\sum_{c}T_{c}T_{c} \propto \identity\), which Equation (14.74) below establishes:

\begin{equation}\tag{14.73} \sum_{c}\tr\left(\comm{T_{a}}{T_{c}}\comm{T_{b}}{T_{c}}\right) = 2\sum_{c}\tr\left(T_{a}T_{c}T_{b}T_{c}\right) - 2\sum_{c}\tr\left(T_{a}T_{b}T_{c}T_{c}\right)\ep \end{equation}

Both sums follow from the completeness of the basis \(\set{\identity,T_{1},\ldots,T_{8}}\) of the complex \(3\times3\) matrices already used in Proposition 14.52. Expanding an arbitrary \(M\) in that basis, with the coefficients read off by Equation (14.66) and \(\tr T_{a} = 0\), gives \(M = \tfrac{1}{3}\left(\tr M\right)\identity + 2\,\tr\left(MT_{a}\right)T_{a}\), and taking for \(M\) the matrix whose only nonzero entry is a \(1\) in position \((j,k)\) turns that into

\begin{equation}\tag{14.74} \boxed{\sum_{a=1}^{8}\left(T_{a}\right)_{il}\left(T_{a}\right)_{kj} = \tfrac{1}{2}\left(\delta_{ij}\delta_{lk} - \tfrac{1}{3}\delta_{il}\delta_{jk}\right)}\ep \end{equation}

Setting \(l = k\) and summing gives \(\sum_{c}T_{c}T_{c} = \tfrac{1}{2}\left(3-\tfrac{1}{3}\right)\identity = \tfrac{4}{3}\identity\), so that \(\sum_{c}\tr(T_{a}T_{b}T_{c}T_{c}) = \tfrac{4}{3}\tr(T_{a}T_{b}) = \tfrac{2}{3}\delta_{ab}\); and contracting Equation (14.74) into \(\tr(T_{a}T_{c}T_{b}T_{c}) = \left(T_{a}\right)_{ij}\left(T_{c}\right)_{jk} \left(T_{b}\right)_{kl}\left(T_{c}\right)_{li}\) gives

\begin{equation*} \sum_{c}\tr\left(T_{a}T_{c}T_{b}T_{c}\right) = \tfrac{1}{2}\left(\tr T_{a}\,\tr T_{b} - \tfrac{1}{3}\tr\left(T_{a}T_{b}\right)\right) = -\tfrac{1}{12}\delta_{ab}\ep \end{equation*}

Substituting both into Equation (14.73),

\begin{equation*} \sum_{c}\tr\left(\comm{T_{a}}{T_{c}}\comm{T_{b}}{T_{c}}\right) = -\tfrac{1}{6}\delta_{ab} - \tfrac{4}{3}\delta_{ab} = -\tfrac{3}{2}\delta_{ab}\ec \end{equation*}

and Equation (14.72) gives \(\kappa_{ab} = 3\delta_{ab}\).

As a check on one entry, take \(a=b=3\) and read the nonzero \(f_{3cd}\) from Equation (14.69) — \(f_{312} = 1\), \(f_{345} = \tfrac{1}{2}\), \(f_{367} = -\tfrac{1}{2}\), each contributing twice because both index orders are summed:

\begin{equation*} \kappa_{33} = f_{3cd}f_{3cd} = 2\left(1\right)^{2} + 2\left(\tfrac{1}{2}\right)^{2} + 2\left(\tfrac{1}{2}\right)^{2} = 3\ep \end{equation*}

The simplicity of \(\mathfrak{su}(3)\) is used below, where the adjoint representation is asserted to be irreducible — that is, where the traceless part of the endomorphisms of the fundamental is identified with the eight-dimensional representation. The argument decomposes a putative ideal into eigenspaces of the two-dimensional Cartan subalgebra and then uses the fact that the six roots of the algebra — the hexagon exhibited below — are connected, every root being reachable from every other by adding roots, so that a single root vector in the ideal drags the whole algebra in with it.

Derives Theorem 14.57.

Proposition 14.54 (The two Casimir operators of $\mathfrak{su}(3)$).
\begin{equation}\tag{14.75} \boxed{C_{2} = \sum_{a=1}^{8} T_{a}T_{a}\ec\qquad C_{3} = d_{abc}\,T_{a}T_{b}T_{c}} \end{equation}

are Casimir elements of \(U\left(\mathfrak{su}(3)\right)\), and by Theorem 14.21 together with Proposition 14.50 they generate all of them: \(\ell = 2\), so there are exactly two. Rests on Corollary 14.17, Theorem 14.16, Proposition 14.52, Proposition 14.53, Theorem 14.110, Theorem 14.21 and Proposition 14.50.

Proof.

Derives Proposition 14.54. For \(C_{2}\) this is Corollary 14.17 with \(\kappa^{ab} = \tfrac{1}{3}\delta^{ab}\) from Equation (14.71), up to the harmless factor of \(3\). For \(C_{3}\), apply Theorem 14.16 with \(r = 3\) and \(k^{abc} = d_{abc}\): the tensor is totally symmetric by Proposition 14.52, and its invariance in the sense of Equation (14.18) is Equation (14.148) for the \(G\)-invariant polynomial \(\gen{T_{a},T_{b},T_{c}} = \tfrac{1}{2}d_{abc}\) — symmetric and invariant because \(\tr\left(\acomm{T_{a}}{T_{b}}T_{c}\right)\) is a trace of a product of generators and therefore unchanged under \(T_{a} \mapsto gT_{a}g^{-1}\), which is Definition 14.109. Written out, the invariance reads

\begin{equation}\tag{14.76} f_{ead}\,d_{dbc} + f_{ebd}\,d_{adc} + f_{ecd}\,d_{abd} = 0\ec \end{equation}

and it is exactly this identity that the proof of Theorem 14.16 consumes.

Proposition 14.55 (The fundamental, the antifundamental and the adjoint).

\(\SU(3)\) has three representations of immediate use.

  1. The fundamental \(\vect{3}\): the defining action on \(\C^{3}\), generated by the \(T_{a}\) themselves. Its quadratic Casimir is

    \begin{equation}\tag{14.77} C_{2}\left(\vect{3}\right) = \tfrac{4}{3}\ep \end{equation}
  2. The antifundamental \(\bar{\vect{3}}\): the action on the complex-conjugate space, generated by \(-T_{a}^{\ast}\). It has the same \(C_{2}\) and is not equivalent to \(\vect{3}\).

  3. The adjoint \(\vect{8}\): the action on \(\mathfrak{su}(3)\) itself, with generators \(\left(T_{a}^{\mathrm{ad}}\right)_{bc} = -\ii f_{abc}\), of dimension \(8\) and quadratic Casimir

    \begin{equation}\tag{14.78} C_{2}\left(\vect{8}\right) = 3\ep \end{equation}

Rests on Definition 14.51, Corollary 14.18 and Proposition 14.53.

Proof.

Derives Proposition 14.55. For Equation (14.77): \(C_{2}\) is central, so by Corollary 14.18 it acts as a number \(c\) on the irreducible \(\vect{3}\); taking the trace, \(3c = \sum_{a}\tr\left(T_{a}T_{a}\right) = 8 \cdot \tfrac{1}{2} = 4\) by Equation (14.66), so \(c = 4/3\). For Equation (14.78): the same argument with \(\tr\left(T_{a}^{\mathrm{ad}}T_{b}^{\mathrm{ad}}\right) = -f_{acd}f_{bdc} = f_{acd}f_{bcd} = 3\delta_{ab}\) from Equation (14.71) gives \(8c = \sum_{a}3 = 24\), so \(c = 3\).

That \(\bar{\vect{3}}\) is a representation is the observation that complex-conjugating \(\comm{T_{a}}{T_{b}} = \ii f_{abc}T_{c}\) and multiplying by \(-1\) reproduces the same brackets. That it is inequivalent to \(\vect{3}\) is read off the weights. Take the Cartan subalgebra spanned by \(T_{3}\) and \(T_{8}\), which are diagonal; the weights of a representation are the pairs of simultaneous eigenvalues of \(\left(T_{3},T_{8}\right)\), and being eigenvalues they are unchanged by any equivalence. For \(\vect{3}\) they are, from Equations (14.63) and (14.65),

\begin{equation}\tag{14.79} \left(\tfrac{1}{2},\tfrac{1}{2\sqrt{3}}\right)\ec\quad \left(-\tfrac{1}{2},\tfrac{1}{2\sqrt{3}}\right)\ec\quad \left(0,-\tfrac{1}{\sqrt{3}}\right)\ec \end{equation}

summing to zero as tracelessness demands, and for \(\bar{\vect{3}}\) they are the three negatives of these. The two sets are different — an upward-pointing triangle and a downward-pointing one — so no equivalence can carry one representation to the other.

Remark 14.56 (Roots, and why $\vect{3}$ and $\bar{\vect{3}}$ differ here but not for $\SU(2)$).

The roots are the weights of the adjoint representation, that is the eigenvalues of \(\ad_{T_{3}}\) and \(\ad_{T_{8}}\). They are the differences of the weights Equation (14.79),

\begin{equation}\tag{14.80} \pm\left(1,0\right)\ec\qquad \pm\left(\tfrac{1}{2},\tfrac{\sqrt{3}}{2}\right)\ec\qquad \pm\left(-\tfrac{1}{2},\tfrac{\sqrt{3}}{2}\right)\ec \end{equation}

six vectors of unit length at \(60\) degrees to one another, forming a regular hexagon, together with two zeros for the Cartan directions themselves: \(6+2 = 8\) weights for an eight-dimensional representation, as required. For \(\mathfrak{su}(2)\) the rank is one, the two roots are \(\pm1\), and the weight set of any representation is symmetric about zero, which is why there the fundamental and its conjugate are equivalent. The asymmetry of the \(\SU(3)\) weight diagram under \(w \mapsto -w\) is precisely the statement that its fundamental is complex.

Theorem 14.57 ($\vect{3}\otimes\bar{\vect{3}} = \vect{1}\oplus\vect{8}$).
\begin{equation}\tag{14.81} \boxed{\vect{3}\otimes\bar{\vect{3}} = \vect{1}\oplus\vect{8}}\ec \qquad 3 \times 3 = 1 + 8\ep \end{equation}

Rests on Proposition 14.55.

Proof.

Derives Theorem 14.57. Identify \(\C^{3}\otimes\left(\C^{3}\right)^{\ast}\) with the space of complex \(3\times3\) matrices \(\operatorname{End}(\C^{3})\); the action induced by \(\vect{3}\) on the first factor and \(\bar{\vect{3}}\) on the second is \(M \mapsto UMU^{\dagger}\). The trace is invariant under it, so \(\operatorname{End}(\C^{3})\) splits as an \(\SU(3)\)-invariant direct sum

\begin{equation*} \operatorname{End}\left(\C^{3}\right) = \C\,\identity \;\oplus\; \set{M \mid \tr M = 0}\ec \end{equation*}

of dimensions \(1\) and \(8\). The first summand is the trivial representation \(\vect{1}\). The second is the complexification of \(\mathfrak{su}(3)\) — a traceless complex matrix splits uniquely into Hermitian and anti-Hermitian traceless parts — carrying the action \(M \mapsto UMU^{\dagger}\), which is the adjoint action; it is irreducible because an invariant subspace would be an ideal of a simple Lie algebra. Hence it is \(\vect{8}\).

The weights confirm the count independently: the weights of the tensor product are the sums of a weight of \(\vect{3}\) and a weight of \(\bar{\vect{3}}\), that is all differences \(w_{i} - w_{j}\) of the three weights Equation (14.79). The three coincident cases \(i = j\) give the weight zero three times, and the six cases \(i \neq j\) give exactly the six roots Equation (14.80). That is one zero for \(\vect{1}\) plus the six roots and two zeros of \(\vect{8}\), and nothing left over.

Remark 14.58 (What is deliberately not said here).

This subsection has supplied group theory and nothing else. Which physical states, if any, carry the \(\vect{3}\); whether the \(\vect{8}\) is realized by anything observable; what the eight generators have to do with the strong interaction — none of that is a mathematical question, and all of it is settled, where it can be settled, by measurement in Quantum Chromodynamics and Electroweak Unification and the Higgs Boson. The historical route ran the other way: the group was proposed as an organizing symmetry of the observed hadron spectrum [GellMann:1961] [Neeman:1961] before the mathematics above was put to the use it now has [Peskin:1995]. The order in this treatise is the opposite one on purpose — the mathematics is established first and then confronted with evidence, so that agreement counts for something.

The Lie algebra associated with a Lie group

Translations in $D$ dimensions

The algebra of the group of translations in \(D\) dimensions, which we denote by \(\mathfrak{t}(D)\), is generated by

\begin{equation}\tag{14.82} \boxed{P_{A} = \pp_{A}\ec\quad A = 1, \ldots, D} \end{equation}

so that

\begin{equation}\tag{14.83} \boxed{\dim\mathfrak{t}(D) = D}\ep \end{equation}

Let \(f\) be a function of class \(C^{2}\) on \(\R^{D}\).

Then

\begin{align*} \comm{P_{A}}{P_{B}}f &= \comm{\pp_{A}}{\pp_{B}}f\\ &= \pp_{A}\pp_{B}f - \pp_{B}\pp_{A}f\\ \text{($f$ of class $C^{2}$)}\quad &= 0\ec \end{align*}

whence

\begin{equation}\tag{14.84} \boxed{\comm{P_{A}}{P_{B}} = 0}\ep \end{equation}

The algebra $\mathfrak{so}(p,q)$

Generators and commutation relations

Consider \(\R^{p+q}\) endowed with the flat metric \(\eta_{AB}\) of signature \((p,q)\) of Notation 14.1, and in it the unit quadric (a pseudo-sphere) defined by the equation

\begin{equation}\tag{14.85} \left(x^{1}\right)^{2} + \cdots + \left(x^{p}\right)^{2} - \left(x^{p+1}\right)^{2} - \cdots - \left(x^{p+q}\right)^{2} = 1\ep \end{equation}

This space carries the Killing vector fields (cf. Differentiable Manifolds, Tensors, and Curvature and [Wald:1984])

\begin{equation}\tag{14.86} \boxed{J_{AB} = x_{A}\pp_{B} - x_{B}\pp_{A}\ec\quad A, B = 1, \ldots, p+q}\ep \end{equation}

That the fields Equation (14.86) are Killing fields of \(\eta_{AB}\), and that they generate the isometries of the quadric Equation (14.85), is asserted by the source without proof. The source introduces them through the quadric; they are in fact a statement about the flat space itself, and are proved here in that form, the quadric following as a corollary.

Proposition 14.59 (The Killing fields of a flat pseudo-Euclidean space).

Let \(\R^{p+q}\) carry the flat metric \(\eta_{AB}\) of Notation 14.1, \(D = p+q\). Then:

  1. every \(J_{AB}\) of Equation (14.86) is a Killing field;

  2. conversely, every Killing field of \(\eta_{AB}\) is a real linear combination of the \(D\) translations \(P_{A}\) of Equation (14.82) and the \(D(D-1)/2\) fields \(J_{AB}\) with \(A < B\), and those \(D(D+1)/2\) fields are linearly independent. Flat space is therefore maximally symmetric in the sense of Proposition 13.140, and its Killing algebra is the inhomogeneous algebra \(\mathfrak{t}(D)\,\bar{\oplus}\,\mathfrak{so}(p,q)\), which for the Lorentzian signature \(q = 1\) is the Poincaré algebra Equation (14.99);

  3. \(J_{AB}\) annihilates the quadratic form \(x^{C}x_{C}\), so its flow preserves each quadric \(x^{C}x_{C} = \text{const}\), in particular Equation (14.85), and restricts on it to a Killing field of the induced metric.

Rests on Equation (14.86), Proposition 13.140 and Definition 13.139.

Proof.

Derives Proposition 14.59. In the Cartesian coordinates \(x^{A}\) the metric components are constant, so the Christoffel symbols vanish and Killing's equation Equation (13.290) reads

\begin{equation}\tag{14.87} \pp_{C}\xi_{D} + \pp_{D}\xi_{C} = 0\ep \end{equation}

(i) The field \(J_{AB}\) has components \(\xi^{C} = x_{A}\delta^{C}_{B} - x_{B}\delta^{C}_{A}\), hence \(\xi_{D} = x_{A}\eta_{BD} - x_{B}\eta_{AD}\), and with \(\pp_{C}x_{A} = \eta_{AC}\),

\begin{equation*} \pp_{C}\xi_{D} = \eta_{AC}\eta_{BD} - \eta_{BC}\eta_{AD}\ec \end{equation*}

which is antisymmetric under \(C \leftrightarrow D\). So Equation (14.87) holds. That the translations satisfy it is immediate, their components being constant.

(ii) Let \(\xi\) satisfy Equation (14.87) and set \(T_{CDE} = \pp_{C}\pp_{D}\xi_{E}\). It is symmetric in its first two indices, because partial derivatives commute, and antisymmetric in its last two by Equation (14.87). Those two symmetries are incompatible unless \(T\) vanishes:

\begin{equation*} T_{CDE} = T_{DCE} = -T_{DEC} = -T_{EDC} = T_{ECD} = T_{CED} = -T_{CDE}\ec \end{equation*}

each step being one of the two symmetries applied once. Hence \(\pp_{C}\pp_{D}\xi_{E} = 0\), so each \(\xi_{E}\) is a polynomial of degree at most one,

\begin{equation}\tag{14.88} \xi_{E} = a_{E} + \omega_{EC}\,x^{C}\ec\qquad a_{E} = \xi_{E}(0)\ec\quad \omega_{EC} = \pp_{C}\xi_{E}\ec \end{equation}

with constant coefficients, and Equation (14.87) says exactly that \(\omega_{EC} = -\omega_{CE}\). Comparing with the components computed in (i), the constant part is \(a^{E}P_{E}\) and the linear part is \(-\tfrac{1}{2}\omega^{AB}J_{AB}\), a combination of the \(J_{AB}\). Conversely \(\xi\) determines \(a_{E}\) and \(\omega_{EC}\) through Equation (14.88), so distinct pairs \((a,\omega)\) give distinct fields and the \(D + D(D-1)/2 = D(D+1)/2\) listed fields are independent. That number is the maximum allowed by Proposition 13.140. The brackets of these fields are Equations (14.84), (14.90) and (14.97), computed below; they are those of the semidirect sum \(\mathfrak{t}(D)\,\bar{\oplus}\,\mathfrak{so}(p,q)\), which at \(q = 1\) is Equation (14.99).

(iii) Directly,

\begin{equation*} J_{AB}\left(x^{C}x_{C}\right) = \left(x_{A}\pp_{B} - x_{B}\pp_{A}\right)x^{C}x_{C} = 2x_{A}x_{B} - 2x_{B}x_{A} = 0\ec \end{equation*}

so \(J_{AB}\) is tangent to every level set of \(x^{C}x_{C}\) and its flow maps each such quadric to itself. That flow consists of isometries of \(\eta\) by (i), and an isometry of the ambient metric preserving a submanifold restricts to an isometry of the induced metric; by Definition 13.139 the restriction of \(J_{AB}\) is therefore a Killing field of the quadric.

Clearly \(J_{AB} = -J_{BA}\), and therefore

\begin{equation}\tag{14.89} \boxed{\dim\mathfrak{so}(p,q) = \frac{(p+q)(p+q-1)}{2} = \frac{D(D-1)}{2}}\ep \end{equation}

Let us compute the commutation relations among the symmetry generators \(J_{AB}\). Let \(f\) be a function of class \(C^{2}\) on \(\R^{p+q}\). We have

\begin{align*} \comm{J_{AB}}{J_{CD}}f &= \comm{x_{A}\pp_{B} - x_{B}\pp_{A}}{x_{C}\pp_{D} - x_{D}\pp_{C}}f\\ &= \comm{x_{A}\pp_{B}}{x_{C}\pp_{D}}f - \comm{x_{A}\pp_{B}}{x_{D}\pp_{C}}f - \comm{x_{B}\pp_{A}}{x_{C}\pp_{D}}f + \comm{x_{B}\pp_{A}}{x_{D}\pp_{C}}f\ec \end{align*}

where

\begin{align*} \comm{x_{A}\pp_{B}}{x_{C}\pp_{D}}f &= x_{A}\pp_{B}\left(x_{C}\pp_{D}f\right) - x_{C}\pp_{D}\left(x_{A}\pp_{B}f\right)\\ &= x_{A}\eta_{BC}\pp_{D}f + x_{A}x_{C}\pp_{B}\pp_{D}f - x_{C}\eta_{DA}\pp_{B}f - x_{C}x_{A}\pp_{D}\pp_{B}f\\ \text{($f$ of class $C^{2}$)}\quad &= \eta_{BC}\,x_{A}\pp_{D}f - \eta_{DA}\,x_{C}\pp_{B}f\ec \end{align*}

and where we have used that \(\pp_{A}x_{B} = \eta_{BC}\,\pp_{A}x^{C} = \eta_{BC}\,\delta^{C}_{A} = \eta_{AB}\). In this manner we obtain

\begin{align*} \comm{J_{AB}}{J_{CD}} ={}& \eta_{BC}\,x_{A}\pp_{D} - \eta_{DA}\,x_{C}\pp_{B} - \eta_{BD}\,x_{A}\pp_{C} + \eta_{CA}\,x_{D}\pp_{B}\\ &- \eta_{AC}\,x_{B}\pp_{D} + \eta_{DB}\,x_{C}\pp_{A} + \eta_{AD}\,x_{B}\pp_{C} - \eta_{CB}\,x_{D}\pp_{A}\\ ={}& \eta_{AC}\left(x_{D}\pp_{B} - x_{B}\pp_{D}\right) + \eta_{BD}\left(x_{C}\pp_{A} - x_{A}\pp_{C}\right)\\ &+ \eta_{AD}\left(x_{B}\pp_{C} - x_{C}\pp_{B}\right) + \eta_{BC}\left(x_{A}\pp_{D} - x_{D}\pp_{A}\right)\ec \end{align*}

so the commutation relations of the algebra \(\mathfrak{so}(p,q)\) are given by

\begin{equation}\tag{14.90} \boxed{\comm{J_{AB}}{J_{CD}} = \eta_{AC}J_{DB} + \eta_{BD}J_{CA} + \eta_{AD}J_{BC} + \eta_{BC}J_{AD}}\ep \end{equation}

Structure constants

To express the commutation relations Equation (14.90) in the usual form \(\comm{J_{AB}}{J_{CD}} = f_{AB,CD}{}^{EF}J_{EF}\), the symmetries of the indices must be taken into account. A mistake one may commit is to simply attach the corresponding Kronecker deltas, writing

\begin{equation}\tag{14.91} f_{AB,CD}{}^{EF} = \eta_{AC}\delta_{D}^{E}\delta_{B}^{F} + \eta_{BD}\delta_{C}^{E}\delta_{A}^{F} + \eta_{AD}\delta_{B}^{E}\delta_{C}^{F} + \eta_{BC}\delta_{A}^{E}\delta_{D}^{F}\ec \end{equation}

or

\begin{equation}\tag{14.92} f_{AB,CD}{}^{EF} = \eta_{AD}\delta_{B}^{E}\delta_{C}^{F} + \eta_{BC}\delta_{A}^{E}\delta_{D}^{F} - \eta_{AC}\delta_{B}^{E}\delta_{D}^{F} - \eta_{BD}\delta_{A}^{E}\delta_{C}^{F}\ep \end{equation}

However, upon substituting \(J_{AB} = -J_{BA}\) into \(\comm{J_{AB}}{J_{CD}} = f_{AB,CD}{}^{EF}J_{EF}\) it is necessary that \(f_{AB,CD}{}^{EF} = -f_{BA,CD}{}^{EF}\), which is clearly satisfied neither by the quantities Equation (14.91) nor by Equation (14.92). In the same way, one must have \(f_{AB,CD}{}^{EF} = -f_{AB,DC}{}^{EF}\). The only thing that can be said about the upper indices of \(f\) is that \(f\) must not be symmetric under their exchange, since the contraction of symmetric upper indices with the antisymmetric indices of \(J_{EF}\) would force \(\comm{J_{AB}}{J_{CD}} = 0\); moreover, this naive choice of \(f\) does not satisfy the antisymmetry in the lower indices either. Let us see what happens if we antisymmetrize in the indices \(E\) and \(F\):

\begin{align*} \comm{J_{AB}}{J_{CD}} &= \eta_{AC}J_{DB} + \eta_{BD}J_{CA} + \eta_{AD}J_{BC} + \eta_{BC}J_{AD}\\ &= \eta_{AC}\delta_{D}^{E}\delta_{B}^{F}J_{EF} + \eta_{BD}\delta_{C}^{E}\delta_{A}^{F}J_{EF} + \eta_{AD}\delta_{B}^{E}\delta_{C}^{F}J_{EF} + \eta_{BC}\delta_{A}^{E}\delta_{D}^{F}J_{EF}\\ &= \frac{1}{2}\eta_{AC} \left(\delta_{D}^{E}\delta_{B}^{F}+\delta_{D}^{E}\delta_{B}^{F}\right)J_{EF} + \frac{1}{2}\eta_{BD} \left(\delta_{C}^{E}\delta_{A}^{F}+\delta_{C}^{E}\delta_{A}^{F}\right)J_{EF}\\ &\quad+ \frac{1}{2}\eta_{AD} \left(\delta_{B}^{E}\delta_{C}^{F}+\delta_{B}^{E}\delta_{C}^{F}\right)J_{EF} + \frac{1}{2}\eta_{BC} \left(\delta_{A}^{E}\delta_{D}^{F}+\delta_{A}^{E}\delta_{D}^{F}\right)J_{EF}\ec \end{align*}

but, using the antisymmetry of \(J_{EF}\) and relabelling the dummy indices in the second copy of each term,

\begin{align*} \left(\delta_{D}^{E}\delta_{B}^{F}+\delta_{D}^{E}\delta_{B}^{F}\right)J_{EF} &= \delta_{D}^{E}\delta_{B}^{F}J_{EF} + \delta_{D}^{E}\delta_{B}^{F}J_{EF}\\ &= \delta_{D}^{E}\delta_{B}^{F}J_{EF} - \delta_{B}^{F}\delta_{D}^{E}J_{FE}\\ &= \delta_{D}^{E}\delta_{B}^{F}J_{EF} - \delta_{B}^{E}\delta_{D}^{F}J_{EF}\\ &= \left(\delta_{D}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{D}^{F}\right)J_{EF}\ec \end{align*}

whence

\begin{align*} \comm{J_{AB}}{J_{CD}} = \frac{1}{2}\Bigl\{& \eta_{AC}\left(\delta_{D}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{D}^{F}\right) + \eta_{BD}\left(\delta_{C}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{C}^{F}\right)\\ &+ \eta_{AD}\left(\delta_{B}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{B}^{F}\right) + \eta_{BC}\left(\delta_{A}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{A}^{F}\right) \Bigr\}J_{EF}\ep \end{align*}

In this case we have

\begin{align*} f_{AB,CD}{}^{EF} = \frac{1}{2}\Bigl\{& \eta_{AC}\left(\delta_{D}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{D}^{F}\right) + \eta_{BD}\left(\delta_{C}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{C}^{F}\right)\\ &+ \eta_{AD}\left(\delta_{B}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{B}^{F}\right) + \eta_{BC}\left(\delta_{A}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{A}^{F}\right) \Bigr\}\ec \end{align*}

for which we must verify the aforementioned symmetries. In the first place,

\begin{align*} f_{CD,AB}{}^{EF} &= \frac{1}{2}\Bigl\{ \eta_{CA}\left(\delta_{B}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{B}^{F}\right) + \eta_{DB}\left(\delta_{A}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{A}^{F}\right)\\ &\qquad+ \eta_{CB}\left(\delta_{D}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{D}^{F}\right) + \eta_{DA}\left(\delta_{C}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{C}^{F}\right) \Bigr\}\\ &= \frac{1}{2}\Bigl\{ \eta_{AC}\left(\delta_{B}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{B}^{F}\right) + \eta_{BD}\left(\delta_{A}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{A}^{F}\right)\\ &\qquad+ \eta_{BC}\left(\delta_{D}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{D}^{F}\right) + \eta_{AD}\left(\delta_{C}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{C}^{F}\right) \Bigr\}\\ &= -f_{AB,CD}{}^{EF}\ec \end{align*}

so we do have the antisymmetry under exchange of the index pairs. Moreover,

\begin{align*} f_{BA,CD}{}^{EF} &= \frac{1}{2}\Bigl\{ \eta_{BC}\left(\delta_{D}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{D}^{F}\right) + \eta_{AD}\left(\delta_{C}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{C}^{F}\right)\\ &\qquad+ \eta_{BD}\left(\delta_{A}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{A}^{F}\right) + \eta_{AC}\left(\delta_{B}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{B}^{F}\right) \Bigr\}\\ &= -f_{AB,CD}{}^{EF}\ec \end{align*}

so that \(f_{AB,DC}{}^{EF} = -f_{DC,AB}{}^{EF} = f_{CD,AB}{}^{EF} = -f_{AB,CD}{}^{EF}\); and additionally

\begin{align*} f_{AB,CD}{}^{FE} &= \frac{1}{2}\Bigl\{ \eta_{AC}\left(\delta_{D}^{F}\delta_{B}^{E}-\delta_{B}^{F}\delta_{D}^{E}\right) + \eta_{BD}\left(\delta_{C}^{F}\delta_{A}^{E}-\delta_{A}^{F}\delta_{C}^{E}\right)\\ &\qquad+ \eta_{AD}\left(\delta_{B}^{F}\delta_{C}^{E}-\delta_{C}^{F}\delta_{B}^{E}\right) + \eta_{BC}\left(\delta_{A}^{F}\delta_{D}^{E}-\delta_{D}^{F}\delta_{A}^{E}\right) \Bigr\}\\ &= \frac{1}{2}\Bigl\{ \eta_{AC}\left(\delta_{B}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{B}^{F}\right) + \eta_{BD}\left(\delta_{A}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{A}^{F}\right)\\ &\qquad+ \eta_{AD}\left(\delta_{C}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{C}^{F}\right) + \eta_{BC}\left(\delta_{D}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{D}^{F}\right) \Bigr\}\\ &= -f_{AB,CD}{}^{EF}\ep \end{align*}

In summary, the commutation relations of the algebra \(\mathfrak{so}(p,q)\) are given by Equation (14.90) or, equivalently, by

\begin{equation}\tag{14.93} \comm{J_{AB}}{J_{CD}} = f_{AB,CD}{}^{EF}J_{EF}\ec \end{equation}

where the \(f_{AB,CD}{}^{EF}\) are the structure constants, given by

\begin{equation}\tag{14.94} \begin{aligned} f_{AB,CD}{}^{EF} = \frac{1}{2}\Bigl\{& \eta_{AC}\left(\delta_{D}^{E}\delta_{B}^{F}-\delta_{B}^{E}\delta_{D}^{F}\right) + \eta_{BD}\left(\delta_{C}^{E}\delta_{A}^{F}-\delta_{A}^{E}\delta_{C}^{F}\right)\\ &+ \eta_{AD}\left(\delta_{B}^{E}\delta_{C}^{F}-\delta_{C}^{E}\delta_{B}^{F}\right) + \eta_{BC}\left(\delta_{A}^{E}\delta_{D}^{F}-\delta_{D}^{E}\delta_{A}^{F}\right) \Bigr\}\ec \end{aligned} \end{equation}

and which satisfy the following antisymmetry relations:

\begin{equation}\tag{14.95} f_{AB,CD}{}^{EF} = -f_{CD,AB}{}^{EF} = -f_{BA,CD}{}^{EF} = -f_{AB,DC}{}^{EF} = -f_{AB,CD}{}^{FE}\ep \end{equation}

The Lorentz algebra in $D$ dimensions

The Lorentz algebra in \(D\) spacetime dimensions arises as the case of Lorentzian signature, \(\mathfrak{so}(D-1,1)\), of the family \(\mathfrak{so}(p,q)\). At the observed \(D = 4\) it is the algebra \(\mathfrak{so}(3,1)\) of [Wald:1984], whose treatment is four-dimensional throughout; the general-\(D\) statement here is the mathematics of Section 14.3.2 and is not taken from that source. From Equation (14.89),

\begin{equation}\tag{14.96} \dim\mathfrak{so}(D-1,1) = \frac{D(D-1)}{2}\ep \end{equation}

From Equation (14.89) (Lorentzian signature, \(p = D-1\) and \(q = 1\)).

The Poincaré algebra in $D$ dimensions

The Poincaré algebra is the algebra generated by the translations \(P_{A}\) together with the pseudo-rotations \(J_{AB}\). Let us compute the commutation relations between \(J_{AB}\) and \(P_{C}\). Let \(f\) be a function of class \(C^{2}\) on \(\R^{D}\). We have

\begin{align*} \comm{J_{AB}}{P_{C}}f &= \comm{x_{A}\pp_{B} - x_{B}\pp_{A}}{\pp_{C}}f\\ &= \comm{x_{A}\pp_{B}}{\pp_{C}}f - \comm{x_{B}\pp_{A}}{\pp_{C}}f\\ &= \comm{x_{A}}{\pp_{C}}\pp_{B}f + x_{A}\comm{\pp_{B}}{\pp_{C}}f - \comm{x_{B}}{\pp_{C}}\pp_{A}f - x_{B}\comm{\pp_{A}}{\pp_{C}}f\\ \text{($f$ of class $C^{2}$)}\quad &= \comm{x_{A}}{\pp_{C}}\pp_{B}f - \comm{x_{B}}{\pp_{C}}\pp_{A}f\\ &= x_{A}\pp_{C}\pp_{B}f - \pp_{C}\left(x_{A}\pp_{B}f\right) - x_{B}\pp_{C}\pp_{A}f + \pp_{C}\left(x_{B}\pp_{A}f\right)\\ &= x_{A}\pp_{C}\pp_{B}f - \pp_{C}x_{A}\,\pp_{B}f - x_{A}\pp_{C}\pp_{B}f - x_{B}\pp_{C}\pp_{A}f + \pp_{C}x_{B}\,\pp_{A}f + x_{B}\pp_{C}\pp_{A}f\\ &= \pp_{C}x_{B}\,\pp_{A}f - \pp_{C}x_{A}\,\pp_{B}f\\ &= \eta_{CB}\,\pp_{A}f - \eta_{CA}\,\pp_{B}f\ec \end{align*}

so that

\begin{equation}\tag{14.97} \boxed{\comm{J_{AB}}{P_{C}} = \eta_{BC}P_{A} - \eta_{AC}P_{B}}\ec \end{equation}

or equivalently

\begin{equation}\tag{14.98} \boxed{\comm{J_{AB}}{P_{C}} = f_{AB,C}{}^{E}\,P_{E}}\ec \qquad\text{with}\qquad \boxed{f_{AB,C}{}^{E} = \eta_{BC}\,\delta^{E}_{A} - \eta_{AC}\,\delta^{E}_{B}}\ep \end{equation}

Since \(J_{AB}\) and \(P_{A}\) do not commute, the algebra generated by \(\set{P_{A}} \cup \set{J_{AB}}\) is the semidirect sum

\begin{equation}\tag{14.99} \boxed{\mathfrak{iso}(D-1,1) = \mathfrak{t}(D)\,\bar{\oplus}\,\mathfrak{so}(D-1,1)}\ec \end{equation}

whose dimension is

\begin{equation*} \dim\mathfrak{iso}(D-1,1) = D + \frac{D(D-1)}{2}\ec \end{equation*}

that is,

\begin{equation}\tag{14.100} \dim\mathfrak{iso}(D-1,1) = \frac{D(D+1)}{2}\ep \end{equation}

From Equations (14.83) and (14.96) (add the dimensions of the two summands of the semidirect sum).

The Casimir operators of the Poincaré algebra

The Poincaré algebra is not semisimple — Equation (14.99) is a semidirect sum with the abelian ideal \(\mathfrak{t}(D)\) — so Theorem 14.21 does not apply and its invariants must be found by hand (Remark 14.23). Two of them are found here. They matter out of all proportion to the difficulty of the computation, because they are what a particle is: an irreducible representation of this algebra is labelled by their values, and those two numbers are the mass and the spin [Wigner:1939].

Proposition 14.60 (The quadratic invariant).

\(P^{2} = \eta^{AB}P_{A}P_{B}\) is a Casimir element of \(U\left(\mathfrak{iso}(D-1,1)\right)\) in every dimension \(D\). Rests on Equation (14.84), Equation (14.97) and Definition 14.8.

Proof.

Derives Proposition 14.60. \(\comm{P^{2}}{P_{C}} = 0\) is immediate from Equation (14.84). For the pseudo-rotations, use the Leibniz rule and Equation (14.97) in the form \(\comm{J_{AB}}{P^{C}} = \delta^{C}_{B}P_{A} - \delta^{C}_{A}P_{B}\):

\begin{align*} \comm{J_{AB}}{P_{C}P^{C}} &= \comm{J_{AB}}{P_{C}}P^{C} + P_{C}\comm{J_{AB}}{P^{C}}\\ &= \left(\eta_{BC}P_{A} - \eta_{AC}P_{B}\right)P^{C} + P_{C}\left(\delta^{C}_{B}P_{A} - \delta^{C}_{A}P_{B}\right)\\ &= \left(P_{A}P_{B} - P_{B}P_{A}\right) + \left(P_{B}P_{A} - P_{A}P_{B}\right) = 0\ec \end{align*}

each parenthesis vanishing separately by Equation (14.84).

The second invariant is specific to the observed four-dimensional spacetime, because it is built from the Levi-Civita symbol, which carries exactly \(D\) indices and pairs with the two-index \(J\) and the one-index \(P\) only when \(D = 4\). Writing it down requires a decision about index order that the rest of the chapter had no occasion to make, and the decision is recorded before anything is computed.

Notation 14.61 (Signature and index order in this subsubsection).

From here to the end of this subsubsection, and here alone, \(D = 4\) and the physical convention of Particles as Poincaré Representations is adopted: the timelike direction is written \(x^{0}\) and listed first, capital indices run from \(0\) to \(3\), and

\begin{equation}\tag{14.101} \eta_{AB} = \diag(+1,-1,-1,-1)\ec\qquad \epsilon_{0123} = +1\ep \end{equation}

This is the reverse ordering of Notation 14.1, which lists the \(p\) positive directions first and therefore, at \((p,q) = (3,1)\), writes the timelike direction last, as \(x^{4}\) with \(\eta_{44} = -1\). The dictionary between the two is the relabelling \(x^{4} \leftrightarrow x^{0}\) of the timelike coordinate together with an overall change of sign of the metric, \(\eta \mapsto -\eta\).

Nothing in the algebra changes under that dictionary. Lowering an index with \(-\eta\) sends \(x_{A} \mapsto -x_{A}\) and hence \(J_{AB} \mapsto -J_{AB}\), while \(P_{A} = \pp_{A}\) is untouched; both sides of Equation (14.90) are then even and both sides of Equation (14.97) odd, so each relation holds verbatim in either convention. What does change sign is every scalar built by contracting with the inverse metric. \(P^{2} = \eta^{AB}P_{A}P_{B}\) carries one inverse metric and flips. The vector \(W_{A}\) of Equation (14.103) below does not: it carries three raised indices, contributing \((-1)^{3}\), and one factor of \(J\), contributing a fourth sign, so it comes out unchanged. But \(W^{2} = \eta^{AB}W_{A}W_{B}\) again carries one inverse metric, and so flips as well. The two invariants change sign together, so their ratio — which is what carries the spin — is the same in both conventions, and the results of this subsubsection, read in the ordering of Notation 14.1, would appear with both signs reversed: \(\hat{P}^{2} = -m^{2}c^{2}\) and \(\hat{W}^{2} = +m^{2}c^{2}\hbar^{2}s(s+1)\). The convention adopted here is the one under which they agree with Theorem 94.18 and Equation (94.25), which is the form in which the rest of this treatise uses them.

Lemma 14.62 (The Levi-Civita symbol is invariant).

For every antisymmetric \(\omega_{AB} = -\omega_{BA}\) in four dimensions,

\begin{equation}\tag{14.102} \omega^{E}{}_{A}\epsilon_{EBCD} + \omega^{E}{}_{B}\epsilon_{AECD} + \omega^{E}{}_{C}\epsilon_{ABED} + \omega^{E}{}_{D}\epsilon_{ABCE} = 0\ep \end{equation}

Rests on Example 13.94, Equation (13.223) and Notation 14.1.

Proof.

Derives Lemma 14.62. The left-hand side is totally antisymmetric in \(ABCD\): exchanging any two of the four free indices swaps two of the terms with each other and reverses the sign of the remaining two through the antisymmetry of \(\epsilon\). A totally antisymmetric four-index array in four dimensions is a multiple of \(\epsilon_{ABCD}\), and contracting the left-hand side with \(\epsilon^{ABCD}\) shows that the multiple is proportional to \(\omega^{A}{}_{A} = \eta^{AB}\omega_{BA} = 0\), a symmetric array contracted with an antisymmetric one.

Definition 14.63 (Pauli–Lubanski vector).
\begin{equation}\tag{14.103} \boxed{W_{A} = \tfrac{1}{2}\,\epsilon_{ABCD}\,J^{BC}P^{D}}\ep \end{equation}

Rests on Equations (14.82) and (14.86).

Lemma 14.64 (Properties of $W$).
  1. \(\comm{W_{A}}{P_{E}} = 0\);

  2. \(W_{A}P^{A} = P^{A}W_{A} = 0\);

  3. \(W\) transforms as a vector, \(\comm{J_{AB}}{W_{C}} = \eta_{BC}W_{A} - \eta_{AC}W_{B}\).

Rests on Definition 14.63, Equation (14.84), Equation (14.97), Equation (14.90) and Lemma 14.62.

Proof.

Derives Lemma 14.64. (i) By Equation (14.84) only the \(J\) factor contributes, and Equation (14.97) gives

\begin{equation*} \comm{W_{A}}{P_{E}} = \tfrac{1}{2}\epsilon_{ABCD} \left(\delta^{C}_{E}P^{B} - \delta^{B}_{E}P^{C}\right)P^{D} = \tfrac{1}{2}\left(\epsilon_{ABED}P^{B}P^{D} - \epsilon_{AECD}P^{C}P^{D}\right) = 0\ec \end{equation*}

since in each term \(\epsilon\) antisymmetrizes a pair of indices carried by the symmetric product \(P\,P\) of commuting translations.

(ii) \(W_{A}P^{A} = \tfrac{1}{2}\epsilon_{ABCD}J^{BC}P^{D}P^{A}\) vanishes for the same reason, the pair \((A,D)\) now being the symmetric one; the reversed order follows from (i).

(iii) Apply the Leibniz rule to Equation (14.103) using Equation (14.90) for the \(J\) factor and Equation (14.97) for the \(P\) factor. Both brackets are of the form “\(\eta\) times a generator with one index replaced”, that is, exactly the action of an infinitesimal pseudo-rotation on each index of the tensors \(J^{BC}\) and \(P^{D}\). The three contributions are therefore the four terms of Equation (14.102) with \(\omega\) the antisymmetrized index pair \((AB)\), less the one term in which the free index \(C\) of \(W\) is rotated; by Lemma 14.62 the first four cancel and the survivor is

\begin{equation*} \comm{J_{AB}}{W_{C}} = \eta_{BC}W_{A} - \eta_{AC}W_{B}\ep \end{equation*}

In words: \(W\) is a contraction of the invariant tensor \(\epsilon\) with two tensors, so it is a tensor with the one uncontracted index it displays.

Theorem 14.65 (The two Casimir operators of the Poincaré algebra in $3+1$ dimensions).

In \(D = 4\),

\begin{equation}\tag{14.104} \boxed{P^{2} = P_{A}P^{A}\ec\qquad W^{2} = W_{A}W^{A}} \end{equation}

commute with every generator \(P_{C}\) and \(J_{AB}\) of \(\mathfrak{iso}(3,1)\), and they are functionally independent. Rests on Proposition 14.60, Lemma 14.64 and Definition 14.63.

Proof.

Derives Theorem 14.65. \(P^{2}\) is Proposition 14.60. For \(W^{2}\): \(\comm{W^{2}}{P_{E}} = 0\) follows from Lemma 14.64(i) by the Leibniz rule, and \(\comm{W^{2}}{J_{AB}} = 0\) is the computation already done for \(P^{2}\) in Proposition 14.60 with \(P\) replaced by \(W\), which is legitimate because Lemma 14.64(iii) says \(W\) carries its index the same way \(P\) does; the square of a vector is a scalar. Independence: \(W^{2}\) contains the generators \(J^{BC}\), which \(P^{2}\) does not, so neither is a function of the other.

Remark 14.66 (What these two numbers are, in SI units).

The generators above are the differential operators of Equations (14.82) and (14.86), and — unlike the rotation generators of Remark 14.24 — they do not all carry the same dimension. The pseudo-rotation generator \(J_{AB} = x_{A}\pp_{B} - x_{B}\pp_{A}\) is dimensionless, being a coordinate times a derivative with respect to a coordinate; the translation generator \(P_{A} = \pp_{A}\) is not, and carries the SI dimension of inverse length, of unit \(\mathrm{m}^{-1}\). In a quantum theory the observable attached to a generator is \(\hbar\) times it, and the dimensions then come out right in both cases. With \(P_{A} = \pp_{A}\) from Equation (14.82) the momentum operator is \(\hat{P}_{A} = -\ii\hbar\,P_{A}\), of SI unit \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), and with \(J_{AB}\) from Equation (14.86) the angular-momentum operator is \(\hat{J}_{AB} = -\ii\hbar\,J_{AB}\), of unit \(\mathrm{J}\,\mathrm{s}\); both are Hermitian because the generators are anti-Hermitian, and the normalisation is that of Equation (94.3). By Corollary 14.18 each invariant is then a number on an irreducible representation, and the numbers are

\begin{equation}\tag{14.105} \hat{P}^{2} = m^{2}c^{2}\ec\qquad \hat{W}^{2} = -m^{2}c^{2}\,\hbar^{2}\,s(s+1)\ec \end{equation}

defining the mass \(m\), of SI unit \(\mathrm{kg}\), and the spin \(s\), dimensionless, with \(2s \in \N\cup\set{0}\) by the \(\mathfrak{su}(2)\) analysis of Theorem 14.43 applied to the subalgebra that leaves a timelike momentum fixed.

Note the factors, and note in particular that the \(\hbar^{2}\) of the second is not simply one power of \(\hbar\) per generator. Each generator does bring one: \(\hat{P}^{2} = -\hbar^{2}P^{2}\) carries two, and \(\hat{W}_{A} = -\hbar^{2}W_{A}\) carries two, so that \(\hat{W}^{2} = \hbar^{4}W^{2}\) carries four. What survives in the eigenvalues is the ratio,

\begin{equation}\tag{14.106} \frac{\hat{W}^{2}}{\hat{P}^{2}} = -\hbar^{2}\,s(s+1)\ec \end{equation}

and the \(\hbar^{2}\) left standing there is exactly the \(\hbar^{2}\) of the angular-momentum Casimir of Remark 14.24. The mechanism is visible in the rest frame of a timelike momentum, where \(\hat{P}^{D} = mc\,\delta^{D}_{0}\): the definition Equation (14.103) collapses to \(\hat{W}_{0} = 0\) and \(\hat{W}_{i} = -mc\,\hat{\mathcal{J}}_{i}\), with \(\hat{\mathcal{J}}_{i} = \tfrac{1}{2}\epsilon_{ijk}\hat{J}^{jk}\) the angular momentum of the subalgebra that fixes it, of unit \(\mathrm{J}\,\mathrm{s}\), so that \(\hat{W}^{2} = -m^{2}c^{2}\,\hat{\vect{\mathcal{J}}}^{2}\) and the remaining factor is Equation (14.58) with \(j = s\). One factor \(m^{2}c^{2}\) comes from the momentum, one factor \(\hbar^{2}s(s+1)\) from the spin, and that is the whole content of the second line of Equation (14.105). For the same reason the first invariant is \(m^{2}c^{2}\) and not \(m^{2}\): \(\hat{P}_{A}\) is a momentum, and the square of a momentum has the unit \(\mathrm{kg}^{2}\,\mathrm{m}^{2}/\mathrm{s}^{2}\).

The classification of the representations themselves — massive, massless, the little groups, and why a massless particle has helicity rather than spin — is carried out in Particles as Poincaré Representations, whose Theorems 94.18 and 94.29 are Equation (14.105) in their physical setting [Wigner:1939] [Weinberg:1995]. What is established here is only the mathematical fact that the two invariants exist.

Remark 14.67 (Why this is $3+1$ and not general $D$).

Proposition 14.60 holds for every \(D\). The Pauli–Lubanski construction does not: Equation (14.103) contracts a \(D\)-index Levi-Civita symbol with one \(J\) and one \(P\), which balances only when \(D = 4\). In general dimension the higher invariants of \(\mathfrak{iso}(D-1,1)\) are built instead from the antisymmetrized products of \(J\) with \(P\), and how many there are is fixed by the subalgebra that leaves a timelike momentum invariant. That subalgebra is \(\mathfrak{so}(D-1)\), of rank \(\lfloor (D-1)/2 \rfloor\) by Lemma 14.68, and by Theorem 14.21 its representations carry exactly that many labels; together with the mass they make

\begin{equation}\tag{14.107} 1 + \left\lfloor\frac{D-1}{2}\right\rfloor = \left\lceil\frac{D}{2}\right\rceil \end{equation}

labels for a massive irreducible representation — two at \(D = 4\), namely \(P^{2}\) and \(W^{2}\), and three at both \(D = 5\) and \(D = 6\). This is not \(\lfloor D/2 \rfloor\): the two expressions agree in even dimension and differ in odd. This treatise instantiates the observed case, and the observed case is \(D = 4\).

The algebraic half of that count is proved here; what is not proved is identified immediately after it.

Lemma 14.68 (The algebra fixing a timelike momentum).

Let \(D \ge 4\) and let \(p \neq 0\) be a timelike vector of \(\R^{D-1,1}\), that is \(\eta(p,p) < 0\) in the ordering of Notation 14.1. The elements of \(\mathfrak{so}(D-1,1)\) annihilating \(p\) form a subalgebra isomorphic to \(\mathfrak{so}(D-1)\), of rank \(\lfloor (D-1)/2 \rfloor\), and

\begin{equation}\tag{14.108} 1 + \left\lfloor\frac{D-1}{2}\right\rfloor = \left\lceil\frac{D}{2}\right\rceil\ep \end{equation}

Rests on Equation (14.23) and Proposition 14.22.

Proof.

Derives Lemma 14.68. Let \(\mathfrak{h} = \set{X \in \mathfrak{so}(D-1,1) \mid Xp = 0}\); it is a subalgebra, being the kernel of the linear map \(X \mapsto Xp\) and closed under the bracket because \(\comm{X}{Y}p = 0\) whenever \(Xp = Yp = 0\). For \(X \in \mathfrak{h}\) and any \(u\), the defining condition Equation (14.23) gives \(\eta(Xu,p) = -\eta(u,Xp) = 0\), so \(X\) maps the whole space into \(p^{\perp}\) and kills \(p\). Since \(\eta(p,p) \neq 0\) the space splits as \(\R p \oplus p^{\perp}\) with \(\eta\) nondegenerate on each summand, and \(X\) is determined by its restriction to \(p^{\perp}\), which by Equation (14.23) again is an element of the pseudo-orthogonal algebra of \(\eta\) restricted to \(p^{\perp}\). Conversely any such element, extended by \(Xp = 0\), satisfies Equation (14.23) on all of \(\R^{D-1,1}\): both sides vanish when either argument is \(p\). Removing one negative direction from the signature \((D-1,1)\) leaves \(\eta|_{p^{\perp}}\) positive definite in \(D-1\) dimensions, so \(\mathfrak{h} \cong \mathfrak{so}(D-1)\). Its rank is \(\lfloor (D-1)/2 \rfloor\) by Proposition 14.22, applicable because \(D-1 \ge 3\).

For Equation (14.108): if \(D = 2k\) then \(\lfloor (D-1)/2 \rfloor = k-1\) and \(\lceil D/2 \rceil = k\); if \(D = 2k+1\) then \(\lfloor (D-1)/2 \rfloor = k\) and \(\lceil D/2 \rceil = k+1\). In both cases the two sides differ by one.

What the lemma supplies is the algebra: the little algebra and its rank. Turning a rank into a count of labels is representation theory — that a massive irreducible representation of the inhomogeneous algebra is determined by the mass together with an irreducible representation of the subalgebra fixing a timelike momentum, and that no further invariant of the enveloping algebra escapes those labels. That is the method of induced representations. Only the case \(D = 4\), proved in full above, is used anywhere in this book; the general statement is derived in the appendix.

Derives Equation (14.107).

The de Sitter algebras in $D$ dimensions

The anti-de Sitter (AdS) and de Sitter (dS) algebras in \(D\) spacetime dimensions are

\begin{equation*} \operatorname{AdS}:\ \mathfrak{so}(D-1,2)\ec\qquad \operatorname{dS}:\ \mathfrak{so}(D,1)\ec \end{equation*}

with, by Equation (14.89),

\begin{align} \dim\mathfrak{so}(D-1,2) &= \frac{D(D+1)}{2}\ec\tag{14.109}\\ \dim\mathfrak{so}(D,1) &= \frac{D(D+1)}{2}\ep\tag{14.110} \end{align}

The dS and AdS algebras therefore have the same dimension as the Poincaré algebra Equation (14.100). That coincidence of dimensions is not an accident, and the reason is that all three algebras are the Killing algebras of maximally symmetric \(D\)-dimensional spacetimes. The identification of the two written above is proved next; the geometry of the two spaces, and the curvature that distinguishes them, belong to Section 13.13.

Proposition 14.69 (The de~Sitter algebras are isometry algebras).

Let \(D \ge 3\) and \(\ell > 0\). In the flat space of signature \((D-1,2)\) let \(\operatorname{AdS}_{D}\) be the quadric

\begin{equation}\tag{14.111} \eta_{MN}X^{M}X^{N} = -\ell^{2}\ec \end{equation}

and in the flat space of signature \((D,1)\) let \(\operatorname{dS}_{D}\) be the quadric

\begin{equation}\tag{14.112} \eta_{MN}X^{M}X^{N} = +\ell^{2}\ep \end{equation}

Each is a \(D\)-dimensional submanifold whose induced metric has the Lorentzian signature \((D-1,1)\), and the restrictions of the ambient fields \(J_{MN}\) of Equation (14.86) are linearly independent Killing fields of it spanning \(\mathfrak{so}(D-1,2)\) and \(\mathfrak{so}(D,1)\) respectively. These exhaust the Killing fields: the two algebras are the full isometry algebras. Rests on Proposition 14.59, Equation (14.89) and Proposition 13.140.

Proof.

Derives Proposition 14.69. Write \(f(X) = \eta_{MN}X^{M}X^{N}\), so that \(\dd f = 2\eta_{MN}X^{M}\dd X^{N}\) is nonzero at every point of either quadric; each is therefore a regular level set, a submanifold of dimension \(D\), with tangent space at \(X\) the orthogonal complement \(X^{\perp}\). Since \(f(X) \neq 0\) there, the ambient space splits as \(\R X \oplus X^{\perp}\) with \(\eta\) nondegenerate on each summand, so the signature of the induced metric is that of the ambient space less one direction: less a negative one for Equation (14.111) in signature \((D-1,2)\), and less a positive one for Equation (14.112) in signature \((D,1)\). Both give \((D-1,1)\).

By Proposition 14.59(iii) each \(J_{MN}\) is tangent to the quadric and restricts to a Killing field of the induced metric. The restrictions are still independent. Indeed a vanishing combination \(\omega^{MN}J_{MN}\) has ambient components proportional to \(\omega^{M}{}_{N}X^{N}\), so it vanishes on the quadric only if the matrix \(\omega^{M}{}_{N}\) annihilates every point of the quadric, hence the whole span of the quadric. That span is everything. In an indefinite signature the vectors of negative norm already span: a negative basis direction has negative norm, and a positive basis direction \(e\) is the half-sum of \(e + 2u\) and \(e - 2u\) for any negative unit \(u\), both of norm \(-3\). Every vector of negative norm is a real multiple of a point of Equation (14.111), since rescaling by \(\lambda\) multiplies the norm by \(\lambda^{2}\). Interchanging the two signs throughout gives the same conclusion for Equation (14.112). Hence \(\omega = 0\).

The count closes the argument. By Equation (14.89) with \(p + q = D+1\), both algebras have dimension \(D(D+1)/2\), which is the largest number of independent Killing fields a \(D\)-dimensional manifold can carry by Proposition 13.140. No further Killing field can be independent of them, so the isometry algebra is exactly what has been exhibited.

The conformal algebra in $D$ dimensions

The conformal algebra in \(D\) spacetime dimensions is

\begin{equation}\tag{14.113} \mathcal{C}_{D} = \mathfrak{so}(D,2)\ec \end{equation}

with

\begin{equation}\tag{14.114} \dim\mathcal{C}_{D} = \frac{(D+1)(D+2)}{2}\ep \end{equation}

From Equation (14.89) (signature \((D,2)\), so the underlying space has dimension \(D+2\)).

Equation (14.113) is an assertion, and what it asserts is that an algebra defined by a conformal condition on \(D\)-dimensional flat space is the isometry algebra of a flat space with two more dimensions. The conformal Killing fields themselves are written down in Equation (13.333); what remains, and is done here because it is pure Lie algebra, is to compute their brackets and exhibit the isomorphism. The dilatation generator is written \(\Delta\) throughout, the letter \(D\) being reserved for the dimension by Notation 14.1.

Proposition 14.70 (The conformal algebra of flat space).

Let \(D = p+q \ge 3\) and, on \(\R^{p,q}\) with the flat metric of Notation 14.1, put

\begin{equation}\tag{14.115} \boxed{P_{A} = \pp_{A}\ec\quad J_{AB} = x_{A}\pp_{B} - x_{B}\pp_{A}\ec\quad \Delta = x^{C}\pp_{C}\ec\quad K_{A} = 2x_{A}\Delta - x^{2}P_{A}}\ec \end{equation}

with \(x^{2} = x^{C}x_{C}\). These \((D+1)(D+2)/2\) vector fields are the conformal Killing fields Equation (13.333) of the flat metric; they close into a Lie algebra, with brackets Equations (14.84), (14.90) and (14.97) together with

\begin{equation}\tag{14.116} \begin{aligned} \comm{\Delta}{P_{A}} &= -P_{A}\ec & \comm{\Delta}{K_{A}} &= +K_{A}\ec & \comm{\Delta}{J_{AB}} &= 0\ec\\ \comm{K_{A}}{K_{B}} &= 0\ec & \comm{J_{AB}}{K_{C}} &= \eta_{BC}K_{A} - \eta_{AC}K_{B}\ec & \comm{P_{A}}{K_{B}} &= 2\eta_{AB}\Delta - 2J_{AB}\ep \end{aligned} \end{equation}

That algebra is isomorphic to \(\mathfrak{so}(p+1,q+1)\). For Minkowski space, \((p,q) = (D-1,1)\), it is \(\mathfrak{so}(D,2)\), which is Equation (14.113). Rests on Equations (13.333), (14.90) and (14.97).

Proof.

Derives Proposition 14.70. The fields are conformal Killing fields. In Cartesian coordinates the conformal Killing equation reads \(\pp_{B}\xi_{C} + \pp_{C}\xi_{B} = \tfrac{2}{D}\left(\pp^{E}\xi_{E}\right)\eta_{BC}\). For \(P_{A}\) and \(J_{AB}\) both sides vanish, by Proposition 14.59. For \(\Delta\), \(\xi_{C} = x_{C}\) gives \(\pp_{B}\xi_{C} + \pp_{C}\xi_{B} = 2\eta_{BC}\) and \(\pp^{E}\xi_{E} = D\), and the equation holds. For \(K_{A}\), \(\xi_{C} = 2x_{A}x_{C} - x^{2}\eta_{AC}\) gives

\begin{equation*} \pp_{B}\xi_{C} = 2\eta_{AB}x_{C} + 2x_{A}\eta_{BC} - 2x_{B}\eta_{AC}\ec \end{equation*}

whose symmetric part is \(4x_{A}\eta_{BC}\) and whose trace is \(\pp^{E}\xi_{E} = 2x_{A} + 2Dx_{A} - 2x_{A} = 2Dx_{A}\); again the equation holds. That these are all of them is Equation (13.333), and the count is \(D + D(D-1)/2 + 1 + D = (D+1)(D+2)/2\).

The brackets. Every computation below uses only the derivation rule \(\comm{X}{fY} = X(f)\,Y + f\comm{X}{Y}\) for a function \(f\), the values

\begin{equation*} P_{A}(x_{B}) = \eta_{AB}\ec\quad \Delta(x_{A}) = x_{A}\ec\quad \Delta(x^{2}) = 2x^{2}\ec\quad J_{AB}(x_{C}) = x_{A}\eta_{BC} - x_{B}\eta_{AC}\ec\quad J_{AB}(x^{2}) = 0\ec \end{equation*}

the last from Proposition 14.59(iii), and the identity \(x_{A}P_{B} - x_{B}P_{A} = J_{AB}\).

For \(\comm{\Delta}{P_{A}}\): \(\comm{x^{C}\pp_{C}}{\pp_{A}} = -\left(\pp_{A}x^{C}\right)\pp_{C} = -P_{A}\). For \(\comm{\Delta}{J_{AB}}\), the same computation on each of the two terms of \(J_{AB}\) gives \(x_{A}P_{B} - x_{A}P_{B}\) and so zero. For \(\comm{\Delta}{K_{A}}\),

\begin{equation*} \comm{\Delta}{2x_{A}\Delta - x^{2}P_{A}} = 2x_{A}\Delta - 2x^{2}P_{A} - x^{2}\comm{\Delta}{P_{A}} = 2x_{A}\Delta - x^{2}P_{A} = K_{A}\ep \end{equation*}

For \(\comm{P_{A}}{K_{B}}\), using \(\comm{P_{A}}{\Delta} = P_{A}\) and \(P_{A}(x^{2}) = 2x_{A}\),

\begin{equation*} \comm{P_{A}}{K_{B}} = 2\eta_{AB}\Delta + 2x_{B}P_{A} - 2x_{A}P_{B} = 2\eta_{AB}\Delta - 2J_{AB}\ep \end{equation*}

For \(\comm{J_{AB}}{K_{C}}\), using \(\comm{J_{AB}}{\Delta} = 0\), \(J_{AB}(x^{2}) = 0\) and Equation (14.97),

\begin{align*} \comm{J_{AB}}{K_{C}} &= 2\left(x_{A}\eta_{BC} - x_{B}\eta_{AC}\right)\Delta - x^{2}\left(\eta_{BC}P_{A} - \eta_{AC}P_{B}\right)\\ &= \eta_{BC}\left(2x_{A}\Delta - x^{2}P_{A}\right) - \eta_{AC}\left(2x_{B}\Delta - x^{2}P_{B}\right) = \eta_{BC}K_{A} - \eta_{AC}K_{B}\ep \end{align*}

Finally, with \(K_{A}(x_{B}) = 2x_{A}x_{B} - x^{2}\eta_{AB}\), \(K_{A}(x^{2}) = 4x_{A}x^{2} - 2x_{A}x^{2} = 2x_{A}x^{2}\), \(\comm{K_{A}}{\Delta} = -K_{A}\) and \(\comm{K_{A}}{P_{B}} = -2\eta_{AB}\Delta - 2J_{AB}\),

\begin{align*} \comm{K_{A}}{K_{B}} &= 2\left(2x_{A}x_{B} - x^{2}\eta_{AB}\right)\Delta - 2x_{B}K_{A} - 2x_{A}x^{2}P_{B} + x^{2}\left(2\eta_{AB}\Delta + 2J_{AB}\right)\\ &= 4x_{A}x_{B}\Delta - 2x_{B}\left(2x_{A}\Delta - x^{2}P_{A}\right) - 2x^{2}x_{A}P_{B} + 2x^{2}J_{AB}\\ &= 2x^{2}\left(x_{B}P_{A} - x_{A}P_{B}\right) + 2x^{2}J_{AB} = 0\ep \end{align*}

The isomorphism. Let the indices \(M,N\) run over \(1,\ldots,D+2\), extend \(\eta\) by

\begin{equation}\tag{14.117} \eta_{D+1,D+1} = -1\ec\qquad \eta_{D+2,D+2} = +1\ec \end{equation}

all other new components vanishing, and set

\begin{equation}\tag{14.118} \boxed{J_{A,D+1} = \tfrac{1}{2}\left(K_{A}-P_{A}\right)\ec\quad J_{A,D+2} = \tfrac{1}{2}\left(K_{A}+P_{A}\right)\ec\quad J_{D+1,D+2} = \Delta}\ec \end{equation}

the remaining \(J_{MN}\) being the \(J_{AB}\) already defined and \(J_{NM} = -J_{MN}\). This is a linear bijection onto the span, since \(P_{A}\) and \(K_{A}\) are recovered as \(J_{A,D+2} \mp J_{A,D+1}\). It remains to check that Equation (14.90) holds for every choice of the four indices among \(1,\ldots,D+2\). With four indices in \(1,\ldots,D\) it is Equation (14.90) itself. With exactly one index equal to \(D+1\) or \(D+2\) it is the bracket \(\comm{J_{AB}}{K_{C}}\) of Equation (14.116) together with Equation (14.97), both of which say that \(K\) and \(P\) carry their index as vectors. The case \(\comm{J_{AB}}{J_{D+1,D+2}} = \comm{J_{AB}}{\Delta} = 0\) agrees with Equation (14.90), every term of which carries a vanishing \(\eta\). There remain the pairs built from Equation (14.118). Using \(\comm{K_{A}}{P_{B}} = -2\eta_{AB}\Delta - 2J_{AB}\) and \(\comm{P_{A}}{K_{B}} = 2\eta_{AB}\Delta - 2J_{AB}\),

\begin{align*} \comm{J_{A,D+1}}{J_{B,D+1}} &= \tfrac{1}{4}\comm{K_{A}-P_{A}}{K_{B}-P_{B}} = J_{AB}\ec\\ \comm{J_{A,D+2}}{J_{B,D+2}} &= \tfrac{1}{4}\comm{K_{A}+P_{A}}{K_{B}+P_{B}} = -J_{AB}\ec\\ \comm{J_{A,D+1}}{J_{B,D+2}} &= \tfrac{1}{4}\comm{K_{A}-P_{A}}{K_{B}+P_{B}} = -\eta_{AB}\Delta\ec \end{align*}

while Equation (14.90) demands, for these three, \(\eta_{D+1,D+1}J_{BA} = J_{AB}\), \(\eta_{D+2,D+2}J_{BA} = -J_{AB}\) and \(\eta_{AB}J_{D+2,D+1} = -\eta_{AB}\Delta\), which are the same three results. Likewise

\begin{equation*} \comm{J_{D+1,D+2}}{J_{A,D+1}} = \comm{\Delta}{\tfrac{1}{2}\left(K_{A}-P_{A}\right)} = J_{A,D+2}\ec\qquad \comm{J_{D+1,D+2}}{J_{A,D+2}} = J_{A,D+1}\ec \end{equation*}

against the demands \(\eta_{D+1,D+1}J_{D+2,A} = J_{A,D+2}\) and \(\eta_{D+2,D+2}J_{A,D+1} = J_{A,D+1}\) of Equation (14.90). All cases agree, so the fields Equation (14.115) satisfy the brackets of \(\mathfrak{so}\) of the extended metric Equation (14.117). That metric adds one positive and one negative direction, so its signature is \((p+1,q+1)\) after the coordinates are reordered to list the positive directions first.

Remark 14.71 (Two spaces of two extra dimensions, and neither is physical).

Propositions 14.69 and 14.70 both realize a \(D\)-dimensional symmetry linearly on a flat space of more dimensions — one more for the de Sitter quadrics, two more for the conformal algebra. The extra directions are a device for linearizing an action, not a claim about the world: the manifold in both cases remains \(D\)-dimensional, and this treatise instantiates \(D = 4\). Nothing here asserts the existence of extra dimensions, which have no evidence base and are excluded by the scope rule of Epistemology and the Scientific Method. The geometric side of the same statement, including the projective null-cone realization of the conformal action, is Section 13.13.1.

Central extensions

An algebra of generators need not close on itself. It may close only on itself together with an extra element that commutes with everything and generates no transformation—a central charge. This is not a pathology: two of the most important constants in physics, the mass in nonrelativistic mechanics and Planck's constant in the canonical commutation relations, appear in exactly this position, and neither can be removed by any redefinition of the generators. This section says what the obstruction is and how it is classified.

Extensions and cocycles

Definition 14.72 (Central extension).

A central extension of a Lie algebra \(\mathfrak{g}\) by \(\R\) is a Lie algebra \(\tilde{\mathfrak{g}}\) fitting into the exact sequence

\begin{equation}\tag{14.119} 0\longrightarrow\R\longrightarrow\tilde{\mathfrak{g}} \longrightarrow\mathfrak{g}\longrightarrow0\ec \end{equation}

with the image of \(\R\) contained in the centre of \(\tilde{\mathfrak{g}}\). Choosing a linear splitting \(\xi\mapsto\tilde{\xi}\) of the projection, the bracket of \(\tilde{\mathfrak{g}}\) takes the form

\begin{equation}\tag{14.120} \boxed{\comm{\tilde{\xi}}{\tilde{\eta}} =\widetilde{\comm{\xi}{\eta}}+c(\xi,\eta)\,Z}\ec \end{equation}

where \(Z\) is the central generator and \(c:\mathfrak{g}\times\mathfrak{g}\to\R\) is a bilinear map.

Proposition 14.73 (The extension datum is a $2$-cocycle).

The map \(c\) of Equation (14.120) is antisymmetric and satisfies

\begin{equation}\tag{14.121} \boxed{c\left(\comm{\xi}{\eta},\zeta\right) +c\left(\comm{\eta}{\zeta},\xi\right) +c\left(\comm{\zeta}{\xi},\eta\right)=0}\ec \end{equation}

for all \(\xi,\eta,\zeta\in\mathfrak{g}\). A bilinear map with these two properties is called a \(2\)-cocycle on \(\mathfrak{g}\). Rests on Definition 14.72.

Proof.

Derives Proposition 14.73. Antisymmetry of \(c\) follows from antisymmetry of the bracket in Equation (14.120). Imposing the Jacobi identity on \(\tilde{\mathfrak{g}}\) for the three elements \(\tilde{\xi},\tilde{\eta},\tilde{\zeta}\), the terms not involving \(Z\) cancel because the Jacobi identity already holds in \(\mathfrak{g}\), and \(Z\) is central so it contributes nothing to a nested bracket; what survives is the coefficient of \(Z\), which is exactly Equation (14.121).

Definition 14.74 (Coboundary and triviality).

A cocycle is a coboundary if there is a linear map \(b:\mathfrak{g}\to\R\) with

\begin{equation}\tag{14.122} c(\xi,\eta)=b\left(\comm{\xi}{\eta}\right)\ep \end{equation}

An extension whose cocycle is a coboundary is called trivial: the shift \(\tilde{\xi}\mapsto\tilde{\xi}+b(\xi)Z\) removes \(c\) from Equation (14.120) entirely, so that \(\tilde{\mathfrak{g}}\cong\mathfrak{g}\oplus\R\) as a direct sum of Lie algebras and the central charge decouples. Rests on Proposition 14.73.

Definition 14.75 (The classifying group).

The \(2\)-cocycles of Proposition 14.73 form a real vector space and the coboundaries of Definition 14.74 a subspace of it. The quotient is the second Chevalley–Eilenberg cohomology group \(H^{2}(\mathfrak{g},\R)\). Two central extensions of \(\mathfrak{g}\) by \(\R\) are called equivalent when there is an isomorphism of Lie algebras between them commuting with the inclusion of \(\R\) and with the projection onto \(\mathfrak{g}\) in Equation (14.119). Rests on Proposition 14.73 and Definition 14.74.

That \(H^{2}\) deserves to be called the classifying group is a theorem, not part of the definition, and it is proved next.

Proposition 14.76 ($H^{2}$ classifies the central extensions).

The assignment Equation (14.120) of a cocycle to an extension induces a bijection between \(H^{2}(\mathfrak{g},\R)\) and the set of central extensions of \(\mathfrak{g}\) by \(\R\) up to equivalence, the zero class corresponding to the trivial extension. In particular

\begin{equation}\tag{14.123} \boxed{H^{2}(\mathfrak{g},\R)=0 \iff \text{every central extension of }\mathfrak{g}\text{ is trivial}}\ep \end{equation}

Rests on Definition 14.75, Definition 14.72, Proposition 14.73 and Definition 14.74.

Proof.

Derives Proposition 14.76. Write \(\partial b\) for the coboundary of Equation (14.122), \((\partial b)(\xi,\eta)=b\left(\comm{\xi}{\eta}\right)\). Three things have to be checked.

The class does not depend on the splitting. Two linear splittings of the projection differ by a linear map \(\mathfrak{g}\to\R\): if \(\xi\mapsto\tilde{\xi}\) and \(\xi\mapsto\tilde{\xi}'\) both project to \(\xi\), their difference lies in the kernel of the projection, which is the line spanned by \(Z\), so \(\tilde{\xi}'=\tilde{\xi}+b(\xi)Z\) with \(b\) linear. Since \(Z\) is central, \(\comm{\tilde{\xi}'}{\tilde{\eta}'}=\comm{\tilde{\xi}}{\tilde{\eta}}\), while the first term of Equation (14.120) shifts by \(b\left(\comm{\xi}{\eta}\right)Z\). Hence \(c'=c-\partial b\) and the two cocycles define the same class. An isomorphism of extensions in the sense of Definition 14.75 is the identity on \(\R\) and induces the identity on \(\mathfrak{g}\), so it acts on a splitting exactly as a change of splitting does; equivalent extensions therefore have equal classes.

Every class is realized. Given a cocycle \(c\), put \(\tilde{\mathfrak{g}}=\mathfrak{g}\oplus\R Z\) as a vector space, with \(Z\) central and the bracket Equation (14.120). It is antisymmetric because \(c\) is, and it satisfies the Jacobi identity: the \(\mathfrak{g}\)-component of the Jacobi expression is the Jacobi identity of \(\mathfrak{g}\), and the coefficient of \(Z\) is Equation (14.121). The obvious inclusion and projection make it a central extension, whose cocycle for the obvious splitting is \(c\) again.

Cohomologous cocycles give equivalent extensions, and only those. A map between the extensions built from \(c\) and from \(c'\) that commutes with the inclusion and the projection is the identity on \(\R Z\) and induces the identity on \(\mathfrak{g}\), so it sends \(\tilde{\xi}\mapsto\tilde{\xi}+\beta(\xi)Z\) for some linear \(\beta:\mathfrak{g}\to\R\) and is automatically bijective. Requiring it to respect the two brackets and reading off the coefficient of \(Z\) gives \(c'=c+\partial\beta\). So the two extensions are equivalent precisely when \(c'-c\) is a coboundary, that is, precisely when \(c\) and \(c'\) define the same class. The trivial extension is the one with \(c=0\), which is Definition 14.74, and it represents the zero class.

Proposition 14.77 (Semisimple algebras admit no nontrivial extension).

If \(\mathfrak{g}\) is a finite-dimensional semisimple Lie algebra over a field of characteristic zero, then \(H^{1}(\mathfrak{g},\R)=H^{2}(\mathfrak{g},\R)=0\). In particular \(\mathfrak{so}(p,q)\), and hence the Lorentz algebra \(\mathfrak{so}(D-1,1)\) of Equation (14.90), admits no nontrivial central extension. Rests on Definition 14.75 and Theorem 14.13.

These are Whitehead's first and second lemmas. Both are proved by averaging a cochain against the Casimir element of the algebra, which Theorem 14.13 makes invertible on any module with no trivial summand; the corollary for the Lorentz algebra then follows from its semisimplicity, which is itself the nondegeneracy of the Killing form.

Derives Proposition 14.77.

Remark 14.78 (Where the obstruction can live).

Proposition 14.77 localises the interesting cases. A semisimple algebra is rigid in this respect, so nontrivial central charges occur only where the algebra is not semisimple: in abelian algebras, in the inhomogeneous algebras built as semidirect sums such as Equation (14.99), in algebras obtained by the contractions of Section 14.5, and in infinite-dimensional algebras. All four cases occur in physics.

The instances that occur in physics

Example 14.79 (The Heisenberg algebra).

Take \(\mathfrak{g}=\R^{2f}\), abelian, spanned by the phase-space translations. Every antisymmetric bilinear form on an abelian algebra is a cocycle, since Equation (14.121) is vacuous, and no coboundary is available because \(\comm{\xi}{\eta}=0\) makes the right-hand side of Equation (14.122) vanish identically. Hence \(H^{2}(\R^{2f},\R)=\Lambda^{2}(\R^{2f})^{*}\), which is as far from zero as possible. The extension defined by the symplectic form is the Heisenberg algebra

\begin{equation}\tag{14.124} \comm{X_{i}}{P_{j}}=\delta_{ij}Z\ec\qquad \comm{X_{i}}{X_{j}}=\comm{P_{i}}{P_{j}}=\comm{Z}{\cdot}=0\ec \end{equation}

and the canonical commutation relations of quantum mechanics are Equation (14.124) with \(Z=\ii\hbar\). Planck's constant is a central charge. Rests on Proposition 14.73 and Definition 14.74.

Example 14.80 (The Galilei algebra).

The ten-generator Galilei algebra of three space dimensions carries the cocycle \(c(K_{i},P_{j})=\delta_{ij}\), whose central charge is the mass, and \(H^{2}\) of that algebra is one-dimensional, so up to a scale this is the only central extension it admits. The classical mechanics of a free particle realizes the extension and not the Galilei algebra itself; the computation, and the proof that this cocycle is not a coboundary, are carried out in Proposition 25.14. That proposition establishes only that the class is nonzero; that it exhausts \(H^{2}\) is the stronger statement, and it needs the cocycle condition solved on the whole ten-dimensional algebra and quotiented by the coboundaries. Rests on Definition 14.75 and Proposition 25.14.

Derives Example 14.80.

Example 14.81 (Infinite-dimensional cases).

The Witt algebra of vector fields on the circle has a one-dimensional \(H^{2}\), whose extension is the Virasoro algebra with central charge \(c\); loop algebras extend to affine Kac–Moody algebras with the level as central charge; and the equal-time current-algebra commutators of a field theory acquire Schwinger terms of the same nature.

All three are computations rather than citations: the Virasoro cocycle and its uniqueness up to coboundary, the Kac–Moody cocycle built from the Killing form and the residue pairing, and the Schwinger term as the statement that a formally central extension survives regularization.

Derives Example 14.81.

Remark 14.82 (Central charges are measured, not postulated).

The four examples above are not on the same evidential footing, and the distinction matters under the scope rule of this treatise. Planck's constant and the mass are measured quantities appearing as central charges of algebras whose representations describe observed systems in the observed \(3+1\) dimensions; nothing stands between the algebra and the instrument.

The Virasoro central charge is not on that footing, and this treatise records no measurement of it. Where it is called observable, the claim runs through the universal critical behaviour of a laboratory sample — a three-dimensional piece of matter, in the regime treated in Phase Transitions and Critical Phenomena, whose long-wavelength behaviour is that of an effectively two-dimensional system when one direction is frozen out by geometry or by temperature. What is measured there are critical exponents and amplitude ratios; the step from those numbers to a value of \(c\) is an identification made inside a conformal field theory, not a reading of an instrument, and this book does not make it. The system is \(3+1\)-dimensional and the effective description is not a spacetime.

The Schwinger terms are on a third footing again. They are the equal-time-commutator face of the anomalies [Adler:1969] [Bell:1969], and it is the anomaly that carries the measurable consequence — the cleanest case being the two-photon decay of the neutral pion, whose rate has been measured directly [Larin:2020]. That the Schwinger term and the anomaly are two faces of one obstruction is asserted here and not proved; what can be said without qualification is that the decay rate is measured and that an anomaly-free theory predicts the wrong one.

What is not evidence in any of these cases is the existence of an algebra as such: writing down a consistent extension is a mathematical act, and whether nature realizes it is a separate question, settled only by experiment.

Contraction of Lie algebras

The Inönü–Wigner limit

A symmetry algebra is rarely exact. More often it is the limit of another one as some parameter—a velocity ratio, a curvature radius—is taken to an extreme, and the limit is singular: the generators must be rescaled as it is taken, or the brackets diverge. Inönü and Wigner made the procedure precise [Inonu:1953].

Definition 14.83 (Contraction).

Let \(\mathfrak{g}\) be a Lie algebra with basis \(\set{T_{\alpha}}\) and let \(U(\varepsilon)\) be a family of invertible linear maps, singular as \(\varepsilon\to0\). If the limit

\begin{equation}\tag{14.125} \boxed{\comm{T_{\alpha}}{T_{\beta}}_{0} =\lim_{\varepsilon\to0}U(\varepsilon)^{-1} \comm{U(\varepsilon)T_{\alpha}}{U(\varepsilon)T_{\beta}}} \end{equation}

exists for all \(\alpha,\beta\), it defines a Lie algebra \(\mathfrak{g}_{0}\) of the same dimension, called a contraction of \(\mathfrak{g}\).

Proposition 14.84 (Properties of a contraction).

Let \(\mathfrak{g}_{0}\) be a contraction of \(\mathfrak{g}\) in the sense of Definition 14.83, and fix a basis \(\set{T_{\alpha}}\), in which \(U(\varepsilon)\) has matrix \(U\) and \(\mathfrak{g}\) has structure constants \(C^{\gamma}{}_{\alpha\beta}\) and Killing form \(\kappa\). Then:

  1. \(\comm{\cdot}{\cdot}_{0}\) is a Lie bracket, and \(\dim\mathfrak{g}_{0} = \dim\mathfrak{g}\);

  2. the derivation algebras obey \(\dim\operatorname{Der}(\mathfrak{g}_{0}) \ge \dim\operatorname{Der}(\mathfrak{g})\), so that a contraction is not reversible: whenever the inequality is strict, the two algebras are not isomorphic and \(\mathfrak{g}\) is not a contraction of \(\mathfrak{g}_{0}\);

  3. the Killing form of \(\mathfrak{g}_{0}\) is \(\kappa_{0} = \lim_{\varepsilon\to0}\kappa_{\varepsilon}\) with \(\kappa_{\varepsilon}(X,Y) = \kappa\left(U(\varepsilon)X, U(\varepsilon)Y\right)\), so that \(\det\kappa_{0} = \lim_{\varepsilon\to0}\left(\det U\right)^{2}\det\kappa\). If \(\det U(\varepsilon) \rightarrow 0\) and \(\mathfrak{g}\) is semisimple, then \(\mathfrak{g}_{0}\) is not;

  4. a contraction can create the cohomology that a central charge needs. Every Lie algebra of dimension \(n\) contracts to the abelian algebra of the same dimension, whose second cohomology is the whole space \(\Lambda^{2}\left(\R^{n}\right)^{\ast}\) of antisymmetric bilinear forms by the argument of Example 14.79, while \(H^{2}\) vanishes for a semisimple \(\mathfrak{g}\) by Proposition 14.77.

Rests on Definition 14.83, Definition 14.10 and Proposition 14.77.

Proof.

Derives Proposition 14.84. Throughout, write \(\comm{X}{Y}_{\varepsilon} = U(\varepsilon)^{-1}\comm{U(\varepsilon)X}{U(\varepsilon)Y}\) for the bracket at finite \(\varepsilon\). It is the pullback of the bracket of \(\mathfrak{g}\) along the linear isomorphism \(U(\varepsilon)\), so \(\left(\mathfrak{g},\comm{\cdot}{\cdot}_{\varepsilon}\right)\) is a Lie algebra isomorphic to \(\mathfrak{g}\) for every \(\varepsilon \neq 0\), and its structure constants \(C^{\gamma}{}_{\alpha\beta}(\varepsilon)\) converge, by hypothesis, to those of \(\mathfrak{g}_{0}\).

(i) Bilinearity and antisymmetry pass to the limit because they are linear conditions on the structure constants. The Jacobi identity is the system of quadratic equations \(C^{\delta}{}_{\alpha\beta}(\varepsilon)C^{\sigma}{}_{\delta\gamma} (\varepsilon) + C^{\delta}{}_{\beta\gamma}(\varepsilon)C^{\sigma}{}_{\delta\alpha} (\varepsilon) + C^{\delta}{}_{\gamma\alpha}(\varepsilon)C^{\sigma}{}_{\delta\beta} (\varepsilon) = 0\), valid for every \(\varepsilon \neq 0\); a polynomial is continuous, so the same equations hold in the limit. The underlying vector space never changed, whence the dimension.

(ii) A derivation is a solution of the linear system \(d\left(\comm{T_{\alpha}}{T_{\beta}}\right) = \comm{dT_{\alpha}}{T_{\beta}} + \comm{T_{\alpha}}{dT_{\beta}}\) in the unknown matrix \(d\), and the coefficients of that system are linear in the structure constants. Write \(L(\varepsilon)\) for its matrix. Since \(\operatorname{Der}\) of isomorphic algebras are conjugate, \(\dim\ker L(\varepsilon) = \dim\operatorname{Der}(\mathfrak{g})\) for all \(\varepsilon \neq 0\). Rank is lower semicontinuous: if \(L(0)\) carries a nonvanishing minor of size \(r\), that same minor is a continuous function of \(\varepsilon\) and so is nonvanishing for small \(\varepsilon\), whence \(\operatorname{rank}\,L(0) \le \operatorname{rank}\,L(\varepsilon)\) there, and the kernels satisfy the reverse inequality. That is the claim. If it is strict, then \(\mathfrak{g}_{0}\) is not isomorphic to \(\mathfrak{g}\), and a contraction the other way would give \(\dim\operatorname{Der}(\mathfrak{g}) \ge \dim\operatorname{Der}(\mathfrak{g}_{0})\) and hence equality, which is excluded.

(iii) From \(\ad^{\varepsilon}_{X} = U(\varepsilon)^{-1}\ad_{U(\varepsilon)X} U(\varepsilon)\) and cyclicity of the trace,

\begin{equation*} \kappa_{\varepsilon}(X,Y) = \tr\left(\ad^{\varepsilon}_{X}\ad^{\varepsilon}_{Y}\right) = \tr\left(\ad_{U(\varepsilon)X}\ad_{U(\varepsilon)Y}\right) = \kappa\left(U(\varepsilon)X,U(\varepsilon)Y\right)\ec \end{equation*}

which in the fixed basis is the matrix \(\bigl(U\bigr)\transpose\kappa\,U\), of determinant \(\left(\det U\right)^{2}\det\kappa\). The Killing form is a quadratic polynomial in the structure constants by Equation (14.15), so \(\kappa_{0}\) is the limit of \(\kappa_{\varepsilon}\) and \(\det\kappa_{0}\) the limit of the determinants. If \(\det U \rightarrow 0\) then \(\det\kappa_{0} = 0\), the Killing form of \(\mathfrak{g}_{0}\) is degenerate, and Theorem 14.13 denies it semisimplicity.

(iv) Take \(U(\varepsilon) = \varepsilon\,\identity\). Then \(\comm{X}{Y}_{\varepsilon} = \varepsilon^{-1}\comm{\varepsilon X}{\varepsilon Y} = \varepsilon\comm{X}{Y} \rightarrow 0\), so every algebra contracts to the abelian one of the same dimension; and \(\det U = \varepsilon^{n} \rightarrow 0\), consistently with (iii). The second cohomology of an abelian algebra is computed in Example 14.79 and is the whole space of antisymmetric bilinear forms, nonzero as soon as \(n \ge 2\).

Remark 14.85 (Why (iv) is only an existence statement).

Part (iv) proves that a contraction can create a nonvanishing \(H^{2}\), by exhibiting the extreme case. It does not say that any particular contraction does so, and no general theorem does: which classes survive a given limit has to be computed limit by limit. What the physical cases have in common is the structure of Lemma 14.86 — an abelian ideal appears where there was none — and it is on an abelian ideal that a cocycle has the most room, by the same computation as in Example 14.79.

Lemma 14.86 (The Inönü–Wigner contraction along a subalgebra).

Let \(\mathfrak{g} = \mathfrak{h}\oplus\mathfrak{m}\) as a vector space, let \(\pi\) be the projection onto \(\mathfrak{m}\) along \(\mathfrak{h}\), and let \(U(\varepsilon)\) be the identity on \(\mathfrak{h}\) and multiplication by \(\varepsilon\) on \(\mathfrak{m}\). The limit Equation (14.125) exists if and only if \(\mathfrak{h}\) is a subalgebra, and then

\begin{equation}\tag{14.126} \boxed{\comm{X}{Y}_{0} = \comm{X}{Y}\ec\quad \comm{X}{Z}_{0} = \pi\comm{X}{Z}\ec\quad \comm{Z}{Z'}_{0} = 0}\ec \end{equation}

for \(X, Y \in \mathfrak{h}\) and \(Z, Z' \in \mathfrak{m}\). In particular \(\mathfrak{m}\) is an abelian ideal of \(\mathfrak{g}_{0}\) and \(\mathfrak{g}_{0}/\mathfrak{m} \cong \mathfrak{h}\). Rests on Definition 14.83.

Proof.

Derives Lemma 14.86. For \(X, Y \in \mathfrak{h}\), \(U(\varepsilon)\) acts as the identity and \(U(\varepsilon)^{-1}\) multiplies the \(\mathfrak{m}\)-part by \(\varepsilon^{-1}\), so \(\comm{X}{Y}_{\varepsilon} = \left(1-\pi\right)\comm{X}{Y} + \varepsilon^{-1}\pi\comm{X}{Y}\), which converges if and only if \(\pi\comm{X}{Y} = 0\), that is if and only if \(\comm{\mathfrak{h}}{\mathfrak{h}} \subseteq \mathfrak{h}\); the limit is then \(\comm{X}{Y}\). For \(X \in \mathfrak{h}\) and \(Z \in \mathfrak{m}\),

\begin{equation*} \comm{X}{Z}_{\varepsilon} = U(\varepsilon)^{-1}\comm{X}{\varepsilon Z} = \varepsilon\left(1-\pi\right)\comm{X}{Z} + \pi\comm{X}{Z} \longrightarrow \pi\comm{X}{Z}\ec \end{equation*}

and for \(Z, Z' \in \mathfrak{m}\), \(\comm{Z}{Z'}_{\varepsilon} = \varepsilon^{2}\left(1-\pi\right)\comm{Z}{Z'} + \varepsilon\pi\comm{Z}{Z'} \rightarrow 0\). Both limits exist unconditionally, so the condition on \(\mathfrak{h}\) is the whole condition. Equation (14.126) then puts \(\comm{\mathfrak{g}_{0}}{\mathfrak{m}}_{0} \subseteq \mathfrak{m}\) and \(\comm{\mathfrak{m}}{\mathfrak{m}}_{0} = 0\), and the quotient bracket is that of \(\mathfrak{h}\).

Lemma 14.87 (A rescaled Casimir stays central).

Let \(C \in U(\mathfrak{g})\) be a Casimir element in the sense of Definition 14.8, and for \(\varepsilon \neq 0\) let \(C_{\varepsilon} = \lambda(\varepsilon)\,\hat{U}(\varepsilon)^{-1}C\), where \(\hat{U}(\varepsilon)\) is the isomorphism of enveloping algebras induced by \(U(\varepsilon)\) and \(\lambda(\varepsilon) \neq 0\) is a scalar. If the coefficients of \(C_{\varepsilon}\) in the ordered monomials of Theorem 14.7 converge as \(\varepsilon\to0\), their limit defines an element \(C_{0} \in U(\mathfrak{g}_{0})\), and \(C_{0}\) is a Casimir element of \(\mathfrak{g}_{0}\). Rests on Definition 14.8, Theorem 14.7 and Definition 14.83.

Proof.

Derives Lemma 14.87. Because \(U(\varepsilon)\) is an isomorphism of Lie algebras from \(\left(\mathfrak{g},\comm{\cdot}{\cdot}_{\varepsilon}\right)\) to \(\mathfrak{g}\), it extends to an isomorphism \(\hat{U}(\varepsilon)\) of the corresponding enveloping algebras, which carries the centre to the centre. Hence \(\comm{C_{\varepsilon}}{T_{\alpha}}_{\varepsilon} = 0\) for every \(\alpha\) and every \(\varepsilon \neq 0\), the scalar \(\lambda(\varepsilon)\) being irrelevant to that.

Reduce both sides to the ordered basis of Theorem 14.7. Every reordering step uses the relations Equation (14.12) for the bracket \(\comm{\cdot}{\cdot}_{\varepsilon}\), so the coefficients of \(\comm{C_{\varepsilon}}{T_{\alpha}}_{\varepsilon}\) in that basis are polynomials in the coefficients of \(C_{\varepsilon}\) and in the structure constants \(C^{\gamma}{}_{\alpha\beta}(\varepsilon)\), with the same polynomials for every \(\varepsilon\) — the degree of \(C_{\varepsilon}\) being fixed. Both arguments converge, the polynomials are continuous, and the value is identically zero; therefore \(\comm{C_{0}}{T_{\alpha}}_{0} = 0\) for every \(\alpha\), which is Equation (14.13) for \(\mathfrak{g}_{0}\).

Remark 14.88 (The Galilei mass, precisely).

Lemma 14.87 is the mechanism by which an invariant of the parent algebra reappears as a central quantity of the contracted one, and the nonrelativistic limit is its standard instance — but the instance has to be stated carefully, because the plain contraction of the Poincaré algebra does not produce a central charge. Take \(D = 4\) in the ordering of Notation 14.1, so that \(x^{4} = ct\) and the generator of time translation is \(H = c\,P_{4} = \pp_{t}\), of SI unit \(\mathrm{s}^{-1}\), and put \(G_{i} = J_{i4}/c\) for the rescaled boost, which tends to the Galilean \(t\,\pp_{i}\). Then Equations (14.90) and (14.97) give

\begin{equation}\tag{14.127} \comm{G_{i}}{P_{j}} = -\frac{\delta_{ij}}{c^{2}}\,H\ec\qquad \comm{G_{i}}{H} = -P_{i}\ec\qquad \comm{G_{i}}{G_{j}} = \frac{1}{c^{2}}J_{ij}\ec \end{equation}

and the limit \(c \rightarrow \infty\) is the Galilei algebra with \(\comm{G_{i}}{P_{j}} = 0\) and no central charge at all. The mass appears when the same limit is taken in the trivially extended algebra \(\mathfrak{iso}(3,1)\oplus\R M\), with \(M\) central, after the \(c\)-dependent redefinition \(H' = H - c^{2}M\): the first bracket of Equation (14.127) then reads \(\comm{G_{i}}{P_{j}} = -\delta_{ij}\left(H'/c^{2} + M\right)\) and tends to \(-\delta_{ij}M\). The extension that survives is nontrivial, which is the content of Proposition 25.14; this treatise proves that nontriviality directly, from the Poisson brackets of a free particle, rather than through the limit. Remark 25.15 is to be read in the sense set out here.

Example 14.89 (Poincaré from de~Sitter, Galilei from Poincaré).

Two contractions are used throughout this treatise.

  1. Vanishing curvature. Rescaling the translation generators of \(\mathfrak{so}(D-1,2)\) or \(\mathfrak{so}(D,1)\) by the inverse radius and letting the radius diverge contracts either de Sitter algebra to the Poincaré algebra Equation (14.99). This is the statement that spacetime looks flat on scales small compared with the curvature radius set by the cosmological constant.

  2. Vanishing velocity ratio. Rescaling the boosts of the Poincaré algebra by \(1/c\) and letting \(c\to\infty\) contracts it to the Galilei algebra. The Casimir \(p_{\mu}p^{\mu}=m^{2}c^{2}\) that labelled a Poincaré representation becomes, in the limit, the central charge of Example 14.80.

The classification of kinematical algebras

A kinematical algebra is one generated by the transformations any spacetime symmetry must contain: rotations \(J_{i}\), boosts \(K_{i}\), space translations \(P_{i}\), and a time translation \(H\)—ten generators in four spacetime dimensions. Bacry and Lévy-Leblond asked which Lie algebras these can form, and answered the question completely [Bacry:1968].

Theorem 14.90 (Bacry–Lévy-Leblond).

Assume that

  1. space is isotropic, so that the \(J_{i}\) generate an \(\mathfrak{so}(3)\) under which \(K_{i}\) and \(P_{i}\) are vectors and \(H\) a scalar;

  2. parity and time reversal are automorphisms of the algebra; and

  3. each boost generates a noncompact one-parameter subgroup.

Then there are exactly eleven such algebras, up to a rescaling of the generators, and they are organized by three constants \((\gamma,\mu,\nu)\) appearing in

\begin{equation}\tag{14.128} \boxed{\comm{K_{i}}{P_{j}}=\gamma\,\delta_{ij}H\ec\qquad \comm{K_{i}}{H}=\mu\,P_{i}\ec\qquad \comm{P_{i}}{H}=\nu\,K_{i}}\ec \end{equation}

the two remaining brackets being fixed by the Jacobi identity to

\begin{equation}\tag{14.129} \comm{K_{i}}{K_{j}}=-\gamma\mu\,\epsilon_{ijk}J_{k}\ec\qquad \comm{P_{i}}{P_{j}}=\gamma\nu\,\epsilon_{ijk}J_{k}\ep \end{equation}

Setting each of \(\gamma,\mu,\nu\) to zero is an Inönü–Wigner contraction, so the eight possible vanishing patterns form a cube of contractions; the eleven algebras are those eight cases, three of which split into two according to a sign. Rests on Equation (14.34) and Definition 14.83.

Derives Theorem 14.90.

Remark 14.91 (Eleven kinematics, not eleven isomorphism classes).

The theorem must not be stated as “eleven algebras up to isomorphism”. Exchanging \(K_{i}\leftrightarrow P_{i}\) is an isomorphism of Lie algebras carrying Poincaré to para-Poincaré and Galilei to para-Galilei (Corollary A.60), but it is not an equivalence of kinematics, because it does not commute with the time reversal of hypothesis (2) and because it exchanges the physical roles of a boost and a translation. Two entries of Table 14.1 are therefore abstractly the same algebra as another entry while describing different physics.

AlgebraSym.$(\gamma,\mu,\nu)$Regime it describesRealized in nature?
de SitterdS$(\ast,\ast,\ast)$relativistic, positive curvatureyes, asymptotically: the measured $\Lambda$ is positive
anti-de SitterAdS$(\ast,\ast,\ast)$relativistic, negative curvatureno: the measured $\Lambda$ has the opposite sign
PoincaréP$(\ast,\ast,0)$relativistic, flatyes: the symmetry of Part IV, tested directly
GalileiG$(0,\ast,0)$nonrelativistic, flatyes, as the $c\to\infty$ limit used throughout Part III
Newton–HookeNH$_{+}$$(0,\ast,\ast)$nonrelativistic, positive curvatureonly as an effective description of $\Lambda$ in local dynamics
anti-Newton–HookeNH$_{-}$$(0,\ast,\ast)$nonrelativistic, negative curvatureno, for the same reason as AdS
CarrollC$(\ast,0,0)$ultrarelativistic, $c\to0$, flatno laboratory regime; it is the intrinsic symmetry of a null hypersurface
para-PoincaréP$'$$(\ast,0,\ast)$ultrarelativistic, curvedno
para-GalileiG$'$$(0,0,\ast)$mixed limitno
inhomogeneous $\SO(4)$E$'$$(\ast,0,\ast)$Euclidean signaturenot a spacetime kinematics at all
staticS$(0,0,0)$no relative motion whateverdegenerate limit only
The eleven kinematical algebras of Theorem 14.90, what each describes, and whether this treatise records a measurement in its regime. The third column gives the vanishing pattern of the three constants of Equation (14.128), $\ast$ marking a nonzero one; the eight patterns are the vertices of the contraction cube, and the three cases carrying two entries differ only by a sign. The last column is a verdict about the evidence base, not about the mathematics — all eleven are consistent Lie algebras — and the evidence for the three realized cases is cited in Remark 14.92.
Remark 14.92 (What the table is claiming).

Three of the eleven are realized in the strong sense that measurements have been made in their regime and agree with them. Poincaré is the subject of Part IV and is tested to the precision recorded in Experiments: Light, the Aether, and Time and Experiment: Time Dilation and Relativistic Kinematics. Galilei is its low-velocity limit, which is what Part III uses throughout. De Sitter qualifies in a weaker and more careful sense: the measured cosmological constant is positive — the supernova evidence for an accelerating expansion [Riess:1998] [Perlmutter:1999], with the value fixed by the cosmic microwave background [Aghanim:2020] — so the asymptotic symmetry of the observed universe is de Sitter rather than anti-de Sitter or Poincaré — but this is an inference from the expansion history of Evidence-Based Cosmology, not a symmetry test. The remaining eight are mathematically consistent and, at present, physically unoccupied.

The Carroll algebra is the interesting borderline case. It has no laboratory regime, being the limit \(c\to0\), yet it is the intrinsic symmetry of a null hypersurface, and null hypersurfaces are observed — the horizons of Experiment: Black-Hole Observations. That is a statement about where an algebra appears in a description, which is not the same as a measurement of the algebra, and the table's last column is deliberately worded to keep the two apart.

Remark 14.93 (The cube is an algebra, not only a picture).

The eight vanishing patterns of \((\gamma,\mu,\nu)\) were called the vertices of a cube of contractions, which so far is a way of drawing the theorem rather than a statement about it. It is more than that. Section 15.7 of the next chapter shows that the four generator species \(J_{i}\), \(K_{i}\), \(H\), \(P_{i}\) — the boosts are written \(K_{i}\) here and in Theorem 14.90, and \(G_{i}\) there — are the four components of a \(\Z_{2}\times\Z_{2}\)-grading of the algebra, that the three constants \((\gamma,\mu,\nu)\) are the three structure constants of a single four-dimensional commutative associative algebra \(A\) carrying the same grading, and that the eight vertices are the eight ways those three constants can vanish — so each edge of the cube is a degeneration of \(A\) and nothing else. The three rescaling-invariant signs of that algebra are the squares of its three graded generators; hypothesis (3) above, that the boosts be noncompact, is exactly the condition that one of those three squares be nonnegative; and counting the vertices against the signs that survive it returns \(8+3=11\). The two accounts are complementary rather than redundant: this chapter's, completed in The Classification of Kinematical Algebras, proves that the list is exhaustive, which no construction can do; the next chapter's says what the list is — one algebra and its degenerations.

Remark 14.94 (Extensions of the classification, and the scope rule).

Theorem 14.90 is not the end of the subject. Dropping the discrete automorphisms, or working in general dimension, enlarges the list; extending the algebras—by further generators or by central charges—produces infinite families, and the construction of gauge theories from them is an active programme. Supersymmetric extensions are the other standard enlargement, and this treatise names them only to say what its scope rule requires: no supersymmetric partner of any known particle has been observed, so supersymmetry has no evidence base and nothing built on it belongs in this book. A recent example builds Maxwell extensions of the kinematical algebras by a semigroup expansion method that generalizes contraction, and uses the resulting invariant tensors to construct three-dimensional Chern–Simons gravity theories [Concha:2026]. These constructions are mathematics, and this treatise treats them as such: gravity in three spacetime dimensions has no propagating degrees of freedom and no observational status, and the paper claims none. The Chern–Simons machinery they use is developed in Section 14.8.1 of this chapter, so the mathematical apparatus is present in the book; what is absent, and is stated here rather than left implicit, is any measurement that selects such an extension. Section 133.7 records the general policy.

Fibre bundles

The invariant-polynomial machinery of the following sections is formulated on a principal fibre bundle, and so is every gauge theory in Part XI. This section supplies the definitions and proves the one identity those constructions consume, the Bianchi identity. Nothing beyond that is developed: the classification of bundles, characteristic numbers as topological invariants, and holonomy are not needed anywhere in this treatise and are not treated.

Bundles, sections and principal bundles

Definition 14.95 (Fibre bundle).

A fibre bundle consists of differentiable manifolds \(E\) (the total space), \(M\) (the base) and \(F\) (the typical fibre), a Lie group \(G\) (the structure group) acting effectively on \(F\) on the left, and a smooth surjection \(\pi : E \longrightarrow M\) (the projection), such that \(M\) carries an open cover \(\set{U_{i}}\) and there are diffeomorphisms

\begin{equation}\tag{14.130} \varphi_{i} : \pi^{-1}(U_{i}) \longrightarrow U_{i}\times F\ec \qquad \operatorname{pr}_{1}\circ\varphi_{i} = \pi\ec \end{equation}

called local trivializations, whose overlaps act on the fibre through \(G\):

\begin{equation}\tag{14.131} \varphi_{i}\circ\varphi_{j}^{-1}(x,f) = \left(x,t_{ij}(x)f\right)\ec \qquad t_{ij} : U_{i}\cap U_{j} \longrightarrow G \end{equation}

smooth. The \(t_{ij}\) are the transition functions; composing Equation (14.131) on triple overlaps gives the cocycle conditions

\begin{equation}\tag{14.132} t_{ii} = e\ec\qquad t_{ij}t_{ji} = e\ec\qquad t_{ij}t_{jk} = t_{ik}\ep \end{equation}

The bundle is trivial if a single trivialization with \(U_{1} = M\) exists, in which case \(E \cong M\times F\).

Definition 14.96 (Section).

A section of \(E\) over an open \(U \subseteq M\) is a smooth map \(s : U \longrightarrow E\) with \(\pi\circ s = \id_{U}\); it is global if \(U = M\). Local sections always exist — take \(s(x) = \varphi_{i}^{-1}(x,f_{0})\) inside a trivializing neighbourhood — and a global section need not. Rests on Definition 14.95.

Definition 14.97 (Principal bundle).

A principal \(G\)-bundle \(P(M,G)\) is a fibre bundle whose typical fibre is the group \(G\) itself, acting on itself by left translation, together with the right action of \(G\) on \(P\) defined in each trivialization by

\begin{equation}\tag{14.133} \varphi_{i}(p) = \left(x,h\right) \quad\longmapsto\quad \varphi_{i}(p\cdot g) = \left(x,hg\right)\ep \end{equation}

The action is well defined — right translation commutes with the left translation by \(t_{ij}\) of Equation (14.131), so Equation (14.133) does not depend on the trivialization — it is free, and its orbits are exactly the fibres of \(\pi\). Rests on Definition 14.95.

Proposition 14.98 (A principal bundle is trivial if and only if it has a global section).

\(P(M,G)\) is trivial if and only if it admits a global section. Rests on Definitions 14.96 and 14.97.

Proof.

Derives Proposition 14.98. If \(P \cong M\times G\), then \(x \mapsto (x,e)\) is a global section. Conversely let \(s\) be one and define

\begin{equation*} \Phi : M\times G \longrightarrow P\ec\qquad \Phi(x,g) = s(x)\cdot g\ep \end{equation*}

It is smooth, satisfies \(\pi\circ\Phi = \operatorname{pr}_{1}\), and is bijective: surjective because \(G\) acts transitively on each fibre, so every \(p\) with \(\pi(p) = x\) is \(s(x)g\) for some \(g\); injective because the action is free. Its inverse is smooth, as one sees in any local trivialization \(\varphi_{i}\): writing \(\varphi_{i}(s(x)) = (x,\sigma_{i}(x))\) with \(\sigma_{i}\) smooth, \(\varphi_{i}\circ\Phi(x,g) = (x,\sigma_{i}(x)g)\), whose inverse \((x,h) \mapsto \left(x,\sigma_{i}(x)^{-1}h\right)\) is smooth because inversion and multiplication in \(G\) are. Hence \(\Phi\) is a diffeomorphism commuting with the right action, and \(P\) is trivial.

Definition 14.99 (Associated bundle).

Let \(P(M,G)\) be principal and let \(G\) act on a manifold \(F\) on the left. The associated bundle \(P\times_{G}F\) is the quotient of \(P\times F\) by the equivalence

\begin{equation}\tag{14.134} \left(p,f\right) \sim \left(p\cdot g,\ g^{-1}f\right)\ec \qquad g \in G\ec \end{equation}

with projection \([p,f]\mapsto\pi(p)\). It is a fibre bundle with typical fibre \(F\) and the same transition functions Equation (14.131). When \(F\) is a vector space and the action is a linear representation, \(P\times_{G}F\) is a vector bundle, and a matter field is a section of it. Rests on Definitions 14.95 and 14.97.

Remark 14.100 (Where the nontriviality is felt).

The definitions above would be idle if every bundle were trivial, since a trivial bundle reduces every construction to a global choice of gauge. Two places in this treatise turn on their not being. A gauge-fixing prescription, such as the Faddeev–Popov construction of Section 106.4, is a section of a principal bundle over the space of gauge orbits, and Proposition 14.98 is the statement quoted in support of Theorem 106.65: for a non-abelian structure group that bundle is not trivial, so it admits no global section — no global section, no global gauge fixing, and the Faddeev–Popov prescription is a local construction used globally [Singer:1978]. And the connection of Definition 14.103 below is what makes the phase acquired around a closed path a property of the path rather than of the coordinates, the effect predicted in [Aharonov:1959], first measured by Chambers with an iron whisker behind an electron biprism [Chambers:1960] and put beyond the leakage-field objection by Tonomura and collaborators, who enclosed the flux in a superconducting shield [Tonomura:1986]. The non-abelian case, in which the structure group is not \(\U(1)\), was introduced for the isotopic-spin symmetry [Yang:1954].

Connection and curvature

Notation 14.101 (Algebra-valued forms and their bracket).

A \(\mathfrak{g}\)-valued \(p\)-form on a manifold is \(\vect{\alpha} = \alpha^{a}\vect{T}_{a}\) with \(\alpha^{a}\) ordinary \(p\)-forms and \(\set{\vect{T}_{a}}\) a basis of \(\mathfrak{g}\). For a \(p\)-form \(\vect{\alpha}\) and a \(q\)-form \(\vect{\beta}\) put

\begin{equation}\tag{14.135} \comm{\vect{\alpha}}{\vect{\beta}} = \alpha^{a}\wedge\beta^{b}\comm{\vect{T}_{a}}{\vect{T}_{b}}\ec \end{equation}

a \(\mathfrak{g}\)-valued \((p+q)\)-form, and \(\dd\vect{\alpha} = \dd\alpha^{a}\,\vect{T}_{a}\). Expanding the wedge product of two \(1\)-forms on a pair of vectors gives

\begin{equation}\tag{14.136} \comm{\vect{\alpha}}{\vect{\beta}}(V,W) = \comm{\vect{\alpha}(V)}{\vect{\beta}(W)} - \comm{\vect{\alpha}(W)}{\vect{\beta}(V)}\ep \end{equation}
Lemma 14.102 (Two identities for a $\mathfrak{g}$-valued $1$-form).

For every \(\mathfrak{g}\)-valued \(1\)-form \(\vect{\alpha}\),

\begin{equation}\tag{14.137} \dd\comm{\vect{\alpha}}{\vect{\alpha}} = 2\comm{\dd\vect{\alpha}}{\vect{\alpha}}\ec\qquad \comm{\vect{\alpha}}{\comm{\vect{\alpha}}{\vect{\alpha}}} = 0\ep \end{equation}

Rests on Equations (14.5) and (14.135).

Proof.

Derives Lemma 14.102. For the first, the graded Leibniz rule gives \(\dd\left(\alpha^{a}\wedge\alpha^{b}\right) = \dd\alpha^{a}\wedge\alpha^{b} - \alpha^{a}\wedge\dd\alpha^{b}\), and in the second term a \(1\)-form and a \(2\)-form commute, so \(-\alpha^{a}\wedge\dd\alpha^{b}\comm{\vect{T}_{a}}{\vect{T}_{b}} = -\dd\alpha^{b}\wedge\alpha^{a}\comm{\vect{T}_{a}}{\vect{T}_{b}} = +\dd\alpha^{a}\wedge\alpha^{b}\comm{\vect{T}_{a}}{\vect{T}_{b}}\) after renaming \(a \leftrightarrow b\) and using the antisymmetry of the bracket; the two terms are equal. For the second,

\begin{equation*} \comm{\vect{\alpha}}{\comm{\vect{\alpha}}{\vect{\alpha}}} = \alpha^{c}\wedge\alpha^{a}\wedge\alpha^{b}\, C^{d}{}_{ab}C^{e}{}_{cd}\,\vect{T}_{e}\ec \end{equation*}

and the \(3\)-form \(\alpha^{c}\wedge\alpha^{a}\wedge\alpha^{b}\) is totally antisymmetric in \((c,a,b)\), so only the totally antisymmetric part of the coefficient survives. That part is the cyclic sum \(C^{d}{}_{ab}C^{e}{}_{cd} + C^{d}{}_{bc}C^{e}{}_{ad} + C^{d}{}_{ca}C^{e}{}_{bd}\), which vanishes: it is the Jacobi identity for \(\mathfrak{g}\) written in the structure constants of Equation (14.5).

Definition 14.103 (Connection $1$-form).

Let \(P(M,G)\) be a principal bundle. For \(X \in \mathfrak{g}\) the fundamental vector field \(X^{\sharp}\) is the generator of the flow \(p \mapsto p\cdot\ee^{tX}\); it is tangent to the fibres, and since the action is free and transitive on each fibre, \(X \mapsto X^{\sharp}_{p}\) is a linear isomorphism of \(\mathfrak{g}\) onto the vertical subspace \(V_{p} = \ker\dd\pi_{p}\). A connection on \(P\) is a \(\mathfrak{g}\)-valued \(1\)-form \(\vect{A}\) on \(P\) with

\begin{equation}\tag{14.138} \boxed{\vect{A}\left(X^{\sharp}\right) = X \quad \text{for all } X \in \mathfrak{g}\ec\qquad R_{g}^{\ast}\vect{A} = g^{-1}\vect{A}g \quad \text{for all } g \in G}\ec \end{equation}

where \(R_{g}\) is the right action Equation (14.133) and the second condition is written for a matrix group, the adjoint action in general. The horizontal subspace is \(H_{p} = \ker\vect{A}_{p}\). Rests on Definition 14.97 and Notation 14.101.

Proposition 14.104 (A connection splits the tangent space).

\(T_{p}P = V_{p}\oplus H_{p}\) for every \(p \in P\), and the projections onto the two summands are \(v \mapsto \left(\vect{A}(v)\right)^{\sharp}\) and \(v \mapsto hv = v - \left(\vect{A}(v)\right)^{\sharp}\). Rests on Definition 14.103.

Proof.

Derives Proposition 14.104. For \(v \in T_{p}P\) write \(v = \left(\vect{A}(v)\right)^{\sharp} + hv\) with \(hv\) defined by that equation. The first term is vertical by construction, and \(\vect{A}(hv) = \vect{A}(v) - \vect{A}\left(\left(\vect{A}(v)\right)^{\sharp}\right) = \vect{A}(v) - \vect{A}(v) = 0\) by the first condition of Equation (14.138), so \(hv \in H_{p}\). The sum is direct: a vector in \(V_{p}\cap H_{p}\) is \(X^{\sharp}\) for some \(X\), and \(0 = \vect{A}(X^{\sharp}) = X\) forces it to vanish.

Definition 14.105 (Curvature $2$-form).

The curvature of the connection \(\vect{A}\) is the \(\mathfrak{g}\)-valued \(2\)-form

\begin{equation}\tag{14.139} \boxed{\vect{F} = \dd\vect{A} + \frac{1}{2}\comm{\vect{A}}{\vect{A}}}\ec \end{equation}

and the covariant exterior derivative of a \(\mathfrak{g}\)-valued form is \(\mathrm{D} = \dd + \comm{\vect{A}}{\ \cdot\ }\), the upright \(\mathrm{D}\) being distinguished from the dimension \(D\) as in Notation 14.1. Rests on Definition 14.103 and Notation 14.101.

Proposition 14.106 (The curvature is horizontal, and measures non-integrability).

For all \(v,w \in T_{p}P\),

\begin{equation}\tag{14.140} \vect{F}(v,w) = \dd\vect{A}\left(hv,hw\right)\ep \end{equation}

In particular \(\vect{F}\) vanishes whenever either argument is vertical, and for horizontal vector fields \(V,W\),

\begin{equation}\tag{14.141} \vect{F}(V,W) = -\vect{A}\left(\comm{V}{W}\right)\ec \end{equation}

so that \(\vect{F} = 0\) if and only if the horizontal distribution is closed under the Lie bracket. Rests on Definition 14.105 and Proposition 14.104.

Proof.

Derives Proposition 14.106. By Equation (14.136), \(\tfrac{1}{2}\comm{\vect{A}}{\vect{A}}(v,w) = \comm{\vect{A}(v)}{\vect{A}(w)}\). By bilinearity it is enough to check Equation (14.140) when each argument is either horizontal or vertical, using Proposition 14.104, and the invariant formula \(\dd\vect{A}(V,W) = V\vect{A}(W) - W\vect{A}(V) - \vect{A}\left(\comm{V}{W}\right)\) for extensions to vector fields. That formula is the coordinate identity \(\dd\alpha(V,W) = \left(\pp_{\mu}\alpha_{\nu} - \pp_{\nu}\alpha_{\mu}\right)V^{\mu}W^{\nu}\) of Section 13.6.4, rearranged: the derivatives of the components of \(V\) and \(W\) cancel against the bracket term.

Both horizontal. Then \(\vect{A}(v) = \vect{A}(w) = 0\), the bracket term vanishes and \(hv = v\), \(hw = w\), so both sides are \(\dd\vect{A}(v,w)\); extending to horizontal fields and using the invariant formula gives Equation (14.141).

Both vertical, \(v = X^{\sharp}\), \(w = Y^{\sharp}\). The bracket term is \(\comm{X}{Y}\). For the fundamental fields \(\comm{X^{\sharp}}{Y^{\sharp}} = \comm{X}{Y}^{\sharp}\), because \(X \mapsto X^{\sharp}\) is the differential of a right action, and \(\vect{A}(X^{\sharp}) = X\) is constant along \(P\), so the invariant formula gives \(\dd\vect{A}(X^{\sharp},Y^{\sharp}) = -\comm{X}{Y}\). The two cancel, and the right-hand side of Equation (14.140) vanishes because \(hv = hw = 0\).

Mixed, \(v\) horizontal and \(w = Y^{\sharp}\). The bracket term vanishes, since \(\vect{A}(v) = 0\). Extend \(v\) to the horizontal lift \(V\) of a vector field on \(M\), which is invariant under every \(R_{g}\) by the second condition of Equation (14.138); then \(\comm{Y^{\sharp}}{V} = 0\), being the derivative along the flow of \(Y^{\sharp}\) — that is, along the action of \(\ee^{tY}\) — of an invariant field. The invariant formula leaves \(V\vect{A}(Y^{\sharp}) = V(Y) = 0\). Both sides vanish, the right-hand one because \(hw = 0\).

Theorem 14.107 (Bianchi identity).

Every connection satisfies

\begin{equation}\tag{14.142} \boxed{\mathrm{D}\vect{F} = \dd\vect{F} + \comm{\vect{A}}{\vect{F}} = 0}\ep \end{equation}

Rests on Definition 14.105 and Lemma 14.102.

Proof.

Derives Theorem 14.107. Differentiate Equation (14.139). Since the exterior derivative squares to zero, and by the first identity of Equation (14.137),

\begin{equation*} \dd\vect{F} = \tfrac{1}{2}\,\dd\comm{\vect{A}}{\vect{A}} = \comm{\dd\vect{A}}{\vect{A}}\ep \end{equation*}

On the other hand, by the second identity of Equation (14.137),

\begin{equation*} \comm{\vect{A}}{\vect{F}} = \comm{\vect{A}}{\dd\vect{A}} + \tfrac{1}{2}\comm{\vect{A}}{\comm{\vect{A}}{\vect{A}}} = \comm{\vect{A}}{\dd\vect{A}}\ep \end{equation*}

Finally \(\comm{\vect{A}}{\dd\vect{A}} = A^{a}\wedge\dd A^{b}\comm{\vect{T}_{a}}{\vect{T}_{b}} = \dd A^{b}\wedge A^{a}\comm{\vect{T}_{a}}{\vect{T}_{b}} = -\comm{\dd\vect{A}}{\vect{A}}\), moving the \(1\)-form past the \(2\)-form at no cost in sign and then renaming \(a \leftrightarrow b\). The two contributions cancel.

Remark 14.108 (The same objects on the base).

A local section \(s : U \longrightarrow P\) (Definition 14.96) pulls the connection and its curvature back to \(\mathfrak{g}\)-valued forms \(s^{\ast}\vect{A}\) and \(s^{\ast}\vect{F}\) on the patch \(U \subseteq M\); the first is what a physicist calls the gauge potential and the second the field strength. Pullback commutes with \(\dd\) and with the pointwise bracket Equation (14.135), so Equations (14.139) and (14.142) hold verbatim for the pulled-back forms. They depend on the section chosen, which is why a gauge potential is not a tensor on \(M\), and by Proposition 14.98 a single section covering all of \(M\) exists only when the bundle is trivial. The sections that follow work on \(P\), where no choice has to be made.

Invariant polynomials

Consider a principal bundle \(P(M,G)\), where \(M\) is a \(D\)-dimensional differentiable manifold and \(G\) a Lie group. We denote the Lie algebra associated with \(G\) by \(\mathfrak{g} = \operatorname{span}\set{\vect{T}_{A}}\), with \(A = 1, \ldots, \dim\mathfrak{g}\). The possible linear combinations of the \(\vect{T}_{A}\) form a vector space, which, when no ambiguity is possible, we shall also denote by \(\mathfrak{g}\). Let, moreover, \([\ ] : G \rightarrow \GL(\mathfrak{g})\) be the linear representation of the Lie group \(G\) defined by \([g] : \vect{B} \mapsto [g]\vect{B} = g\vect{B}g^{-1}\), with \(g \in G\) and \(\vect{B} = B^{A}\vect{T}_{A} \in \mathfrak{g}\). Finally, let \(S_{r}(\mathfrak{g},\R)\) be the space of real symmetric \(r\)-linear functionals, and let \(P \in S_{r}(\mathfrak{g},\R)\), i.e. a mapping of the form

\begin{equation}\tag{14.143} P : \underbrace{\mathfrak{g}\otimes\cdots\otimes\mathfrak{g}}_{r\ \text{times}} \longrightarrow \R \end{equation}

satisfying the following axioms:

  1. the mapping is an \(r\)-linear transformation, i.e.

    \begin{equation}\tag{14.144} P(\vect{B}_{1},\ldots,a\vect{B}+b\vect{C},\ldots,\vect{B}_{r}) = a\,P(\vect{B}_{1},\ldots,\vect{B},\ldots,\vect{B}_{r}) + b\,P(\vect{B}_{1},\ldots,\vect{C},\ldots,\vect{B}_{r})\ec \end{equation}
  2. the mapping is totally symmetric:

    \begin{equation}\tag{14.145} P(\vect{B}_{1},\ldots,\vect{B}_{i},\ldots,\vect{B}_{j},\ldots,\vect{B}_{r}) = P(\vect{B}_{1},\ldots,\vect{B}_{j},\ldots,\vect{B}_{i},\ldots,\vect{B}_{r}) \ec\quad \forall\, i, j\ep \end{equation}

Of all these totally symmetric polynomials, we shall concentrate on those that are invariant under the action induced by the representation \([\ ]\) of the Lie group \(G\), which is enshrined in the following definition.

Definition 14.109 ($G$-invariant polynomial).

We say that \(P \in S_{r}(\mathfrak{g},\R)\) is a \(G\)-invariant polynomial if

\begin{equation*} \forall\, g \in G,\quad \forall\, \vect{B}_{l} \in \mathfrak{g},\ l = 1, \ldots, r,\quad P\left([g]\vect{B}_{1},\ldots,[g]\vect{B}_{r}\right) = P\left(\vect{B}_{1},\ldots,\vect{B}_{r}\right)\ec \end{equation*}

where \([g]\vect{B} = g\vect{B}g^{-1}\). We introduce, moreover, the notation \(P\left(\vect{B}_{1},\ldots,\vect{B}_{r}\right) = \gen{\vect{B}_{1},\ldots,\vect{B}_{r}}\). Throughout this chapter the angle brackets denote this multilinear functional and nothing else; a linear span is written \(\operatorname{span}\). The brackets appear once before this definition, in the proof of Proposition 14.54, and there they already carry the meaning fixed here. Rests on Equations (14.144) and (14.145).

The polynomial extends to \(\mathfrak{g}\)-valued differential forms by multilinearity: if \(\vect{B}_{l} = B_{l}^{A_{l}}\vect{T}_{A_{l}}\) with \(B_{l}^{A_{l}}\) ordinary \(p_{l}\)-forms, then \(\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} = B_{1}^{A_{1}}\cdots B_{r}^{A_{r}} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{r}}}\), where the product of the component forms is the wedge product.

Theorem 14.110.

Let \(\gen{\vect{B}_{1},\ldots,\vect{B}_{r}}\) be a \(G\)-invariant polynomial evaluated on \(\mathfrak{g}\)-valued differential forms of degrees \(p_{1},\ldots,p_{r}\) defined on the principal bundle \(P(M,G)\).

  1. One has

    \begin{align} \dd\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} &= \sum_{l=1}^{r}(-1)^{p_{1}+\cdots+p_{l-1}} \gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\dd\vect{B}_{l}, \vect{B}_{l+1},\ldots,\vect{B}_{r}}\ec\tag{14.146}\\ \mathrm{D}\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} &= \sum_{l=1}^{r}(-1)^{p_{1}+\cdots+p_{l-1}} \gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\mathrm{D}\vect{B}_{l}, \vect{B}_{l+1},\ldots,\vect{B}_{r}}\ec\tag{14.147} \end{align}

    where \(\dd\) is the exterior derivative and \(\mathrm{D} = \dd + \comm{\vect{A}}{\ \cdot\ }\) is the covariant exterior derivative, with \(\vect{A}\) a connection defined on \(P(M,G)\).

  2. Let \(\set{\vect{T}_{A}}\) be a basis of the Lie algebra \(\mathfrak{g}\) associated with the Lie group \(G\). One has

    \begin{equation}\tag{14.148} \sum_{l=1}^{r} C_{AA_{l}}{}^{B} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{l-1}},\vect{T}_{B}, \vect{T}_{A_{l+1}},\ldots,\vect{T}_{A_{r}}} = 0\ep \end{equation}
  3. Let, moreover, \(\vect{B} = B^{A}\vect{T}_{A}\) be a \(\mathfrak{g}\)-valued \(p\)-form. Then

    \begin{equation}\tag{14.149} \sum_{l=1}^{r}(-1)^{p(p_{1}+\cdots+p_{l-1})} \gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\comm{\vect{B}}{\vect{B}_{l}}, \vect{B}_{l+1},\ldots,\vect{B}_{r}} = 0\ep \end{equation}
  4. One has

    \begin{equation}\tag{14.150} \mathrm{D}\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} = \dd\gen{\vect{B}_{1},\ldots,\vect{B}_{r}}\ep \end{equation}

Rests on Definition 14.109 and Equation (14.5).

Proof.

Derives Theorem 14.110. We prove the four statements in turn.

  1. Expanding \(\vect{B}_{1},\ldots,\vect{B}_{r}\) in the basis \(\set{\vect{T}_{A}}\) of the Lie algebra \(\mathfrak{g}\) of the group \(G\) and using the linearity of the polynomial, we have \(\dd\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} = \dd\left(B_{1}^{A_{1}}\cdots B_{r}^{A_{r}}\right) \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{r}}}\). By the graded Leibniz rule for the exterior derivative,

    \begin{align*} \dd\left(B_{1}^{A_{1}}\cdots B_{r}^{A_{r}}\right) &= \dd B_{1}^{A_{1}}\,B_{2}^{A_{2}}\cdots B_{r}^{A_{r}} + (-1)^{p_{1}}B_{1}^{A_{1}}\,\dd B_{2}^{A_{2}}\,B_{3}^{A_{3}}\cdots B_{r}^{A_{r}}\\ &\quad+ \cdots + (-1)^{p_{1}+\cdots+p_{r-1}}B_{1}^{A_{1}}\cdots B_{r-1}^{A_{r-1}}\, \dd B_{r}^{A_{r}}\\ &= \sum_{l=1}^{r}(-1)^{p_{1}+\cdots+p_{l-1}} B_{1}^{A_{1}}\cdots B_{l-1}^{A_{l-1}}\,\dd B_{l}^{A_{l}}\, B_{l+1}^{A_{l+1}}\cdots B_{r}^{A_{r}}\ec \end{align*}

    and hence

    \begin{align*} &\dd\gen{\vect{B}_{1},\ldots,\vect{B}_{r}}\\ &= \sum_{l=1}^{r}(-1)^{p_{1}+\cdots+p_{l-1}} B_{1}^{A_{1}}\cdots B_{l-1}^{A_{l-1}}\,\dd B_{l}^{A_{l}}\, B_{l+1}^{A_{l+1}}\cdots B_{r}^{A_{r}} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{r}}}\\ &= \sum_{l=1}^{r}(-1)^{p_{1}+\cdots+p_{l-1}} \gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\dd\vect{B}_{l}, \vect{B}_{l+1},\ldots,\vect{B}_{r}}\ep \end{align*}

    The expression Equation (14.147) for the covariant exterior derivative is proved by an entirely analogous procedure, yielding the claim.

  2. From the definition of a \(G\)-invariant polynomial we have

    \begin{equation}\tag{14.151} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{r}}} = \gen{[g]\vect{T}_{A_{1}},\ldots,[g]\vect{T}_{A_{r}}} = \gen{g\vect{T}_{A_{1}}g^{-1},\ldots,g\vect{T}_{A_{r}}g^{-1}}\ec \end{equation}

    where \(g = \ee^{t^{A}\vect{T}_{A}}\) is an element of the Lie group \(G\). Furthermore, since the generators \(\vect{T}_{A}\) do not depend on the parameters of the infinitesimal transformations, neither does the polynomial \(\gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{r}}}\); therefore \(\left.\frac{\dd}{\dd t^{A}} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{r}}}\right|_{t^{A}=0} = 0\), and from Equation (14.151) we arrive at

    \begin{equation}\tag{14.152} \left.\frac{\dd}{\dd t^{A}} \gen{g\vect{T}_{A_{1}}g^{-1},\ldots,g\vect{T}_{A_{r}}g^{-1}} \right|_{t^{A}=0} = 0\ep \end{equation}

    From part 1 — analogously to the procedure used to prove Equation (14.146) — one obtains

    \begin{align} &\frac{\dd}{\dd t^{A}} \gen{g\vect{T}_{A_{1}}g^{-1},\ldots,g\vect{T}_{A_{r}}g^{-1}}\notag\\ &= \sum_{l=1}^{r} \gen{g\vect{T}_{A_{1}}g^{-1},\ldots,g\vect{T}_{A_{l-1}}g^{-1}, \frac{\dd}{\dd t^{A}}\left(g\vect{T}_{A_{l}}g^{-1}\right), g\vect{T}_{A_{l+1}}g^{-1},\ldots,g\vect{T}_{A_{r}}g^{-1}}\ec \tag{14.153} \end{align}

    but \(\left.g\vect{T}_{A}g^{-1}\right|_{t^{A}=0} = \left.\ee^{t^{B}\vect{T}_{B}}\vect{T}_{A}\, \ee^{-t^{C}\vect{T}_{C}}\right|_{t^{A}=0} = \vect{T}_{A}\), and moreover

    \begin{align*} \left.\frac{\dd}{\dd t^{A}} \left(g\vect{T}_{B}g^{-1}\right)\right|_{t^{A}=0} &= \left.\left(\frac{\dd}{\dd t^{A}} \left(\ee^{t^{C}\vect{T}_{C}}\right)\vect{T}_{B}\, \ee^{-t^{E}\vect{T}_{E}} + \ee^{t^{C}\vect{T}_{C}}\vect{T}_{B}\, \frac{\dd}{\dd t^{A}}\left(\ee^{-t^{E}\vect{T}_{E}}\right) \right)\right|_{t^{A}=0}\\ &= \left.\left(\ee^{t^{C}\vect{T}_{C}}\vect{T}_{A}\vect{T}_{B}\, \ee^{-t^{E}\vect{T}_{E}} - \ee^{t^{C}\vect{T}_{C}}\vect{T}_{B}\, \ee^{-t^{E}\vect{T}_{E}}\vect{T}_{A}\right)\right|_{t^{A}=0}\\ &= \comm{\vect{T}_{A}}{\vect{T}_{B}}\\ &= C_{AB}{}^{C}\,\vect{T}_{C}\ec \end{align*}

    so that we finally conclude from Equation (14.152) and Equation (14.153) that the claim holds.

  3. In the basis \(\set{\vect{T}_{A}}\) we have

    \begin{align*} &\gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\comm{\vect{B}}{\vect{B}_{l}}, \vect{B}_{l+1},\ldots,\vect{B}_{r}}\\ &= B_{1}^{A_{1}}\cdots B_{l-1}^{A_{l-1}}\,B^{A}B_{l}^{A_{l}}\, B_{l+1}^{A_{l+1}}\cdots B_{r}^{A_{r}} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{l-1}}, \comm{\vect{T}_{A}}{\vect{T}_{A_{l}}}, \vect{T}_{A_{l+1}},\ldots,\vect{T}_{A_{r}}}\\ &= (-1)^{p(p_{1}+\cdots+p_{l-1})} B^{A}B_{1}^{A_{1}}\cdots B_{l-1}^{A_{l-1}}B_{l}^{A_{l}} B_{l+1}^{A_{l+1}}\cdots B_{r}^{A_{r}}\\ &\quad\times \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{l-1}}, C_{AA_{l}}{}^{B}\vect{T}_{B}, \vect{T}_{A_{l+1}},\ldots,\vect{T}_{A_{r}}}\ec \end{align*}

    where the sign arises from commuting the \(p\)-form components \(B^{A}\) past the forms \(B_{1}^{A_{1}},\ldots,B_{l-1}^{A_{l-1}}\) of total degree \(p_{1}+\cdots+p_{l-1}\). Hence

    \begin{align*} &(-1)^{p(p_{1}+\cdots+p_{l-1})} \gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\comm{\vect{B}}{\vect{B}_{l}}, \vect{B}_{l+1},\ldots,\vect{B}_{r}}\\ &= B^{A}B_{1}^{A_{1}}\cdots B_{r}^{A_{r}}\,C_{AA_{l}}{}^{B} \gen{\vect{T}_{A_{1}},\ldots,\vect{T}_{A_{l-1}},\vect{T}_{B}, \vect{T}_{A_{l+1}},\ldots,\vect{T}_{A_{r}}}\ec \end{align*}

    so that, summing over all values of \(l = 1, \ldots, r\), part 2 yields the desired result.

  4. Since \(\mathrm{D} = \dd + \comm{\vect{A}}{\ \cdot\ }\), from part 1 we see that

    \begin{equation*} \mathrm{D}\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} = \dd\gen{\vect{B}_{1},\ldots,\vect{B}_{r}} + \sum_{l=1}^{r}(-1)^{p_{1}+\cdots+p_{l-1}} \gen{\vect{B}_{1},\ldots,\vect{B}_{l-1},\comm{\vect{A}}{\vect{B}_{l}}, \vect{B}_{l+1},\ldots,\vect{B}_{r}}\ep \end{equation*}

    From part 3 with \(\vect{B} = \vect{A}\) and \(p = 1\) — since the connection is a \(1\)-form — the second term vanishes.

The Chern–Weil theorem and transgression forms

Theorem 14.111 (Chern–Weil).

Let \(P(M,G)\) be a principal bundle equipped with a connection \(1\)-form \(\vect{A}\) of curvature \(\vect{F} = \dd\vect{A} + \frac{1}{2}\comm{\vect{A}}{\vect{A}}\), and let \(P\) be a \(G\)-invariant polynomial of degree \(r\). Then:

  1. \(P(\vect{F},\ldots,\vect{F}) = \gen{\vect{F},\ldots,\vect{F}}\) is a closed form, i.e. \(\dd\gen{\vect{F},\ldots,\vect{F}} = 0\).

  2. If \(\vect{F}\) and \(\bar{\vect{F}}\) are the curvature \(2\)-forms corresponding to the connections \(\vect{A}\) and \(\bar{\vect{A}}\) respectively, both defined on \(P(M,G)\), then \(\gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}}\) is an exact form, i.e.

    \begin{equation}\tag{14.154} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} = \dd Q^{(2r-1)}(\vect{A},\bar{\vect{A}})\ec \end{equation}

    where the \((2r-1)\)-form \(Q^{(2r-1)}(\vect{A},\bar{\vect{A}})\) is called the transgression form of the polynomial \(P\) and is explicitly given by

    \begin{equation}\tag{14.155} Q^{(2r-1)}(\vect{A},\bar{\vect{A}}) = r\int_{0}^{1} \gen{\vect{A}-\bar{\vect{A}},\vect{F}_{t},\ldots,\vect{F}_{t}} \dd t\ec \end{equation}

    where \(\vect{F}_{t} = \dd\vect{A}_{t} + \frac{1}{2}\comm{\vect{A}_{t}}{\vect{A}_{t}}\) is the interpolating field strength corresponding to the interpolating connection \(\vect{A}_{t}\) defined as

    \begin{equation}\tag{14.156} \vect{A}_{t}(\vect{A},\bar{\vect{A}}) = \bar{\vect{A}} + t\left(\vect{A}-\bar{\vect{A}}\right)\ec \qquad 0 \leq t \leq 1\ep \end{equation}

Rests on Theorem 14.110, Equation (14.139) and Equation (14.142).

Proof.

Derives Theorem 14.111. Theorem 14.110 is the principal tool.

  1. Consider part 1 of Theorem 14.110, in its covariant form Equation (14.147), for \(\vect{B}_{1},\ldots,\vect{B}_{r} = \vect{F}\), so that \(p_{1},\ldots,p_{r} = 2\). We obtain

    \begin{align} \mathrm{D}\gen{\vect{F},\ldots,\vect{F}} &= \sum_{l=1}^{r}(-1)^{2+\cdots+2} \gen{\vect{F},\ldots,\vect{F},\mathrm{D}\vect{F}, \vect{F},\ldots,\vect{F}}\notag\\ &= \sum_{l=1}^{r} \gen{\vect{F},\ldots,\vect{F},\mathrm{D}\vect{F}, \vect{F},\ldots,\vect{F}}\notag\\ &= r\gen{\vect{F},\ldots,\vect{F},\mathrm{D}\vect{F}, \vect{F},\ldots,\vect{F}}\ep\tag{14.157} \end{align}

    From the Bianchi identity we have \(\mathrm{D}\vect{F} = \dd\vect{F} + \comm{\vect{A}}{\vect{F}} = 0\), and \(\gen{\vect{F},\ldots,\vect{F},0,\vect{F},\ldots,\vect{F}} = 0\) as a direct consequence of multilinearity:

    \begin{align*} \gen{\vect{F},\ldots,\vect{F},0,\vect{F},\ldots,\vect{F}} &= \gen{\vect{F},\ldots,\vect{F},\vect{B}-\vect{B}, \vect{F},\ldots,\vect{F}}\\ &= \gen{\vect{F},\ldots,\vect{F},\vect{B},\vect{F},\ldots,\vect{F}} + \gen{\vect{F},\ldots,\vect{F},-\vect{B},\vect{F},\ldots,\vect{F}}\\ &= \gen{\vect{F},\ldots,\vect{F},\vect{B},\vect{F},\ldots,\vect{F}} - \gen{\vect{F},\ldots,\vect{F},\vect{B},\vect{F},\ldots,\vect{F}}\\ &= 0\ep \end{align*}

    Together with part 4 of Theorem 14.110, which gives \(\dd\gen{\vect{F},\ldots,\vect{F}} = \mathrm{D}\gen{\vect{F},\ldots,\vect{F}} = 0\), this proves the statement.

  2. Write \(\vect{\theta} = \vect{A} - \bar{\vect{A}}\), so that \(\vect{A}_{0} = \bar{\vect{A}}\) and \(\vect{A}_{1} = \bar{\vect{A}} + \vect{\theta} = \vect{A}\), and denote by \(\bar{\mathrm{D}} = \dd + \comm{\bar{\vect{A}}}{\ \cdot\ }\) the covariant exterior derivative of the connection \(\bar{\vect{A}}\). The field strength corresponding to \(\vect{A}_{t}\) is given by \(\vect{F}_{t} = \dd\vect{A}_{t} + \frac{1}{2}\comm{\vect{A}_{t}}{\vect{A}_{t}}\), so that \(\vect{F}_{0} = \bar{\vect{F}}\) and \(\vect{F}_{1} = \vect{F}\). We have

    \begin{equation*} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} = \gen{\vect{F}_{1},\ldots,\vect{F}_{1}} - \gen{\vect{F}_{0},\ldots,\vect{F}_{0}} = \int_{0}^{1}\frac{\dd}{\dd t} \gen{\vect{F}_{t},\ldots,\vect{F}_{t}}\,\dd t\ep \end{equation*}

    Analogously to the procedure used to obtain Equation (14.146), one finds

    \begin{equation*} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} = \int_{0}^{1}\sum_{l=1}^{r} \gen{\vect{F}_{t},\ldots,\vect{F}_{t}, \frac{\dd\vect{F}_{t}}{\dd t}, \vect{F}_{t},\ldots,\vect{F}_{t}}\,\dd t\ep \end{equation*}

    Since the summand does not depend on the index \(l\), and since a \(G\)-invariant polynomial is by definition totally symmetric, we obtain \(r\) equal terms, arriving at

    \begin{equation*} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} = r\int_{0}^{1} \gen{\frac{\dd\vect{F}_{t}}{\dd t}, \vect{F}_{t},\ldots,\vect{F}_{t}}\,\dd t\ep \end{equation*}

    To compute the derivative of the interpolating field strength \(\vect{F}_{t}\) with respect to \(t\), observe that

    \begin{align*} \vect{F}_{t} &= \dd\vect{A}_{t} + \frac{1}{2}\comm{\vect{A}_{t}}{\vect{A}_{t}}\\ &= \dd\bar{\vect{A}} + t\,\dd\vect{\theta} + \frac{1}{2}\comm{\bar{\vect{A}}+t\vect{\theta}} {\bar{\vect{A}}+t\vect{\theta}}\\ &= \dd\bar{\vect{A}} + t\,\dd\vect{\theta} + \frac{1}{2}\comm{\bar{\vect{A}}}{\bar{\vect{A}}} + \frac{t}{2}\comm{\bar{\vect{A}}}{\vect{\theta}} + \frac{t}{2}\comm{\vect{\theta}}{\bar{\vect{A}}} + \frac{t^{2}}{2}\comm{\vect{\theta}}{\vect{\theta}}\\ &= \dd\bar{\vect{A}} + \frac{1}{2}\comm{\bar{\vect{A}}}{\bar{\vect{A}}} + t\,\dd\vect{\theta} + t\comm{\bar{\vect{A}}}{\vect{\theta}} + \frac{t^{2}}{2}\comm{\vect{\theta}}{\vect{\theta}}\ec \end{align*}

    that is,

    \begin{equation}\tag{14.158} \vect{F}_{t} = \bar{\vect{F}} + t\,\bar{\mathrm{D}}\vect{\theta} + \frac{t^{2}}{2}\comm{\vect{\theta}}{\vect{\theta}}\ec \end{equation}

    so that \(\frac{\dd\vect{F}_{t}}{\dd t} = \bar{\mathrm{D}}\vect{\theta} + t\comm{\vect{\theta}}{\vect{\theta}}\), and thus

    \begin{equation*} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} = r\int_{0}^{1} \gen{\vect{F}_{t},\ldots,\vect{F}_{t}, \bar{\mathrm{D}}\vect{\theta} + t\comm{\vect{\theta}}{\vect{\theta}}, \vect{F}_{t},\ldots,\vect{F}_{t}}\,\dd t\ep \end{equation*}

    Since the polynomial is totally symmetric we may write

    \begin{equation}\tag{14.159} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} = r\int_{0}^{1} \gen{\bar{\mathrm{D}}\vect{\theta} + t\comm{\vect{\theta}}{\vect{\theta}}, \vect{F}_{t},\ldots,\vect{F}_{t}}\,\dd t\ep \end{equation}

    On the other hand, applying part 1 of Theorem 14.110 — for the connection \(\bar{\vect{A}}\) — with \(\vect{B}_{1} = \vect{\theta}\) and \(\vect{B}_{2},\ldots,\vect{B}_{r} = \vect{F}_{t}\), so that \(p_{1} = 1\) and \(p_{2},\ldots,p_{r} = 2\), together with part 4, we observe that

    \begin{align} \dd\gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}} &= \bar{\mathrm{D}} \gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}}\notag\\ &= \gen{\bar{\mathrm{D}}\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}} + \sum_{l=2}^{r}(-1)^{1+2+\cdots+2} \gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}, \bar{\mathrm{D}}\vect{F}_{t}, \vect{F}_{t},\ldots,\vect{F}_{t}}\notag\\ &= \gen{\bar{\mathrm{D}}\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}} - (r-1)\gen{\vect{\theta},\bar{\mathrm{D}}\vect{F}_{t}, \vect{F}_{t},\ldots,\vect{F}_{t}}\ep\tag{14.160} \end{align}

    Recall that the difference of two connections is a tensor, and that a connection plus a tensor defines a new connection. In this way \(\vect{A}_{t}\) determines a connection on the bundle, to which corresponds a covariant derivative \(\mathrm{D}_{\vect{A}_{t}} = \dd + \comm{\vect{A}_{t}}{\ \cdot\ }\). For the field strength \(\vect{F}_{t}\) we have the Bianchi identity \(\mathrm{D}_{\vect{A}_{t}}\vect{F}_{t} = 0\), which implies

    \begin{align*} \dd\vect{F}_{t} + \comm{\vect{A}_{t}}{\vect{F}_{t}} &= 0\ec\\ \dd\vect{F}_{t} + \comm{\bar{\vect{A}}+t\vect{\theta}}{\vect{F}_{t}} &= 0\ec\\ \dd\vect{F}_{t} + \comm{\bar{\vect{A}}}{\vect{F}_{t}} + t\comm{\vect{\theta}}{\vect{F}_{t}} &= 0\ec\\ \implies \bar{\mathrm{D}}\vect{F}_{t} &= -t\comm{\vect{\theta}}{\vect{F}_{t}}\ec \end{align*}

    so that

    \begin{equation}\tag{14.161} \dd\gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}} = \gen{\bar{\mathrm{D}}\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}} + t(r-1)\gen{\vect{\theta},\comm{\vect{\theta}}{\vect{F}_{t}}, \vect{F}_{t},\ldots,\vect{F}_{t}}\ep \end{equation}

    Finally, applying part 3 of Theorem 14.110 with \(\vect{B} = \vect{\theta}\), \(\vect{B}_{1} = \vect{\theta}\) and \(\vect{B}_{2},\ldots,\vect{B}_{r} = \vect{F}_{t}\), so that \(p = 1\), \(p_{1} = 1\) and \(p_{2},\ldots,p_{r} = 2\), we obtain

    \begin{align*} \gen{\comm{\vect{\theta}}{\vect{\theta}}, \vect{F}_{t},\ldots,\vect{F}_{t}} + \sum_{l=2}^{r}(-1)^{1\cdot(1+2+\cdots+2)} \gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}, \comm{\vect{\theta}}{\vect{F}_{t}}, \vect{F}_{t},\ldots,\vect{F}_{t}} &= 0\ec\\ \implies \gen{\comm{\vect{\theta}}{\vect{\theta}}, \vect{F}_{t},\ldots,\vect{F}_{t}} - (r-1)\gen{\vect{\theta},\comm{\vect{\theta}}{\vect{F}_{t}}, \vect{F}_{t},\ldots,\vect{F}_{t}} &= 0\ec \end{align*}

    from which we see that Equation (14.161) becomes \(\dd\gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}} = \gen{\bar{\mathrm{D}}\vect{\theta} + t\comm{\vect{\theta}}{\vect{\theta}}, \vect{F}_{t},\ldots,\vect{F}_{t}}\). Substituting this result into Equation (14.159), we obtain

    \begin{align*} \gen{\vect{F},\ldots,\vect{F}} - \gen{\bar{\vect{F}},\ldots,\bar{\vect{F}}} &= r\int_{0}^{1} \dd\gen{\vect{\theta},\vect{F}_{t},\ldots,\vect{F}_{t}}\,\dd t\\ &= \dd\left(r\int_{0}^{1} \gen{\vect{A}-\bar{\vect{A}},\vect{F}_{t},\ldots,\vect{F}_{t}}\, \dd t\right)\\ &= \dd Q^{(2r-1)}(\vect{A},\bar{\vect{A}})\ep \end{align*}

Chern–Simons forms

Definition 14.112 (Chern–Simons form).

Let \(P(\vect{F}) = \gen{\vect{F},\ldots,\vect{F}}\) be a \(G\)-invariant polynomial of degree \(r\). By the Chern–Weil theorem we know that \(\gen{\vect{F},\ldots,\vect{F}}\) is a closed form, and it can therefore be written locally as \(\gen{\vect{F},\ldots,\vect{F}} = \dd Q^{(2r-1)}(\vect{A})\). The \((2r-1)\)-form \(Q^{(2r-1)}(\vect{A})\) is called the Chern–Simons form corresponding to the polynomial \(P\). Rests on Theorem 14.111 and Definition 14.109.

By the Chern–Weil theorem we have

\begin{equation*} \dd Q^{(2r-1)}(\vect{A}) - \dd Q^{(2r-1)}(\bar{\vect{A}}) = \dd\left(r\int_{0}^{1}\dd t \gen{\vect{A}-\bar{\vect{A}},\vect{F}_{t},\ldots,\vect{F}_{t}}\right)\ec \end{equation*}

so that for \(\bar{\vect{A}} = 0\) one has \(Q^{(2r-1)}(\bar{\vect{A}}) = 0\) and thus

\begin{equation}\tag{14.162} Q^{(2r-1)}(\vect{A}) = r\int_{0}^{1}\dd t \gen{\vect{A},\vect{F}_{t},\ldots,\vect{F}_{t}}\ep \end{equation}

However, an important warning must be made: since the coordinate transformation law of a connection is inhomogeneous, it is impossible to choose a connection that vanishes globally. Note, moreover, that locally a transgression form with \(\bar{\vect{A}} = 0\) corresponds to a Chern–Simons form (which is, of course, defined locally):

\begin{equation}\tag{14.163} Q^{(2r-1)}(\vect{A},0) = Q^{(2r-1)}(\vect{A})\ep \end{equation}

When \(\bar{\vect{A}} = 0\), the interpolating connection is given by \(\vect{A}_{t} = t\vect{A}\), with \(0 \leq t \leq 1\), so that \(\vect{A}_{0} = 0\) and \(\vect{A}_{1} = \vect{A}\). The interpolating field strength corresponding to \(\vect{A}_{t}\) is given by

\begin{align*} \vect{F}_{t} &= \dd\vect{A}_{t} + \frac{1}{2}\comm{\vect{A}_{t}}{\vect{A}_{t}}\\ &= t\,\dd\vect{A} + \frac{t^{2}}{2}\comm{\vect{A}}{\vect{A}}\\ &= t\left(\vect{F} - \frac{1}{2}\comm{\vect{A}}{\vect{A}}\right) + \frac{t^{2}}{2}\comm{\vect{A}}{\vect{A}}\\ &= t\vect{F} + \frac{1}{2}t(t-1)\comm{\vect{A}}{\vect{A}}\ep \end{align*}

From Equation (14.155) we then find that the sought Chern–Simons form is

\begin{equation}\tag{14.164} Q^{(2r-1)}(\vect{A}) = r\int_{0}^{1}\dd t\, t^{r-1} \gen{\vect{A}, \vect{F}+\frac{1}{2}(t-1)\comm{\vect{A}}{\vect{A}}, \ldots, \vect{F}+\frac{1}{2}(t-1)\comm{\vect{A}}{\vect{A}}}\ep \end{equation}

The source manuscript stops at Equation (14.164). The three things it was evidently going to do next — evaluate the Chern–Simons forms in the degrees that occur, name the invariant polynomials whose Chern–Weil forms are the characteristic classes, and apply the construction to an \(\SO(p,q)\) connection — are done here, and the chapter closes with them.

Proposition 14.113 (The Chern–Simons forms of degree $3$ and $5$).

For \(r = 2\) and \(r = 3\), Equation (14.164) evaluates to

\begin{align} Q^{(3)}(\vect{A}) &= \gen{\vect{A},\vect{F}} - \tfrac{1}{6}\gen{\vect{A},\comm{\vect{A}}{\vect{A}}}\ec \tag{14.165}\\ Q^{(5)}(\vect{A}) &= \gen{\vect{A},\vect{F},\vect{F}} - \tfrac{1}{4}\gen{\vect{A},\vect{F},\comm{\vect{A}}{\vect{A}}} + \tfrac{1}{40}\gen{\vect{A},\comm{\vect{A}}{\vect{A}}, \comm{\vect{A}}{\vect{A}}}\ep \tag{14.166} \end{align}

When the invariant polynomial is the trace of a product of matrices in some representation, \(\gen{\vect{X},\vect{Y}} = \tr(\vect{X}\vect{Y})\), Equation (14.165) takes the familiar form

\begin{equation}\tag{14.167} \boxed{Q^{(3)}(\vect{A}) = \tr\left(\vect{A}\wedge\dd\vect{A} + \tfrac{2}{3}\,\vect{A}\wedge\vect{A}\wedge\vect{A}\right)}\ep \end{equation}

Rests on Equation (14.164), Definition 14.112 and Definition 14.109.

Proof.

Derives Proposition 14.113. Abbreviate \(\vect{G}_{t} = \vect{F} + \tfrac{1}{2}(t-1)\comm{\vect{A}}{\vect{A}}\), which is the entry appearing in every slot but the first of Equation (14.164). For \(r = 2\) there is one such slot, and linearity gives

\begin{equation*} Q^{(3)} = 2\int_{0}^{1}\dd t\,t\,\gen{\vect{A},\vect{G}_{t}} = 2\gen{\vect{A},\vect{F}}\int_{0}^{1}t\,\dd t + \gen{\vect{A},\comm{\vect{A}}{\vect{A}}} \int_{0}^{1}t(t-1)\,\dd t\ec \end{equation*}

and the two integrals are \(1/2\) and \(1/3-1/2 = -1/6\), which is Equation (14.165). For \(r = 3\) there are two such slots, and by total symmetry

\begin{equation*} \gen{\vect{A},\vect{G}_{t},\vect{G}_{t}} = \gen{\vect{A},\vect{F},\vect{F}} + (t-1)\gen{\vect{A},\vect{F},\comm{\vect{A}}{\vect{A}}} + \tfrac{1}{4}(t-1)^{2} \gen{\vect{A},\comm{\vect{A}}{\vect{A}},\comm{\vect{A}}{\vect{A}}}\ec \end{equation*}

so that the three coefficients of Equation (14.166) are \(3\int_{0}^{1}t^{2}\dd t = 1\), \(3\int_{0}^{1}t^{2}(t-1)\dd t = 3\left(\tfrac{1}{4}-\tfrac{1}{3}\right) = -\tfrac{1}{4}\) and \(\tfrac{3}{4}\int_{0}^{1}t^{2}(t-1)^{2}\dd t = \tfrac{3}{4}\left(\tfrac{1}{5}-\tfrac{1}{2}+\tfrac{1}{3}\right) = \tfrac{1}{40}\).

For Equation (14.167), matrix-valued \(1\)-forms obey \(\comm{\vect{A}}{\vect{A}} = 2\,\vect{A}\wedge\vect{A}\) by Equation (14.135), so \(\gen{\vect{A},\vect{F}} = \tr\left(\vect{A}\wedge\dd\vect{A}\right) + \tr\left(\vect{A}\wedge\vect{A}\wedge\vect{A}\right)\) and \(\tfrac{1}{6}\gen{\vect{A},\comm{\vect{A}}{\vect{A}}} = \tfrac{1}{3}\tr\left(\vect{A}\wedge\vect{A}\wedge\vect{A}\right)\).

Characteristic classes

The Chern–Weil forms of Theorem 14.111 live on the total space \(P\). They are forms on the base as well, and that is what makes them invariants of \(M\) rather than of a choice of gauge.

Proposition 14.114 (Chern–Weil forms descend to the base).

Let \(\vect{A}\) be a connection on \(P(M,G)\) with curvature \(\vect{F}\), and let \(\gen{\ \cdot\ }\) be a \(G\)-invariant polynomial of degree \(r\). Then, for a matrix group,

\begin{equation}\tag{14.168} R_{g}^{\ast}\vect{F} = g^{-1}\vect{F}g\ec \end{equation}

and \(\gen{\vect{F},\ldots,\vect{F}}\) is the pullback \(\pi^{\ast}\rho\) of a unique closed \(2r\)-form \(\rho\) on \(M\). Rests on Theorem 14.111, Proposition 14.106 and Definition 14.109.

Proof.

Derives Proposition 14.114. Pullback commutes with \(\dd\) and with the bracket Equation (14.135), and \(g\) is a fixed group element, so

\begin{equation*} R_{g}^{\ast}\vect{F} = \dd\left(g^{-1}\vect{A}g\right) + \tfrac{1}{2}\comm{g^{-1}\vect{A}g}{g^{-1}\vect{A}g} = g^{-1}\left(\dd\vect{A} + \tfrac{1}{2}\comm{\vect{A}}{\vect{A}}\right)g\ec \end{equation*}

using Equation (14.138) in the first step. Hence \(R_{g}^{\ast}\gen{\vect{F},\ldots,\vect{F}} = \gen{g^{-1}\vect{F}g,\ldots,g^{-1}\vect{F}g} = \gen{\vect{F},\ldots,\vect{F}}\) by Definition 14.109: the form is invariant. It is also horizontal, being a wedge of the component forms of \(\vect{F}\), each of which vanishes as soon as one argument is vertical by Proposition 14.106.

An invariant horizontal form descends. Define \(\rho_{x}(v_{1},\ldots,v_{2r}) = \gen{\vect{F},\ldots,\vect{F}}_{p} \left(\tilde{v}_{1},\ldots,\tilde{v}_{2r}\right)\) for any \(p \in \pi^{-1}(x)\) and any \(\tilde{v}_{i}\) with \(\dd\pi\,\tilde{v}_{i} = v_{i}\). Two lifts at the same \(p\) differ by vertical vectors, which horizontality kills, so the value does not depend on the choice; two points of the same fibre differ by some \(g\), and invariance makes the two values agree. Smoothness follows by computing \(\rho\) on a local section, \(\rho|_{U} = s^{\ast} \gen{\vect{F},\ldots,\vect{F}}\). Uniqueness holds because \(\dd\pi\) is onto. Finally \(\dd\rho\) pulls back to \(\dd\gen{\vect{F},\ldots,\vect{F}} = 0\) by Theorem 14.111, and \(\pi^{\ast}\) is injective, so \(\rho\) is closed.

Definition 14.115 (Characteristic forms).

Let \(G\) act by matrices in a representation of dimension \(N\), and define the polynomials \(\sigma_{r}\) on \(\mathfrak{g}\) by

\begin{equation}\tag{14.169} \det\left(\identity + \lambda\vect{X}\right) = \sum_{r=0}^{N}\lambda^{r}\,\sigma_{r}(\vect{X})\ec \end{equation}

and, for \(\dim = 2m\) and \(\vect{X}\) antisymmetric, the Pfaffian

\begin{equation}\tag{14.170} \operatorname{Pf}(\vect{X}) = \frac{1}{2^{m}m!}\,\epsilon^{a_{1}\cdots a_{2m}} X_{a_{1}a_{2}}\cdots X_{a_{2m-1}a_{2m}}\ep \end{equation}

By Lemma 14.116 these are \(G\)-invariant polynomials, so Theorem 14.111 and Proposition 14.114 attach to each of them a closed form on \(M\) whose de Rham class does not depend on the connection. The Chern forms are those of \(\sigma_{r}\) for \(G\) unitary, evaluated at \(\vect{X} = \ii\vect{F}/2\pi\); the Pontryagin forms are those of \(\sigma_{r}\) for \(G\) orthogonal, evaluated at \(\vect{X} = \vect{F}/2\pi\), only the even \(r\) surviving; and the Euler form is that of the Pfaffian, for \(G = \SO(2m)\). Rests on Definition 14.109, Theorem 14.111 and Proposition 14.114.

Lemma 14.116 (Invariance of the characteristic coefficients).

The \(\sigma_{r}\) of Equation (14.169) are \(G\)-invariant polynomials of degree \(r\) in the sense of Definition 14.109, and \(\sigma_{r}\) vanishes identically on antisymmetric matrices for odd \(r\). The Pfaffian Equation (14.170) obeys \(\operatorname{Pf}\left(g\vect{X}g\transpose\right) = \det(g)\operatorname{Pf}(\vect{X})\) for orthogonal \(g\), so it is invariant under \(\SO(2m)\) and changes sign under an orientation-reversing orthogonal transformation; the Euler form therefore belongs to an oriented bundle. Rests on Definition 14.109.

Proof.

Derives Lemma 14.116. Invariance is \(\det\left(\identity+\lambda\,g\vect{X}g^{-1}\right) = \det\left(g\left(\identity+\lambda\vect{X}\right)g^{-1}\right) = \det\left(\identity+\lambda\vect{X}\right)\), true coefficient by coefficient in \(\lambda\); each \(\sigma_{r}\) is a homogeneous polynomial of degree \(r\) in the entries of \(\vect{X}\), and the associated symmetric multilinear form is its polarization. For antisymmetric \(\vect{X}\),

\begin{equation*} \det\left(\identity+\lambda\vect{X}\right) = \det\left(\bigl(\identity+\lambda\vect{X}\bigr)\transpose\right) = \det\left(\identity-\lambda\vect{X}\right)\ec \end{equation*}

so the polynomial is even in \(\lambda\) and the odd \(\sigma_{r}\) vanish. For the Pfaffian, an orthogonal \(g\) has \(g\vect{X}g^{-1} = g\vect{X}g\transpose\), each factor in Equation (14.170) carries two of the \(g\)'s, and

\begin{equation*} \epsilon^{b_{1}\cdots b_{2m}}\,g^{a_{1}}{}_{b_{1}}\cdots g^{a_{2m}}{}_{b_{2m}} = \det(g)\,\epsilon^{a_{1}\cdots a_{2m}}\ec \end{equation*}

which is the definition of the determinant. The stated transformation follows, and \(\det g = \pm1\) for an orthogonal matrix.

The transgression of an $\mathfrak{so}(p,q)$ connection

Proposition 14.117 (The quadratic invariant and its Chern–Simons form).

Let \(\vect{\omega} = \tfrac{1}{2}\omega^{AB}J_{AB}\) be a connection whose structure group is \(\SO(p,q)\), taken in the defining representation of Equation (14.23), and let \(\vect{R} = \tfrac{1}{2}R^{AB}J_{AB}\) be its curvature. With the invariant polynomial \(\gen{\vect{X},\vect{Y}} = \tr(\vect{X}\vect{Y})\),

\begin{equation}\tag{14.171} \gen{\vect{R},\vect{R}} = -R^{AB}\wedge R_{AB} = -8\pi^{2}\,\sigma_{2}\left(\vect{R}/2\pi\right)\ec \end{equation}

that is, \(-8\pi^{2}\) times the first Pontryagin form \(\sigma_{2}(\vect{R}/2\pi)\) of Definition 14.115, the invariant \(\gen{\vect{R},\vect{R}}\) itself carrying none of the \(2\pi\) normalization of the characteristic forms. Its Chern–Simons \(3\)-form is

\begin{equation}\tag{14.172} \boxed{Q^{(3)}(\vect{\omega}) = -\omega^{AB}\wedge\dd\omega_{AB} + \tfrac{2}{3}\,\omega^{A}{}_{B}\wedge\omega^{B}{}_{C} \wedge\omega^{C}{}_{A}}\ec \end{equation}

with \(\dd Q^{(3)}(\vect{\omega}) = \gen{\vect{R},\vect{R}}\). Rests on Proposition 14.113, Equation (14.86) and Definition 14.109.

Proof.

Derives Proposition 14.117. In the defining representation the generators are the matrices \(\left(J_{AB}\right)^{M}{}_{N} = \delta^{M}_{A}\eta_{BN} - \delta^{M}_{B}\eta_{AN}\) already used in Proposition 14.22, so

\begin{equation*} \tr\left(J_{AB}J_{CD}\right) = \left(\delta^{M}_{A}\eta_{BN} - \delta^{M}_{B}\eta_{AN}\right) \left(\delta^{N}_{C}\eta_{DM} - \delta^{N}_{D}\eta_{CM}\right) = 2\left(\eta_{AD}\eta_{BC} - \eta_{AC}\eta_{BD}\right)\ec \end{equation*}

which is symmetric under the exchange of the index pairs and invariant, being a trace of a product of generators. Contracting,

\begin{equation*} \gen{\vect{R},\vect{R}} = \tfrac{1}{4}R^{AB}\wedge R^{CD}\,\tr\left(J_{AB}J_{CD}\right) = \tfrac{1}{2}\left(R^{AB}\wedge R_{BA} - R^{AB}\wedge R_{AB}\right) = -R^{AB}\wedge R_{AB}\ec \end{equation*}

the two \(2\)-forms commuting. For the second equality in Equation (14.171): the elements of \(\mathfrak{so}(p,q)\) are traceless, so \(\sigma_{1}(\vect{X}) = \tr\vect{X}\) vanishes, and expanding Equation (14.169) to second order gives \(\sigma_{2}(\vect{X}) = \tfrac{1}{2}\left[(\tr\vect{X})^{2} - \tr(\vect{X}\vect{X})\right] = -\tfrac{1}{2}\tr(\vect{X}\vect{X})\); at \(\vect{X} = \vect{R}/2\pi\), the value prescribed for the Pontryagin forms in Definition 14.115, this is \(-\gen{\vect{R},\vect{R}}/8\pi^{2}\). The same substitution in the matrix components gives \(\left(\vect{\omega}\right)^{M}{}_{N} = \omega^{M}{}_{N}\), so Equation (14.167) reads \(Q^{(3)} = \omega^{M}{}_{N}\wedge\dd\omega^{N}{}_{M} + \tfrac{2}{3}\omega^{M}{}_{N}\wedge\omega^{N}{}_{P} \wedge\omega^{P}{}_{M}\), and lowering the indices in the first term turns \(\omega^{M}{}_{N}\wedge\dd\omega^{N}{}_{M}\) into \(-\omega^{AB}\wedge\dd\omega_{AB}\). That \(\dd Q^{(3)}\) is \(\gen{\vect{R},\vect{R}}\) is Definition 14.112.

Remark 14.118 (Where this ends, and why).

Equation (14.172) is where the source manuscript was heading and is the natural end of the chapter, because it is the last construction this treatise uses. The physical use of the invariant polynomials is made in Part XI, where the obstruction to a classical symmetry surviving quantization is an anomaly (Theorem 106.94); the measured consequence quoted in this chapter is the two-photon decay rate of the neutral pion, with its evidence, in Remark 14.82. The gauge connection of Definition 14.103 is the object every chapter of Part XI works with.

Three directions are deliberately not taken. The classification of bundles over a given base, and the reading of characteristic classes as integers attached to a topology, is algebraic topology and nothing in this book requires it. Index theorems, which relate those integers to the spectra of the operators of Hilbert Spaces, are likewise outside what the evidence under discussion needs. And Chern–Simons actions — as opposed to Chern–Simons forms — are field theories in three spacetime dimensions, which the scope rule of Epistemology and the Scientific Method keeps out of this treatise: Remark 14.94 records the position. The mathematics is here because the anomaly and the gauge principle need it, and stops where they stop.