Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism

Contents
  1. Singular Lagrangians and primary constraints
  2. The Dirac–Bergmann consistency algorithm
  3. First- and second-class constraints
  4. The Dirac bracket
  5. Gauge fixing and the reduced phase space
  6. Worked systems
  7. Quantizing a constrained system
  8. Constraints in field theory

Every construction of Hamiltonian Mechanics rests on one step: that the definition \(p_{a}=\pp\Lag/\pp\dot{q}^{a}\) can be solved for the velocities. When it cannot, the Legendre transformation of Definition 22.2 is not invertible, the momenta are not independent functions on phase space, and the Hamiltonian formalism as stated does not exist. Proposition 22.3 named the condition and Remark 22.4 named the failure; this chapter is what happens next.

This is not a rare degeneracy to be noted and set aside. It is the generic situation for every theory with a gauge symmetry, and therefore for electromagnetism, for the whole Standard Model, and for general relativity. The reason is structural and is worth stating before any formalism: a gauge symmetry means that distinct sets of variables describe the same physical state, so the equations of motion cannot determine the evolution of all the variables—some of them must remain arbitrary—and a Legendre transformation that produced a well-posed evolution for all of them would be a contradiction. Degeneracy of the Lagrangian is the mechanical signature of redundancy in the description.

Dirac's constraint formalism, developed with Bergmann in the years around 1950, is the systematic treatment. It sorts the constraints into those that generate gauge transformations and those that merely restrict phase space, replaces the Poisson bracket by a modified bracket on the physical surface, counts the true degrees of freedom, and hands the result to the quantization rule of Section 25.4. It is stated here, in the classical part, because it is classical mechanics; Parts VII and XI use it rather than redevelop it.

Notation 26.1 (Weak and strong equality).

Indices \(a,b=1,\ldots,n\) label the configuration variables and \(c,d=1,\ldots,2n\) the phase-space variables, as in Notation 22.1. Constraints are written \(\phi_{m}\) for primary constraints, \(\chi_{\alpha}\) for second-class constraints, \(\gamma_{A}\) for first-class constraints, and \(\varphi_{j}\) for the full set. The symbol \(\approx\) denotes weak equality: \(F\approx G\) means that \(F\) and \(G\) agree on the surface defined by the constraints, but not necessarily off it. The distinction is essential and is not pedantry — a weak equality may not be used inside a Poisson bracket, because the bracket differentiates in directions leaving the surface. Ordinary equality \(=\) is called strong when the contrast is being drawn. Rests on Notation 22.1 and Definition 22.29.

Remark 26.2 (The primary literature of this chapter is not yet in the bibliography).

The formalism below is due to Dirac and to Bergmann and his collaborators, in papers of 1949–1951 and in Dirac's 1964 lecture notes; the modern standard reference is the monograph of Henneaux and Teitelboim. None of these works currently has an entry in this treatise's bibliography, so the attributions made in the prose of this chapter are, exceptionally, uncited. They are recorded in the source file as a citation debt. Every statement below is proved here or carries an explicit pending note, so nothing in the chapter rests on an unchecked appeal to those sources; what is missing is the historical credit, not the evidence.

Singular Lagrangians and primary constraints

Definition 26.3 (Singular Lagrangian).

A Lagrangian \(\Lag(q,\dot{q},t)\) is singular, or degenerate, if its Hessian with respect to the velocities

\begin{equation}\tag{26.1} W_{ab}=\frac{\pp^{2}\Lag}{\pp\dot{q}^{a}\,\pp\dot{q}^{b}} \end{equation}

is not invertible. If \(\operatorname{rank}W=R<n\), the Lagrangian is singular with \(M=n-R\) degrees of degeneracy. Rests on Equation (22.2) and Proposition 22.3.

Everything below assumes that \(\operatorname{rank}W\) is constant on the region of \(TQ\) under study. This is an assumption about the system, not a triviality: the rank can drop on a lower-dimensional set, and where it does the constraint structure changes and the analysis must be redone patch by patch. All the systems worked in Section 26.6 have constant rank on the region of interest.

Remark 26.4 (The constant-rank theorem is owed by Part II).

The proof of Proposition 26.5 needs the following statement, which belongs to multivariable analysis and to the manifolds chapter (Real Analysis and Differentiable Manifolds, Tensors, and Curvature) and is at present in neither. Constant-rank theorem. Let \(f:U\subseteq\R^{N}\longrightarrow\R^{N'}\) be \(C^{k}\), \(k\ge1\), with \(\operatorname{rank}Df=r\) at every point of \(U\). Then about each point of \(U\) and its image there are \(C^{k}\) coordinate changes in which \(f\) reads \((u^{1},\ldots,u^{N})\longmapsto(u^{1},\ldots,u^{r},0,\ldots,0)\). Consequently \(f(U)\) is locally an \(r\)-dimensional embedded submanifold of \(\R^{N'}\) (Definition 13.54), cut out near each of its points by \(N'-r\) functions with linearly independent differentials. Theorem 13.59 is the codimension-one, maximal-rank special case of the last sentence, and Theorem A.286 is the analytic input from which the theorem is proved. The statement is used here and nowhere restated.

Proposition 26.5 (Primary constraints).

If \(\Lag\) is singular with \(\operatorname{rank}W=R\) constant, the Legendre map \((q,\dot{q})\mapsto(q,p)\) has image of dimension \(n+R\), so the momenta satisfy \(M=n-R\) independent relations

\begin{equation}\tag{26.2} \phi_{m}(q,p)\approx0\ec\qquad m=1,\ldots,M\ec \end{equation}

identically in the velocities. These are the primary constraints; they follow from the definition of the momenta alone, before any equation of motion is used. Rests on Definition 26.3 and Remark 26.4.

Proof.

Derives Proposition 26.5. Write the Legendre map as \(\mathcal{F}:(q^{a},\dot{q}^{a})\longmapsto \left(q^{a},\,\pp\Lag/\pp\dot{q}^{a}\right)\), a map between spaces of dimension \(2n\). In the block ordering \((q,\dot{q})\) its Jacobian is

\begin{equation}\tag{26.3} D\mathcal{F}=\begin{pmatrix} \identity_{n} & 0\\[2pt] \dfrac{\pp^{2}\Lag}{\pp q^{b}\,\pp\dot{q}^{a}} & W_{ab} \end{pmatrix}\ep \end{equation}

The first block column is already independent, and adding it to the second changes no rank, so \(\operatorname{rank}D\mathcal{F}=n+\operatorname{rank}W=n+R\), constant by hypothesis. By the constant-rank theorem of Remark 26.4 the image is locally an embedded submanifold of the \(2n\)-dimensional space of the \((q,p)\), of dimension \(n+R\), and is cut out near each of its points by \(2n-(n+R)=n-R=M\) functions with independent differentials. Those functions are the \(\phi_{m}\), and since the image is by construction the set of values the momenta actually take, \(\phi_{m}\left(q,\pp\Lag/\pp\dot{q}\right)\) vanishes for every \((q,\dot{q})\): the relations Equation (26.2) hold identically in the velocities, which is what distinguishes them from equations of motion. The \(q\) are untouched by the map, so no relation involves the coordinates alone.

Proposition 26.6 (The canonical Hamiltonian is defined on the constraint surface).

The function

\begin{equation}\tag{26.4} \Ham_{\text{c}}=p_{a}\dot{q}^{a}-\Lag \end{equation}

depends on the velocities only through the combinations that the momenta determine: it is a well-defined function of \((q,p)\) on the primary constraint surface, and is ambiguous off it by terms proportional to the primary constraints. Rests on Proposition 26.5 and Equation (26.4).

Proof.

Derives Proposition 26.6. Vary the right-hand side of Equation (26.4) with respect to the velocities at fixed \(q\) and fixed \(p\):

\begin{equation}\tag{26.5} \delta\Ham_{\text{c}} =p_{a}\,\delta\dot{q}^{a} -\pdv{\Lag}{\dot{q}^{a}}\,\delta\dot{q}^{a} =\left(p_{a}-\pdv{\Lag}{\dot{q}^{a}}\right)\delta\dot{q}^{a}=0\ec \end{equation}

the last step because \(p_{a}=\pp\Lag/\pp\dot{q}^{a}\) is exactly the condition defining the point at which the variation is taken. So \(\Ham_{\text{c}}\) is constant along each fibre of the Legendre map, that is, along each set of velocities carried to the same momentum, and therefore descends to a function on the image — the primary constraint surface of Proposition 26.5. Off that surface it is not defined at all; any extension will do. Two extensions differ by a function vanishing on the surface, hence, by the regularity assumption of Remark 26.7, by a combination \(c^{m}(q,p)\phi_{m}\). That ambiguity is not a defect to be removed but the origin of the undetermined multipliers in Definition 26.8.

Remark 26.7 (The regularity assumption, stated once).

Throughout this chapter the constraint functions are assumed regular: at every point of the surface \(\Sigma\) they cut out, their differentials are linearly independent, and every smooth function vanishing on \(\Sigma\) is a combination \(c^{j}(q,p)\varphi_{j}\) with smooth coefficients. The first half is what makes \(\Sigma\) a submanifold; the second is what lets “\(F\approx0\)” be traded for “\(F=c^{j}\varphi_{j}\)”, which is used in almost every proof below. A constraint set that fails regularity — \(\phi=q^{2}\) rather than \(\phi=q\) is the standard toy — must be replaced by an equivalent regular one before the formalism applies. This is an assumption about the system, and it is not free: it is exactly what fails at the points where \(\operatorname{rank}W\) drops.

Definition 26.8 (Total Hamiltonian).

The total Hamiltonian is

\begin{equation}\tag{26.6} \Ham_{\text{T}}=\Ham_{\text{c}}+u^{m}\phi_{m}\ec \end{equation}

with \(u^{m}\) undetermined multipliers. The evolution of any phase-space function is

\begin{equation}\tag{26.7} \dot{F}\approx\pb{F}{\Ham_{\text{T}}} =\pb{F}{\Ham_{\text{c}}}+u^{m}\pb{F}{\phi_{m}}\ep \end{equation}

Rests on Proposition 26.6 and Definition 22.29.

The Dirac–Bergmann consistency algorithm

The constraints must be preserved by the motion: a system that starts on the constraint surface must stay on it.

Definition 26.9 (Consistency conditions).

The consistency conditions are

\begin{equation}\tag{26.8} \dot{\phi}_{m}=\pb{\phi_{m}}{\Ham_{\text{c}}} +u^{n}\pb{\phi_{m}}{\phi_{n}}\approx0\ec\qquad m=1,\ldots,M\ep \end{equation}

Rests on Equations (26.2) and (26.7).

Proposition 26.10 (The four outcomes).

Each condition Equation (26.8) does exactly one of four things:

  1. it is satisfied identically on the constraint surface, and says nothing;

  2. it reduces to a relation among the \(q\) and \(p\) alone, independent of the multipliers—a new, secondary constraint;

  3. it determines some of the multipliers \(u^{m}\); or

  4. it is inconsistent, in which case the Lagrangian describes no motion at all.

The procedure is then repeated on any new constraint, generating tertiary constraints and so on, until no new condition arises. Rests on Definition 26.9 and Remark 26.7.

Proof.

Derives Proposition 26.10. Abbreviate \(h_{m}:=\pb{\phi_{m}}{\Ham_{\text{c}}}\) and \(C_{mn}:=\pb{\phi_{m}}{\phi_{n}}\), both evaluated on the primary constraint surface. Equation (26.8) is then the linear system

\begin{equation}\tag{26.9} C_{mn}\,u^{n}\approx-h_{m}\ec \end{equation}

in the \(M\) unknowns \(u^{n}\), with an antisymmetric coefficient matrix. Let \(r=\operatorname{rank}C\) on the surface and let \(\set{V^{(k)}_{m}}\), \(k=1,\ldots,M-r\), be a basis of the null vectors of \(C\), so that \(V^{(k)}_{m}C_{mn}\approx0\).

Decompose Equation (26.9) accordingly. Along the row space of \(C\) — \(r\) independent combinations — the system is solvable for \(r\) combinations of the multipliers, and this is outcome 3; the remaining \(M-r\) combinations of the \(u^{n}\) are left arbitrary, which is the point taken up in Theorem 26.14. Contracting Equation (26.9) with a null vector annihilates every multiplier and leaves

\begin{equation}\tag{26.10} V^{(k)}_{m}\,h_{m} =\pb{V^{(k)}_{m}\phi_{m}}{\Ham_{\text{c}}}\approx0\ec \end{equation}

a condition on the phase-space point alone. There are three possibilities for it, and no others. If the left-hand side vanishes on the surface already, the condition is empty: outcome 1. If it is a function that does not vanish there, it is a new restriction on \((q,p)\) — a secondary constraint: outcome 2. If it reduces to a nonzero constant, no phase-space point satisfies it and the theory has no solutions: outcome 4. The exhaustion is complete because Equation (26.9) is a linear system and every linear system either determines an unknown, or produces a compatibility condition, or is empty.

Repeating the argument with the enlarged constraint set gives the recursion. Note that Equation (26.10) identifies the generator of any new constraint as the bracket of a first-class primary combination \(V^{(k)}_{m}\phi_{m}\) with the canonical Hamiltonian, which is how the algorithm is run in practice.

Definition 26.11 (The final constraint set).

The algorithm terminates on a finite-dimensional phase space, since each round either adds an independent constraint—and there are at most \(2n\) of them—or ends. Its output is the full set

\begin{equation}\tag{26.11} \varphi_{j}\approx0\ec\qquad j=1,\ldots,J\ec \end{equation}

comprising the primary constraints and all those generated from them, and the surface \(\Sigma\subset M\) they define is the constraint surface. The distinction primary/secondary refers to how a constraint was found and has no further significance; the classification that matters is the next one. Rests on Proposition 26.10 and Remark 26.7.

First- and second-class constraints

Definition 26.12 (First and second class).

A phase-space function \(F\) is first class if

\begin{equation}\tag{26.12} \pb{F}{\varphi_{j}}\approx0\qquad\text{for every }j\ec \end{equation}

and second class otherwise. The classification applies in particular to the constraints themselves. A complete separation of the constraint set is a choice of basis \(\set{\gamma_{A},\chi_{\alpha}}\) for it in which the \(\gamma_{A}\) are a maximal set of independent first-class combinations and the \(\chi_{\alpha}\) are the rest, so that no nonzero combination of the \(\chi_{\alpha}\) alone is first class. Rests on Definitions 22.29 and 26.11.

Proposition 26.13 (The first-class functions close).

The Poisson bracket of two first-class functions is first class. The first-class constraints therefore form a Lie algebra under the bracket, weakly; its structure functions need not be constants. Rests on Definition 26.12, Equation (22.46) and Remark 26.7.

Proof.

Derives Proposition 26.13. Let \(F\) and \(G\) be first class. By Equation (26.12) and the regularity assumption of Remark 26.7 there are smooth functions with

\begin{equation}\tag{26.13} \pb{F}{\varphi_{j}}=f_{j}{}^{k}\varphi_{k}\ec\qquad \pb{G}{\varphi_{j}}=g_{j}{}^{k}\varphi_{k}\ep \end{equation}

The Jacobi identity Equation (22.46), rearranged, gives

\begin{equation}\tag{26.14} \pb{\pb{F}{G}}{\varphi_{j}} =\pb{F}{\pb{G}{\varphi_{j}}}-\pb{G}{\pb{F}{\varphi_{j}}}\ep \end{equation}

Insert Equation (26.13) and expand each bracket by the Leibniz rule:

\begin{equation}\tag{26.15} \pb{\pb{F}{G}}{\varphi_{j}} =\pb{F}{g_{j}{}^{k}}\varphi_{k}+g_{j}{}^{k}\pb{F}{\varphi_{k}} -\pb{G}{f_{j}{}^{k}}\varphi_{k}-f_{j}{}^{k}\pb{G}{\varphi_{k}}\ep \end{equation}

Every term on the right carries either an explicit \(\varphi_{k}\) or a bracket that vanishes weakly by hypothesis, so the whole expression is weakly zero. Hence \(\pb{F}{G}\) is first class.

Applied to the first-class constraints themselves this says \(\pb{\gamma_{A}}{\gamma_{B}}\approx0\), so again by regularity

\begin{equation}\tag{26.16} \pb{\gamma_{A}}{\gamma_{B}}=f_{AB}{}^{C}\,\gamma_{C}\ec \end{equation}

which is closure. The coefficients \(f_{AB}{}^{C}\) are functions on phase space in general; whether they are constants is a property of the particular theory, and Sections 26.6.4 and 26.6.5 are the two cases that matter — constants for Yang–Mills, genuine functions for gravity. Where they are constants, Equation (26.16) is a Lie algebra in the sense of Theorem 25.3; where they are not, it is not, and the difference is not cosmetic.

Theorem 26.14 (First-class constraints generate gauge transformations).

Let \(\gamma_{k}=V^{(k)}_{m}\phi_{m}\) be the first-class primary constraints, \(V^{(k)}\) the null vectors of Proposition 26.10. Then two solutions of the equations of motion with the same initial data at time \(t\) differ at \(t+\delta t\) by

\begin{equation}\tag{26.17} \delta_{\varepsilon}F=\varepsilon^{k}\pb{F}{\gamma_{k}}\ec \end{equation}

with \(\varepsilon^{k}\) arbitrary. The two therefore describe the same physical history, and the variables they differ in carry no physical information. Rests on Proposition 26.10 and Definition 26.8.

Proof.

Derives Theorem 26.14. By Proposition 26.10 the consistency conditions fix the multipliers only up to the null space of \(C_{mn}\): the general solution of Equation (26.9) is

\begin{equation}\tag{26.18} u^{m}=U^{m}+v^{k}V^{(k)}_{m}\ec \end{equation}

with \(U^{m}\) any particular solution and the \(v^{k}\) arbitrary functions of time. Evolve the same initial point \(z(t)\) through \(\delta t\) with two admissible choices \(u\) and \(u'\). Both results solve the equations of motion, and by Equation (26.7) they differ by

\begin{equation}\tag{26.19} \delta F=\delta t\left(u^{m}-u'^{m}\right)\pb{F}{\phi_{m}} =\delta t\left(v^{k}-v'^{k}\right)V^{(k)}_{m}\pb{F}{\phi_{m}} =\varepsilon^{k}\pb{F}{\gamma_{k}}\ec \end{equation}

writing \(\varepsilon^{k}:=\delta t\,(v^{k}-v'^{k})\), which is arbitrary, and using that the \(V^{(k)}\) are functions of \((q,p)\) so that \(V^{(k)}_{m}\pb{F}{\phi_{m}}=\pb{F}{V^{(k)}_{m}\phi_{m}}\) weakly. That \(\gamma_{k}=V^{(k)}_{m}\phi_{m}\) is first class is Equation (26.10) together with \(V^{(k)}_{m}C_{mn}\approx0\): its bracket with every primary constraint vanishes weakly by the second, and with the canonical Hamiltonian by the first.

Two evolutions of one initial state cannot be two physical histories, on pain of the theory predicting nothing. So the difference Equation (26.19) is unphysical, and any two phase-space points related by Equation (26.17) describe the same state.

Remark 26.15 (Dirac's conjecture, and why the theorem stops where it does).

Theorem 26.14 is proved for the first-class primary constraints, because those are the ones whose multipliers the algorithm demonstrably leaves free. Dirac conjectured that all first-class constraints, secondary ones included, generate gauge transformations. The conjecture is true for every system in Section 26.6, and it is what justifies Definition 26.16; the general construction assembles the gauge generator as a chain \(G=\varepsilon\gamma+\dot{\varepsilon}\gamma'+\cdots\) built from constraints of successive generations, of which Equation (26.61) below is the worked instance. It is nonetheless false in general: counterexamples exist in which a secondary first-class constraint generates no symmetry of the action, and they are not pathological beyond failing the regularity of Remark 26.7 at a point. This treatise does not yet carry a citable reference for either the general construction or the counterexample (Remark 26.2), so the theorem above claims only what is proved here.

Definition 26.16 (Extended Hamiltonian).

The extended Hamiltonian adds all first-class constraints, primary or not, with independent multipliers:

\begin{equation}\tag{26.20} \Ham_{\text{E}}=\Ham_{\text{c}}+v^{A}\gamma_{A}\ep \end{equation}

It generates the same motion of gauge-invariant quantities as \(\Ham_{\text{T}}\), and a larger motion of gauge-variant ones. Rests on Definition 26.8 and Theorem 26.14.

Theorem 26.17 (Counting the physical degrees of freedom).

Let a system have \(2n\) phase-space dimensions, \(F\) independent first-class constraints and \(S\) independent second-class constraints. Then

\begin{equation}\tag{26.21} \dim(\text{physical phase space})=2n-2F-S\ec \end{equation}

and the number of physical configuration-space degrees of freedom is half of that, \(n-F-\tfrac{1}{2}S\). A first-class constraint removes two dimensions—itself, and the gauge direction it generates—while a second-class constraint removes one. Rests on Theorem 26.14, Proposition 26.19 and Remark 26.18.

Proof.

Derives Theorem 26.17. By regularity (Remark 26.7) the \(F+S\) constraints have independent differentials, so the constraint surface \(\Sigma\) of Equation (26.11) is an embedded submanifold (Definition 13.54) of dimension

\begin{equation}\tag{26.22} \dim\Sigma=2n-F-S\ep \end{equation}

Consider on \(\Sigma\) the Hamiltonian vector fields \(X_{A}:=\pb{\cdot}{\gamma_{A}}\) of the first-class constraints. Each is tangent to \(\Sigma\), because \(X_{A}\varphi_{j}=\pb{\varphi_{j}}{\gamma_{A}} =-\pb{\gamma_{A}}{\varphi_{j}}\approx0\) by Equation (26.12), so the flow it generates does not leave the surface; and the \(F\) of them are pointwise independent, because the \(\dd\gamma_{A}\) are independent and the symplectic form is non-degenerate. They therefore span an \(F\)-dimensional distribution on \(\Sigma\). It is involutive: by Proposition 25.4 the bracket of Hamiltonian vector fields is the Hamiltonian vector field of the Poisson bracket, \(\comm{X_{A}}{X_{B}}=X_{\pb{\gamma_{A}}{\gamma_{B}}}\), and \(\pb{\gamma_{A}}{\gamma_{B}}=f_{AB}{}^{C}\gamma_{C}\) by Equation (26.16), whose Hamiltonian vector field on \(\Sigma\) is \(f_{AB}{}^{C}X_{C}\) — again in the span.

By the Frobenius theorem of Remark 26.18 the distribution is integrable, and \(\Sigma\) is foliated by \(F\)-dimensional leaves: the gauge orbits. By Theorem 26.14 all points of one orbit describe one physical state, so the physical phase space is the space of leaves, of dimension

\begin{equation}\tag{26.23} \dim\Sigma-F=2n-F-S-F=2n-2F-S\ec \end{equation}

which is Equation (26.21). It is even, because \(S\) is even by Proposition 26.19, so half of it is an integer; and it is \(2\left(n-F-\tfrac{1}{2}S\right)\), the doubling that a configuration space and its momenta always produce. Finally the reduced space carries a symplectic form: the pullback of \(\omega\) to \(\Sigma\) has the \(X_{A}\) in its kernel and nothing else, so it descends to a non-degenerate form on the quotient — which is Theorem 24.47 in the special case where the momentum map is the constraint set.

Remark 26.18 (The Frobenius theorem is owed by Part II).

The proof of Theorem 26.17 needs the following, which belongs to the manifolds chapter (Differentiable Manifolds, Tensors, and Curvature) and is not there. Frobenius theorem. A smooth distribution \(D\subset TN\) of constant rank \(k\) on a manifold \(N\) is integrable — through each point there passes a \(k\)-dimensional immersed submanifold whose tangent space is \(D\), and \(N\) is foliated by such leaves — if and only if \(D\) is involutive, that is, closed under the Lie bracket of vector fields. The theorem is invoked by name elsewhere in this part as well: the pending note of Theorem 24.38 attributes it to Part II, and Calculus of Variations appeals to it to separate holonomic from nonholonomic constraints. It is stated here because it is used here, not because this is its home.

Proposition 26.19 (Second-class constraints come in pairs).

For a complete separation in the sense of Definition 26.12, the number \(S\) of independent second-class constraints is even, and the matrix

\begin{equation}\tag{26.24} C_{\alpha\beta}=\pb{\chi_{\alpha}}{\chi_{\beta}} \end{equation}

is invertible on the constraint surface. Rests on Definition 26.12 and Proposition 5.119.

Proof.

Derives Proposition 26.19. \(C_{\alpha\beta}\) is antisymmetric, by antisymmetry of the Poisson bracket (Equation (22.43)). Suppose it were degenerate at a point of \(\Sigma\): there would be numbers \(\lambda^{\alpha}\), not all zero, with \(C_{\alpha\beta}\lambda^{\beta}\approx0\). Put \(\chi:=\lambda^{\alpha}\chi_{\alpha}\). Then \(\pb{\chi_{\beta}}{\chi}\approx0\) for every \(\beta\) by construction, and \(\pb{\gamma_{A}}{\chi}\approx0\) for every \(A\) because the \(\gamma_{A}\) are first class. Since \(\set{\gamma_{A},\chi_{\alpha}}\) is a basis of the constraint set, \(\chi\) has weakly vanishing bracket with every constraint: it is first class. But a complete separation admits no nonzero first-class combination of the \(\chi_{\alpha}\) alone, so \(\lambda=0\) — a contradiction. Hence \(C\) is non-degenerate on \(\Sigma\), which is the second assertion.

For the first, \(C\) is then the matrix of a non-degenerate antisymmetric bilinear form on the \(S\)-dimensional real space spanned by the \(\chi_{\alpha}\) at that point. By Proposition 5.119 such a form exists only in even dimension, so \(S=2s\). Equivalently, and this is the one-line version, \(\det C=\det\left(-C\transpose\right) =(-1)^{S}\det C\), so \(\det C=0\) whenever \(S\) is odd.

The Dirac bracket

Second-class constraints cannot be imposed strongly inside a Poisson bracket—setting \(\chi_{\alpha}=0\) in Equation (26.24) would make an invertible matrix vanish. The resolution is to change the bracket.

Definition 26.20 (Dirac bracket).

Let \(\chi_{\alpha}\) be a complete set of second-class constraints and \(\left(C^{-1}\right)^{\alpha\beta}\) the inverse of the matrix Equation (26.24). The Dirac bracket of two phase-space functions is

\begin{equation}\tag{26.25} \pb{A}{B}_{\text{D}} =\pb{A}{B} -\pb{A}{\chi_{\alpha}} \left(C^{-1}\right)^{\alpha\beta} \pb{\chi_{\beta}}{B}\ep \end{equation}

Rests on Proposition 26.19 and Definition 22.29.

Theorem 26.21 (Properties of the Dirac bracket).

The Dirac bracket is bilinear and antisymmetric, obeys the Leibniz rule Equation (24.24) and the Jacobi identity Equation (22.46), and in addition

\begin{align} \pb{A}{\chi_{\alpha}}_{\text{D}}&=0\quad\text{strongly, for every }A \ec\tag{26.26}\\ \pb{A}{B}_{\text{D}}&\approx\pb{A}{B}\quad \text{whenever }A\text{ is first class}\ec \tag{26.27}\\ \dot{F}&\approx\pb{F}{\Ham_{\text{T}}}_{\text{D}}\ep \tag{26.28} \end{align}

Because of Equation (26.26), the second-class constraints may be set to zero strongly once the Dirac bracket is adopted: they become identities, and the phase space is effectively reduced. Rests on Definition 26.20, Proposition 26.19 and Definition 26.8.

Proof.

Derives Theorem 26.21. Bilinearity is immediate: both terms of Equation (26.25) are bilinear in \((A,B)\). For antisymmetry, note first that the inverse of an invertible antisymmetric matrix is antisymmetric, since \(C\transpose=-C\) gives \(\left(C^{-1}\right)\transpose=\left(C\transpose\right)^{-1}=-C^{-1}\). Exchanging \(A\) and \(B\) in the correction term and using \(\pb{B}{\chi_{\alpha}}=-\pb{\chi_{\alpha}}{B}\), \(\pb{\chi_{\beta}}{A}=-\pb{A}{\chi_{\beta}}\) and then relabelling \(\alpha\leftrightarrow\beta\),

\begin{equation}\tag{26.29} -\pb{B}{\chi_{\alpha}}\left(C^{-1}\right)^{\alpha\beta} \pb{\chi_{\beta}}{A} =-\pb{A}{\chi_{\alpha}}\left(C^{-1}\right)^{\beta\alpha} \pb{\chi_{\beta}}{B} =+\pb{A}{\chi_{\alpha}}\left(C^{-1}\right)^{\alpha\beta} \pb{\chi_{\beta}}{B}\ec \end{equation}

so the correction changes sign with the leading term and \(\pb{A}{B}_{\text{D}}=-\pb{B}{A}_{\text{D}}\). The Leibniz rule in the first slot follows because \(A\longmapsto\pb{A}{B}\) and \(A\longmapsto\pb{A}{\chi_{\alpha}}\) are both derivations (Proposition 25.4) and the second factor of the correction does not involve \(A\); antisymmetry then gives it in the second slot.

For Equation (26.26), take \(B=\chi_{\gamma}\):

\begin{equation}\tag{26.30} \pb{A}{\chi_{\gamma}}_{\text{D}} =\pb{A}{\chi_{\gamma}} -\pb{A}{\chi_{\alpha}}\left(C^{-1}\right)^{\alpha\beta}C_{\beta\gamma} =\pb{A}{\chi_{\gamma}}-\pb{A}{\chi_{\alpha}}\delta^{\alpha}_{\gamma} =0\ec \end{equation}

which holds wherever \(C^{-1}\) is defined — strongly, not merely on \(\Sigma\). Equation (26.27) is immediate: if \(A\) is first class then \(\pb{A}{\chi_{\alpha}}\approx0\) and the whole correction vanishes weakly. For Equation (26.28), the consistency algorithm has fixed the multipliers of the second-class primary constraints, and with that fixing \(\Ham_{\text{T}}\) is first class — it is precisely the statement that the motion preserves every constraint — so \(\pb{\chi_{\beta}}{\Ham_{\text{T}}}\approx0\) and Equation (26.27) applies with \(B=\Ham_{\text{T}}\), giving \(\pb{F}{\Ham_{\text{T}}}_{\text{D}}\approx\pb{F}{\Ham_{\text{T}}} \approx\dot{F}\).

The Jacobi identity is the one property that is not a short computation; it is what earns the object the name bracket, and it is deferred.

Derives Theorem 26.21.

Two routes are written out there. The direct one expands \(\pb{A}{\pb{B}{C}_{\text{D}}}_{\text{D}}\) and its two cyclic partners: besides the Poisson Jacobi identity itself, this produces terms carrying derivatives of the inverse matrix, which are reduced by differentiating \(C^{-1}C=\identity\) to express \(\pb{A}{(C^{-1})^{\alpha\beta}}\) through \(\pb{A}{C_{\gamma\delta}}\), after which everything cancels in triples by the Poisson Jacobi identity applied to the constraints themselves. The structural one observes that the Dirac bracket is the Poisson bracket of the symplectic form induced on the second-class surface — a symplectic form precisely because \(C\) is invertible, which is Proposition 26.25 below — and reads the identity off the closure of that form.

Example 26.22 (Holonomic constraints).

A particle of mass \(m\) moving in \(\R^{3}\) under a potential \(V(\vect{x})\), with \(\Ham=\vect{p}^{2}/2m+V\), is confined to the surface \(g(\vect{x})=0\) of Definition 21.2. The Dirac bracket built from the resulting second-class pair reproduces exactly the reduced dynamics obtained in Lagrangian Mechanics by eliminating the constrained coordinate. The elementary treatment of holonomic constraints there and the formalism here agree, which is the check that fixes all signs. Rests on Definitions 21.2 and 26.20.

Derivation. Derives Example 26.22. Confinement to the surface is one condition on the coordinates, but a single \(\chi_{1}=g(\vect{x})\) is not preserved by the motion: the algorithm of Section 26.2 applied to it returns

\begin{equation}\tag{26.31} \chi_{2}:=\pb{g}{\Ham}=\frac{1}{m}\,p_{i}\,\pp_{i}g\ec \end{equation}

the statement that the velocity is tangent to the surface. Their bracket is

\begin{equation}\tag{26.32} \pb{\chi_{1}}{\chi_{2}} =\pb{g}{\tfrac{1}{m}p_{j}\pp_{j}g} =\frac{1}{m}\,\pp_{i}g\,\pp_{i}g =\frac{\abs{\nabla g}^{2}}{m}=:w\ec \end{equation}

nonzero wherever \(\nabla g\neq0\), which is where the surface is regular. The pair is second class, \(S=2\), \(F=0\), and Equation (26.21) gives \(2\times3-2=4\): two configuration degrees of freedom, as Definition 21.9 requires for a point on a surface. With

\begin{equation}\tag{26.33} C=\begin{pmatrix}0&w\\-w&0\end{pmatrix}\ec\qquad C^{-1}=\begin{pmatrix}0&-w^{-1}\\ w^{-1}&0\end{pmatrix}\ec \end{equation}

and the elementary brackets \(\pb{x_{i}}{\chi_{1}}=0\), \(\pb{x_{i}}{\chi_{2}}=\pp_{i}g/m\), \(\pb{p_{i}}{\chi_{1}}=-\pp_{i}g\) and \(\pb{p_{i}}{\chi_{2}}=-p_{k}\pp_{i}\pp_{k}g/m\), Equation (26.25) gives

\begin{align} \pb{x_{i}}{x_{j}}_{\text{D}}&=0\ec \tag{26.34}\\ \pb{x_{i}}{p_{j}}_{\text{D}} &=\delta_{ij}-\frac{\pp_{i}g\,\pp_{j}g}{\abs{\nabla g}^{2}} =\delta_{ij}-n_{i}n_{j}\ec \tag{26.35}\\ \pb{p_{i}}{p_{j}}_{\text{D}} &=\frac{p_{k}}{\abs{\nabla g}^{2}} \left[\left(\pp_{i}\pp_{k}g\right)\pp_{j}g -\left(\pp_{j}\pp_{k}g\right)\pp_{i}g\right]\ec \tag{26.36} \end{align}

with \(\vect{n}=\nabla g/\abs{\nabla g}\) the unit normal. Read these three. Equation (26.35) is the projector onto the tangent plane: the momentum conjugate to a coordinate is only the tangential part of \(\vect{p}\), which is exactly what the reduced Lagrangian treatment produces when the normal coordinate is eliminated. The normal component of \(\vect{p}\) has vanishing Dirac bracket with everything, so it is not an independent variable at all — it is \(\chi_{2}\), set strongly to zero by Equation (26.26). And Equation (26.36) says that the momenta no longer commute: the second derivatives of \(g\) are the second fundamental form of the surface, so the failure is its extrinsic curvature.

For the sphere \(g=\abs{\vect{x}}-R\) one has \(\nabla g=\hat{\vect{r}}\) and \(\pp_{i}\pp_{j}g=\left(\delta_{ij}-\hat{r}_{i}\hat{r}_{j}\right)/r\), and Equation (26.36) collapses on the constraint surface, where \(\vect{p}\cdot\hat{\vect{r}}\approx0\), to

\begin{equation}\tag{26.37} \pb{p_{i}}{p_{j}}_{\text{D}} \approx\frac{p_{i}\hat{r}_{j}-p_{j}\hat{r}_{i}}{R} =-\frac{\epsilon_{ijk}L_{k}}{R^{2}}\ec\qquad \vect{L}=\vect{x}\times\vect{p}\ec \end{equation}

so the momenta on a sphere close on the angular momentum. That is the bracket algebra of the free particle on \(S^{2}\), whose Hamiltonian is \(\vect{L}^{2}/2mR^{2}\) — precisely the reduced system obtained in Lagrangian Mechanics by writing the Lagrangian in the two angles. The two routes agree, and Equation (26.37) is the classical origin of the operator ordering problem for a particle on a curved configuration space, taken up in Section 25.5.1.

Gauge fixing and the reduced phase space

Definition 26.23 (Gauge conditions).

Given \(F\) first-class constraints \(\gamma_{A}\), a set of gauge conditions is a set of \(F\) functions \(\psi^{A}\) such that

\begin{equation}\tag{26.38} \det\pb{\gamma_{A}}{\psi^{B}}\neq0 \end{equation}

on the constraint surface. The combined set \(\left\{\gamma_{A},\psi^{A}\right\}\) is then second class, and the Dirac bracket built from it defines the dynamics on the reduced phase space of dimension \(2n-2F-S\). Rests on Theorem 26.17 and Definition 26.20.

Condition Equation (26.38) says two things at once, and both are needed. It says the gauge is attainable: since \(\pb{\psi^{B}}{\gamma_{A}}\) is the change of \(\psi^{B}\) under the gauge transformation generated by \(\gamma_{A}\), a non-vanishing determinant means the gauge parameters can be chosen to reach \(\psi=0\) from anywhere nearby. And it says the gauge is complete: no gauge freedom survives, because a residual transformation would leave \(\psi\) unchanged and make the determinant vanish. The count then follows from Theorem 26.17 with \(F\to0\) and \(S\to S+2F\).

Proposition 26.24 (The Faddeev–Popov determinant).

Write \(\Gamma_{a}=\left(\gamma_{A},\psi^{B}\right)\) for the combined second-class set of Definition 26.23. Then

\begin{equation}\tag{26.39} \det\pb{\Gamma_{a}}{\Gamma_{b}} \approx\left(\det\pb{\gamma_{A}}{\psi^{B}}\right)^{2}\ec \end{equation}

so that the factor \(\sqrt{\det\pb{\Gamma_{a}}{\Gamma_{b}}}\) by which the Liouville measure of Definition 24.20 is corrected on passing to the gauge-fixed surface is \(\abs{\det\pb{\gamma_{A}}{\psi^{B}}}\). This is the object that appears as the Faddeev–Popov determinant in the path-integral quantization of Path-Integral Quantization [Faddeev:1967]. Its exponentiation as an integral over anticommuting fields is what introduces the ghosts of gauge-fixed perturbation theory. Rests on Definition 26.23 and Proposition 26.19.

Proof.

Derives Proposition 26.24. Order the combined set as \(\left(\gamma_{A},\psi^{B}\right)\) and write \(D_{A}{}^{B}:=\pb{\gamma_{A}}{\psi^{B}}\). Because the \(\gamma_{A}\) are first class, \(\pb{\gamma_{A}}{\gamma_{B}}\approx0\), so on \(\Sigma\)

\begin{equation}\tag{26.40} \pb{\Gamma_{a}}{\Gamma_{b}} \approx\begin{pmatrix}0&D\\-D\transpose&E\end{pmatrix}\ec\qquad E^{AB}=\pb{\psi^{A}}{\psi^{B}}\ec \end{equation}

with all four blocks \(F\times F\) and \(D\) invertible by Equation (26.38). Multiply on the left by \(\begin{pmatrix}\identity&-DE^{-1}\\0&\identity\end{pmatrix}\), of determinant \(1\), where \(E\) is invertible; the product is \(\begin{pmatrix}DE^{-1}D\transpose&0\\-D\transpose&E\end{pmatrix}\), whose determinant is \(\det\left(DE^{-1}D\transpose\right)\det E=(\det D)^{2}\), in which \(E\) has cancelled. Both sides of Equation (26.39) are polynomials in the entries of an arbitrary \(F\times F\) matrix \(E\) — the antisymmetry of \(E\) is nowhere used — and they have just been shown equal on the dense open set of invertible \(E\), so they are equal for every \(E\). In particular the identity holds in the common case \(E=0\), where the gauge conditions commute among themselves, and this matters: for odd \(F\) an antisymmetric \(E\) is never invertible, so the perturbation must be taken outside the antisymmetric matrices, which is what the previous sentence does. Taking the positive square root gives the stated factor.

The second sentence of Proposition 26.24 — that this determinant is the factor by which the Liouville measure is corrected — is a separate statement, and it is the measure-theoretic face of the Dirac bracket.

Proposition 26.25 (The Liouville measure of a second-class surface).

Let \(\chi_{\alpha}\), \(\alpha=1,\ldots,S\), be a second-class set on a phase space of dimension \(2n\), so that \(C_{\alpha\beta}\) of Equation (26.24) is invertible on the surface \(\Sigma_{\chi}\) they define, and write \(\Theta\) for the Liouville form Equation (24.13) of the ambient space — the letter \(\Omega\) being reserved in this chapter for the BRST charge Equation (26.97). Then \(\Sigma_{\chi}\) carries the non-degenerate restriction \(\omega_{\Sigma}\) of the symplectic form, \(\det C>0\), and the Liouville form \(\Theta_{\Sigma}\) built from \(\omega_{\Sigma}\) satisfies

\begin{equation}\tag{26.41} \int_{\Sigma_{\chi}}F\,\Theta_{\Sigma} =\int F\,\sqrt{\det C}\, \prod_{\alpha=1}^{S}\delta\!\left(\chi_{\alpha}\right)\Theta \end{equation}

for every integrable \(F\). The reduced measure is thus the ambient one with a delta function of each constraint and the factor \(\sqrt{\det C}\), which for a gauge-fixed set \(\Gamma_{a}\) is the \(\sqrt{\det\pb{\Gamma_{a}}{\Gamma_{b}}}\) of Proposition 26.24. Rests on Proposition 26.19, Definition 24.20 and Proposition 24.3.

Proof.

Derives Proposition 26.25. Write \(M\) for the phase space, so that \(\dim M=2n\). Everything is pointwise linear algebra in the tangent space \(T_{x}M\) at a point \(x\in\Sigma_{\chi}\), plus one convention about the delta functions.

The splitting. With the conventions Equations (24.5) and (24.7), \(\dd\chi_{\gamma}\!\left(X_{\chi_{\alpha}}\right) =\pb{\chi_{\gamma}}{\chi_{\alpha}}=C_{\gamma\alpha}\), which is invertible; so the \(S\) vectors \(X_{\chi_{\alpha}}\) are independent and span a subspace \(W\), and the \(S\) covectors \(\dd\chi_{\alpha}\) are independent, with common kernel \(T_{x}\Sigma_{\chi}\) of dimension \(2n-S\). The map

\begin{equation}\tag{26.42} P(v)=v-X_{\chi_{\alpha}}\left(C^{-1}\right)^{\alpha\beta} \dd\chi_{\beta}(v) \end{equation}

gives \(\dd\chi_{\gamma}\!\left(P(v)\right) =\dd\chi_{\gamma}(v)-C_{\gamma\alpha}\left(C^{-1}\right)^{\alpha\beta} \dd\chi_{\beta}(v)=0\), so it projects \(T_{x}M\) onto \(T_{x}\Sigma_{\chi}\) along \(W\), and \(T_{x}M=W\oplus T_{x}\Sigma_{\chi}\). The two summands are \(\omega\)-orthogonal: for \(v\in T_{x}\Sigma_{\chi}\), \(\omega\!\left(X_{\chi_{\alpha}},v\right) =\left(\iota_{X_{\chi_{\alpha}}}\omega\right)(v) =\dd\chi_{\alpha}(v)=0\). On \(W\) the form has the invertible matrix \(\omega\!\left(X_{\chi_{\alpha}},X_{\chi_{\beta}}\right) =C_{\alpha\beta}\); and a \(v\in T_{x}\Sigma_{\chi}\) annihilated by \(\omega_{\Sigma}\) is by that orthogonality annihilated by \(\omega\) on all of \(T_{x}M\), hence zero. Both restrictions are therefore non-degenerate, \(\left(\Sigma_{\chi},\omega_{\Sigma}\right)\) is symplectic, and \(S=2s\) is even — which is Proposition 26.19 read geometrically.

The volume factorises. Write \(\omega=\omega_{1}+\omega_{2}\), where \(\omega_{1}\) is the pullback of \(\omega|_{W}\) along \(v\longmapsto v-P(v)\) and \(\omega_{2}\) the pullback of \(\omega_{\Sigma}\) along \(P\); that there is no cross term is the \(\omega\)-orthogonality just proved. Two-forms commute under \(\wedge\), and \(\omega_{1}^{\wedge k}=0\) for \(k>s\) while \(\omega_{2}^{\wedge k}=0\) for \(k>n-s\), each being pulled back from a space of that dimension. Expanding \(\omega^{\wedge n}\) by the binomial theorem therefore leaves a single term:

\begin{equation}\tag{26.43} \Theta=\frac{\omega^{\wedge n}}{n!} =\frac{\omega_{1}^{\wedge s}}{s!}\wedge \frac{\omega_{2}^{\wedge\left(n-s\right)}}{\left(n-s\right)!} =\Theta_{W}\wedge P^{*}\Theta_{\Sigma}\ec \end{equation}

with \(\Theta_{W}\) the Liouville form of the symplectic space \(W\).

Two evaluations. Take the adapted basis \(\left(X_{\chi_{1}},\ldots,X_{\chi_{S}},v_{1},\ldots,v_{2n-S}\right)\) with the \(v_{r}\) a basis of \(T_{x}\Sigma_{\chi}\). By Proposition 24.3 pick a canonical basis \(E_{1},\ldots,E_{S}\) of \(\left(W,\omega|_{W}\right)\), ordered as in Equation (24.13) so that \(\Theta_{W}\) takes the value \(1\) on it, and write \(X_{\chi_{\alpha}}=L_{\alpha}{}^{a}E_{a}\). Then \(C=L\,J\,{L}\transpose\) with \(J\) the matrix Equation (5.144), whose determinant is \(1\), so \(\det C=\left(\det L\right)^{2}\) — positive, as claimed — while \(\Theta_{W}\!\left(X_{\chi_{1}},\ldots,X_{\chi_{S}}\right)=\det L\). Since \(\Theta_{W}\) annihilates every \(v_{r}\) and \(P^{*}\Theta_{\Sigma}\) every \(X_{\chi_{\alpha}}\), only one term of the wedge in Equation (26.43) survives on that basis:

\begin{equation}\tag{26.44} \Theta\left(X_{\chi_{1}},\ldots,v_{2n-S}\right) =\pm\sqrt{\det C}\; \Theta_{\Sigma}\left(v_{1},\ldots,v_{2n-S}\right)\ep \end{equation}

The other top form built from the constraints evaluates on the same basis to

\begin{equation}\tag{26.45} \left(\dd\chi_{1}\wedge\cdots\wedge\dd\chi_{S} \wedge P^{*}\Theta_{\Sigma}\right) \left(X_{\chi_{1}},\ldots,v_{2n-S}\right) =\det C\;\Theta_{\Sigma}\left(v_{1},\ldots,v_{2n-S}\right)\ec \end{equation}

because \(\det\left(\dd\chi_{\alpha}(X_{\chi_{\beta}})\right)=\det C\). Two top forms on a \(2n\)-dimensional space that agree up to a factor on one basis agree up to that factor everywhere, so, orienting \(\Sigma_{\chi}\) so that the sign in Equation (26.44) is positive,

\begin{equation}\tag{26.46} \Theta=\frac{1}{\sqrt{\det C}}\, \dd\chi_{1}\wedge\cdots\wedge\dd\chi_{S} \wedge P^{*}\Theta_{\Sigma}\ep \end{equation}

The convention. A delta function of a constraint means one thing: if a top form is presented as \(\dd\chi_{1}\wedge\cdots\wedge\dd\chi_{S}\wedge\eta\) for a \(\left(2n-S\right)\)-form \(\eta\), then integrating \(\prod_{\alpha}\delta(\chi_{\alpha})\) against it is integrating \(\eta\) pulled back to \(\Sigma_{\chi}\). Applied to Equation (26.46), and using that \(P\) restricts to the identity on \(T\Sigma_{\chi}\), so that \(P^{*}\Theta_{\Sigma}\) pulls back to \(\Theta_{\Sigma}\),

\[ \int F\,\sqrt{\det C}\,\prod_{\alpha}\delta(\chi_{\alpha})\,\Theta =\int_{\Sigma_{\chi}}F\,\sqrt{\det C}\, \frac{\Theta_{\Sigma}}{\sqrt{\det C}} =\int_{\Sigma_{\chi}}F\,\Theta_{\Sigma}\ec \]

which is Equation (26.41).

Remark 26.26 (What the factor is doing).

Equation (26.41) is the reason the Faddeev–Popov determinant is not an arbitrary insertion. The delta functions alone would define a measure on \(\Sigma_{\chi}\) that depends on how the constraints are written: replacing \(\chi_{\alpha}\) by \(A_{\alpha}{}^{\beta}\chi_{\beta}\) for an invertible matrix of functions leaves the surface untouched but multiplies \(\prod\delta(\chi_{\alpha})\) by \(\abs{\det A}^{-1}\). It multiplies \(\sqrt{\det C}\) by \(\abs{\det A}\) as well, since \(C\mapsto A\,C\,{A}\transpose\) on the surface, so the product is invariant — and it is invariant because it equals the intrinsic Liouville volume of the reduced symplectic manifold, which knows nothing about the presentation. The same computation, read on the tangent space rather than on the volume, is what makes the Dirac bracket of Definition 26.20 the Poisson bracket of \(\omega_{\Sigma}\). Rests on Propositions 26.24 and 26.25.

Remark 26.27 (Gauge conditions need not exist globally).

Equation (26.38) is a local condition. For non-abelian gauge theories no gauge condition satisfies it over the whole space of field configurations—the Gribov ambiguity [Gribov:1978]—so the reduced phase space is not covered by one chart, and a gauge orbit meets the gauge-fixing surface more than once. This is a genuine obstruction, not a technical inconvenience: it is invisible in perturbation theory about \(A=0\), which is why the Faddeev–Popov procedure works there, and it is exactly what fails when the coupling is strong. Its consequences for the non-perturbative treatment of Quantum Chromodynamics are noted there.

Worked systems

The formalism above is worth exactly as much as the cases it settles. The following are the ones this treatise needs; each is stated with its constraint structure and its degree-of-freedom count.

Notation 26.28 (Field systems).

From Section 26.6.2 onward the systems are field theories. A plain \(H=\int\dd^{3}x\,\Ham\) is the Hamiltonian and the script \(\Ham\) the corresponding density; the index \(j\) of Equation (26.11) carries a continuous label \(\vect{x}\) and the sums over constraints become integrals; and the Poisson bracket is built from functional derivatives,

\begin{equation}\tag{26.47} \pb{F}{G}=\int\dd^{3}x\left( \frac{\delta F}{\delta\varphi(\vect{x})} \frac{\delta G}{\delta\pi(\vect{x})} -\frac{\delta F}{\delta\pi(\vect{x})} \frac{\delta G}{\delta\varphi(\vect{x})}\right)\ec \end{equation}

so that \(\pb{\varphi(\vect{x})}{\pi(\vect{y})} =\delta^{3}(\vect{x}-\vect{y})\). The evolution parameter is the coordinate time \(t\) of a chosen inertial frame, and \(x^{0}=ct\); a dot means \(\pp_{t}\). Momenta conjugate to a field are taken with respect to \(\pp_{t}\) of the covariant components \(A_{\mu}\), which is what makes the Gauss constraint come out as Equation (26.56) rather than with the opposite sign. The prices and the limits of the passage from finitely many degrees of freedom to a field are collected in Section 26.8. Rests on Notation 26.1 and Definition 22.29.

The relativistic free particle

Proposition 26.29 (Reparametrization invariance and the mass-shell constraint).

For the action of a free relativistic particle parametrized by an arbitrary parameter \(\tau\),

\begin{equation}\tag{26.48} S=-mc\int\sqrt{\eta_{\mu\nu} \dv{x^{\mu}}{\tau}\dv{x^{\nu}}{\tau}}\ \dd\tau\ec \end{equation}

in the signature of Minkowski Space and Its Symmetries, the Legendre transformation is singular of rank three, the canonical Hamiltonian vanishes identically, and the single primary constraint is the mass shell

\begin{equation}\tag{26.49} \phi=\eta^{\mu\nu}p_{\mu}p_{\nu}-m^{2}c^{2}\approx0\ec \end{equation}

which is first class. The count is \(2\times4-2\times1=6\), that is three physical configuration degrees of freedom—the particle's position in space—as it must be. Rests on Proposition 26.5, Theorem 40.5 and Theorem 26.17.

Proof.

Derives Proposition 26.29. Write \(\dot{x}^{\mu}=\dd x^{\mu}/\dd\tau\) and \(\Lag=-mc\left(\eta_{\alpha\beta}\dot{x}^{\alpha}\dot{x}^{\beta} \right)^{1/2}\). The momenta are

\begin{equation}\tag{26.50} p_{\mu}=\pdv{\Lag}{\dot{x}^{\mu}} =-mc\,\frac{\eta_{\mu\nu}\dot{x}^{\nu}} {\sqrt{\eta_{\alpha\beta} \dot{x}^{\alpha}\dot{x}^{\beta}}}\ec \end{equation}

and contracting Equation (26.50) with itself the normalization cancels the denominator exactly: \(\eta^{\mu\nu}p_{\mu}p_{\nu}=m^{2}c^{2}\) for every \(\dot{x}\). That is Equation (26.49), and it holds identically in the velocities, which by Proposition 26.5 makes it a primary constraint. The Hessian

\begin{equation}\tag{26.51} W_{\mu\nu}=\frac{\pp^{2}\Lag}{\pp\dot{x}^{\mu}\pp\dot{x}^{\nu}} =-\frac{mc}{\sqrt{\dot{x}^{2}}} \left(\eta_{\mu\nu} -\frac{\dot{x}_{\mu}\dot{x}_{\nu}}{\dot{x}^{2}}\right) \end{equation}

is a multiple of the projector orthogonal to \(\dot{x}\), of rank three, confirming \(M=4-3=1\): exactly one constraint, no more.

The canonical Hamiltonian is

\begin{equation}\tag{26.52} \Ham_{\text{c}}=p_{\mu}\dot{x}^{\mu}-\Lag =-mc\frac{\dot{x}^{2}}{\sqrt{\dot{x}^{2}}} +mc\sqrt{\dot{x}^{2}}=0\ec \end{equation}

identically, so \(\Ham_{\text{T}}=u\phi\) by Equation (26.6). Consistency is automatic, \(\dot{\phi}=\pb{\phi}{u\phi}=0\): no secondary constraint arises and \(u\) is never determined. By Theorem 26.14 the constraint is first class — there is nothing for its bracket to fail to annihilate — and \(u\) is a gauge parameter. The equations of motion it generates,

\begin{equation}\tag{26.53} \dot{x}^{\mu}=\pb{x^{\mu}}{\Ham_{\text{T}}}=2u\,p^{\mu}\ec\qquad \dot{p}_{\mu}=0\ec \end{equation}

determine the worldline but not its parametrization: changing \(u(\tau)\) reparametrizes \(\tau\) and nothing else. Choosing \(\tau\) to be the proper time of Definition 38.12 fixes \(u=1/2m\) and returns \(p^{\mu}=m\,\dd x^{\mu}/\dd\tau\), the four-momentum of Definition 40.4. Finally \(2n=8\), \(F=1\), \(S=0\), so Equation (26.21) gives \(8-2=6\) and three configuration degrees of freedom.

Remark 26.30 (A sign to be read once).

With the signature \((+,-,-,-)\) and the overall minus in Equation (26.48), Equation (26.50) evaluated in the parametrization \(\tau=t\) gives \(p_{i}=\gamma_{u}m\,u_{i}\) — the ordinary relativistic momentum, with its familiar sign — but \(p_{0}=-E/c\). The canonical momenta conjugate to the \(x^{\mu}\) are thus the components of \(-p_{\mu}\) in the convention of Definition 40.4. Nothing in the constraint analysis depends on this: Equation (26.49) is quadratic and blind to the overall sign, and it is the same statement as Equation (40.5). The sign is recorded because it is a standing trap in the literature, where both conventions are used without comment.

Remark 26.31 (A vanishing Hamiltonian is not a static system).

That \(\Ham_{\text{c}}\approx0\) says only that the evolution in the arbitrary parameter is pure gauge; the physical motion is the relation between the \(x^{\mu}\), which is unaffected. Any theory whose invariance group includes reparametrizations of the evolution parameter has this property, and the confusion it causes when the parameter is called “time” is discussed under Section 26.6.5.

The electromagnetic field

Proposition 26.32 (Constraint structure of the free Maxwell field).

For the Maxwell Lagrangian density of Section 64.5.1,

\begin{equation}\tag{26.54} \Lag=-\frac{1}{4\mu_{0}}F_{\mu\nu}F^{\mu\nu}-A_{\mu}J^{\mu}\ec\qquad F_{\mu\nu}=\pp_{\mu}A_{\nu}-\pp_{\nu}A_{\mu}\ec \end{equation}

the momentum conjugate to \(A_{0}\) vanishes identically, giving the primary constraint

\begin{equation}\tag{26.55} \pi^{0}(\vect{x})\approx0\ec \end{equation}

whose consistency yields the secondary constraint—Gauss's law—

\begin{equation}\tag{26.56} \pp_{i}\pi^{i}(\vect{x})-\rho(\vect{x})\approx0\ep \end{equation}

Both are first class, and they generate the gauge transformations \(A_{\mu}\mapsto A_{\mu}+\pp_{\mu}\Lambda\). The count per point of space is \(2\times4-2\times2=4\), that is two physical field degrees of freedom. The algorithm closes after two rounds if and only if the source is conserved, \(\pp_{\mu}J^{\mu}=0\). Rests on Proposition 26.5, Proposition 26.10 and Theorem 26.17.

Proof.

Derives Proposition 26.32. Since \(F_{\mu\nu}\) is antisymmetric, \(\Lag\) contains no \(\pp_{0}A_{0}\) at all, and

\begin{equation}\tag{26.57} \pi^{\mu}:=\pdv{\Lag}{\dot{A}_{\mu}} =-\frac{1}{\mu_{0}c}F^{0\mu}\ec\qquad\text{so}\qquad \pi^{0}=0\ec\quad \pi^{i}=\epsilon_{0}E^{i}\ec \end{equation}

using \(F^{0i}=-E^{i}/c\) and \(\epsilon_{0}\mu_{0}c^{2}=1\). The first is Equation (26.55), one primary constraint per point; the second identifies the field momentum with the electric displacement, of dimension \(\mathrm{C}/\mathrm{m}^{2}\). The Hessian in the four \(\dot{A}_{\mu}\) has rank three per point, so there is exactly one.

The canonical Hamiltonian follows from Equation (26.4). Using \(\dot{A}_{i}=E_{i}+\pp_{i}\phi\) with \(\phi=cA_{0}\) the electrostatic potential, and integrating the term \(\epsilon_{0}\vect{E}\cdot\nabla\phi\) by parts,

\begin{equation}\tag{26.58} H_{\text{c}}=\int\dd^{3}x\left[ \frac{\pi^{i}\pi^{i}}{2\epsilon_{0}} +\frac{\left(\nabla\times\vect{A}\right)^{2}}{2\mu_{0}} -c\,A_{0}\left(\pp_{i}\pi^{i}-\rho\right) -\vect{J}\cdot\vect{A}\right]\ec \end{equation}

whose first two terms are the familiar field energy density \(\tfrac{1}{2}\epsilon_{0}E^{2}+B^{2}/2\mu_{0}\). The potential \(A_{0}\) appears linearly and without derivatives: it is a multiplier, not a dynamical variable. Consistency of Equation (26.55) is then

\begin{equation}\tag{26.59} \dot{\pi}^{0}(\vect{x}) =\pb{\pi^{0}(\vect{x})}{H_{\text{T}}} =-\frac{\delta H_{\text{c}}}{\delta A_{0}(\vect{x})} =c\left(\pp_{i}\pi^{i}-\rho\right)\approx0\ec \end{equation}

which is outcome 2 of Proposition 26.10: a new constraint containing no multiplier. It is Equation (26.56), and written out it is \(\epsilon_{0}\nabla\cdot\vect{E}=\rho\) — Gauss's law appears here as a constraint on initial data, not as an equation of motion, which is the structural statement this analysis makes about Maxwell's system.

Running the algorithm once more, the remaining Hamilton equation is \(\dot{\pi}^{i}=-\delta H_{\text{c}}/\delta A_{i} =\left(\nabla\times\vect{B}\right)_{i}/\mu_{0}-J_{i}\), the Ampère–Maxwell law, so

\begin{equation}\tag{26.60} \frac{\dd}{\dd t}\left(\pp_{i}\pi^{i}-\rho\right) =\pp_{i}\frac{\left(\nabla\times\vect{B}\right)_{i}}{\mu_{0}} -\nabla\cdot\vect{J}-\dot{\rho} =-\left(\dot{\rho}+\nabla\cdot\vect{J}\right)\ec \end{equation}

the divergence of a curl vanishing identically. The algorithm therefore terminates precisely when \(\pp_{\mu}J^{\mu}=0\), and otherwise reaches outcome 4 of Proposition 26.10: a Maxwell field coupled to a non-conserved source has no solutions at all.

Both constraints are first class. \(\pi^{0}\) has vanishing bracket with everything built from \(\vect{A}\) and \(\vect{\pi}\); and \(\pb{\mathcal{G}(\vect{x})}{\mathcal{G}(\vect{y})}=0\) identically, where \(\mathcal{G}=\pp_{i}\pi^{i}-\rho\), because \(\mathcal{G}\) depends on the momenta alone. So \(F=2\) per point, \(S=0\), \(2n=8\), and Equation (26.21) gives \(8-4=4\): two configuration degrees of freedom per point of space.

It remains to identify the gauge transformation. Take

\begin{equation}\tag{26.61} G[\Lambda]=\int\dd^{3}x\left[ \frac{\dot{\Lambda}}{c}\,\pi^{0} -\Lambda\left(\pp_{i}\pi^{i}-\rho\right)\right]\ec \end{equation}

a combination of the primary and secondary constraints with an arbitrary function \(\Lambda\) and its time derivative. Then \(\pb{A_{0}}{G}=\dot{\Lambda}/c=\pp_{0}\Lambda\) and, integrating the second term by parts, \(\pb{A_{i}}{G}=\pp_{i}\Lambda\), while \(\pb{\pi^{\mu}}{G}=0\). That is \(A_{\mu}\mapsto A_{\mu}+\pp_{\mu}\Lambda\) with \(\vect{E}\) and \(\vect{B}\) unchanged, as claimed.

Remark 26.33 (The count is the observation).

The two degrees of freedom of Proposition 26.32 are the two transverse polarizations of light. They are not a formal result: they are what is measured whenever the polarization of an electromagnetic wave is analysed (Experiment: Wave Optics), and the absence of a third, longitudinal, mode is a statement about nature that the constraint analysis explains rather than assumes. The same count performed for a massless spin-two field gives two polarizations again—the plus and cross modes of Gravitational-Wave Theory, observed as such in the detections of Experiment: Gravitational Waves [Maggiore:2008].

The massive vector field

Proposition 26.34 (Proca constraints are second class).

For the massive vector field of Section 64.6.1 [Proca:1936],

\begin{equation}\tag{26.62} \Lag=-\frac{1}{4\mu_{0}}F_{\mu\nu}F^{\mu\nu} +\frac{1}{2\mu_{0}}\left(\frac{mc}{\hbar}\right)^{2} A_{\mu}A^{\mu}\ec \end{equation}

the primary constraint is again \(\pi^{0}\approx0\) but its consistency returns

\begin{equation}\tag{26.63} \pp_{i}\pi^{i} +\frac{1}{\mu_{0}c}\left(\frac{mc}{\hbar}\right)^{2}A_{0} -\rho\approx0\ec \end{equation}

and the two have the nonvanishing bracket

\begin{equation}\tag{26.64} \pb{\pi^{0}(\vect{x})}{\,\cdot\,} \longrightarrow -\frac{1}{\mu_{0}c}\left(\frac{mc}{\hbar}\right)^{2} \delta^{3}(\vect{x}-\vect{y})\ec \end{equation}

proportional to the mass squared. They are second class rather than first: there is no gauge invariance. The count per point is \(2\times4-2=6\), that is three physical degrees of freedom—the three polarizations of a massive spin-one particle. Rests on Proposition 26.32, Definition 26.12 and Theorem 26.17.

Proof.

Derives Proposition 26.34. The mass term of Equation (26.62) carries no derivatives, so the momenta Equation (26.57) are unchanged and \(\pi^{0}=0\) is still the only primary constraint. The canonical Hamiltonian Equation (26.58) acquires the extra density \(-\left(2\mu_{0}\right)^{-1}\left(mc/\hbar\right)^{2} \left(A_{0}^{2}-A_{i}A_{i}\right)\), and now \(A_{0}\) appears quadratically. Consistency of \(\pi^{0}\approx0\) therefore gives

\begin{equation}\tag{26.65} \dot{\pi}^{0} =-\frac{\delta H_{\text{c}}}{\delta A_{0}} =c\left[\pp_{i}\pi^{i}-\rho +\frac{1}{\mu_{0}c}\left(\frac{mc}{\hbar}\right)^{2}A_{0}\right] \approx0\ec \end{equation}

which is Equation (26.63). Writing \(\chi_{1}=\pi^{0}\) and \(\chi_{2}\) for the left-hand side of Equation (26.63), the only term of \(\chi_{2}\) that fails to commute with \(\pi^{0}\) is the one containing \(A_{0}\), and

\begin{equation}\tag{26.66} \pb{\chi_{1}(\vect{x})}{\chi_{2}(\vect{y})} =-\frac{1}{\mu_{0}c}\left(\frac{mc}{\hbar}\right)^{2} \delta^{3}(\vect{x}-\vect{y})\ec \end{equation}

which is Equation (26.64) and is invertible for every \(m\neq0\). By Definition 26.12 the pair is second class, so \(S=2\) and \(F=0\); consistency of \(\chi_{2}\) then determines the multiplier instead of producing a third constraint, and the algorithm stops. Equation (26.21) gives \(8-0-2=6\) per point, three configuration degrees of freedom.

The content of Equation (26.63) is worth stating separately: it determines \(A_{0}\) algebraically in terms of the momenta,

\begin{equation}\tag{26.67} A_{0}=-\mu_{0}c\left(\frac{\hbar}{mc}\right)^{2} \left(\pp_{i}\pi^{i}-\rho\right)\ec \end{equation}

so \(A_{0}\) is not free, not gauge, and not dynamical — it is an auxiliary field. Combining Equation (26.67) with the equation of motion \(\pp_{\mu}F^{\mu\nu}+\left(mc/\hbar\right)^{2}A^{\nu} =\mu_{0}J^{\nu}\) and taking a divergence returns \(\pp_{\mu}A^{\mu}=0\) for a conserved source: the Lorenz condition, here a consequence rather than a choice, exactly as Section 64.6.1 says.

Remark 26.35 (Why the massive and massless counts differ).

The passage from Proposition 26.32 to Proposition 26.34 changes the classification of the constraints, not their number, and that is what changes the count from two to three. Note precisely where the change enters: Equation (26.66) is proportional to \(m^{2}\), so the bracket that makes the pair second class vanishes discontinuously at \(m=0\); the count is a step function of the mass, and no \(m\to0\) limit of the massive theory has two degrees of freedom. The physical content is that a mass term breaks the gauge invariance and thereby liberates the longitudinal mode. Classically this is the shadow of the van Dam–Veltman–Zakharov discontinuity [vanDam:1970] [Zakharov:1970], where the corresponding step in the spin-two count leaves a massless limit that disagrees with general relativity. Experimentally it is why “the photon is massless” is a bound and never a measurement: the current limit is \(m_{\gamma}<10^{-18}\,\mathrm{eV}/c^{2}\) [Navas:2024], from laboratory tests of Coulomb's law [Williams:1971] and, far more sharply, from planetary and galactic magnetic fields [Goldhaber:2010]. And it is the mechanism behind the massive \(W\) and \(Z\) bosons of Electroweak Unification and the Higgs Boson, with \(M_{W}c^{2}=80.3692(133)\,\mathrm{GeV}\) and \(M_{Z}c^{2}=91.1880(20)\,\mathrm{GeV}\) [Navas:2024], where the liberated modes are supplied by the scalar field rather than imposed by hand.

Yang–Mills fields

The non-abelian case is the one the Standard Model actually uses, and it differs from Proposition 26.32 in exactly one place: the Gauss constraints no longer commute with one another. Everything else — the vanishing time-component momentum, the secondary constraint, the count of two polarizations per generator — survives unchanged.

Notation 26.36 (Gauge fields and their SI dimensions).

Let the gauge group be compact and simple of dimension \(N\), with structure constants \(f^{abc}\) totally antisymmetric in the basis of Notation 102.18, whose conventions are used throughout. The potentials \(A^{a}_{\mu}\) carry the dimension \(\mathrm{V}\,\mathrm{s}/\mathrm{m}\) of the electromagnetic potential and the coupling \(g\) the dimension \(\mathrm{C}\) of a charge, so that

\begin{equation}\tag{26.68} \left[\frac{g}{\hbar}\right]=\mathrm{C}/\mathrm{J}/\mathrm{s}\ec\qquad \left[\frac{g}{\hbar}A^{a}_{\mu}\right]=/\mathrm{m}\ep \end{equation}

The ratio \(g/\hbar\) is the only combination in which the coupling occurs below, and by Equation (26.68) it is what turns a potential into an inverse length; it is the scalar behind the matrix-valued combination Equation (102.18). The field strength is Equation (102.24), \(F^{a}_{\mu\nu}=\pp_{\mu}A^{a}_{\nu}-\pp_{\nu}A^{a}_{\mu} -\frac{g}{\hbar} f^{abc}A^{b}_{\mu}A^{c}_{\nu}\), of dimension \(\mathrm{T}\), and the covariant derivative in the adjoint representation acts on a colour vector \(X^{a}\) as

\begin{equation}\tag{26.69} \left(D_{i}X\right)^{a} =\pp_{i}X^{a}-\frac{g}{\hbar} f^{abc}A^{b}_{i}X^{c}\ep \end{equation}

Rests on Notation 102.18, Equation (102.24) and Notation 26.28.

Proposition 26.37 (Constraint structure of Yang–Mills theory).

For \(\Lag=-\left(4\mu_{0}\right)^{-1}F^{a}_{\mu\nu}F^{a\,\mu\nu}\) coupled to matter of colour charge density \(\rho^{a}\), the momenta are \(\pi^{0}_{a}=0\) and \(\pi^{i}_{a}=\epsilon_{0}E^{i}_{a}\), so there are \(N\) primary constraints

\begin{equation}\tag{26.70} \pi^{0}_{a}(\vect{x})\approx0\ec\qquad a=1,\ldots,N\ec \end{equation}

whose consistency yields the \(N\) secondary constraints — the non-abelian Gauss law —

\begin{equation}\tag{26.71} \mathcal{G}^{a} :=\left(D_{i}\pi^{i}\right)^{a}-\rho^{a}\approx0\ep \end{equation}

All \(2N\) are first class, and the Gauss constraints close on the Lie algebra of the gauge group,

\begin{equation}\tag{26.72} \pb{\mathcal{G}^{a}(\vect{x})}{\mathcal{G}^{b}(\vect{y})} =-\frac{g}{\hbar} f^{abc}\,\mathcal{G}^{c}(\vect{x})\, \delta^{3}(\vect{x}-\vect{y})\ec \end{equation}

with genuine constants \(f^{abc}\) where Equation (26.56) had zero. The count per point of space is \(2\times4N-2\times2N=4N\): two physical degrees of freedom per generator, that is two transverse polarizations for each of the \(N\) gauge bosons. Rests on Proposition 26.32, Proposition 26.13 and Notation 26.36.

Proof.

Derives Proposition 26.37. \(F^{a}_{\mu\nu}\) is antisymmetric in \(\mu\nu\) whatever the nonlinear term does, so no \(\pp_{0}A^{a}_{0}\) occurs in \(\Lag\) and Equation (26.70) follows exactly as Equation (26.55) did. The Legendre transformation gives, by the computation of Proposition 26.32 with \(\pp_{i}\) replaced by \(D_{i}\) throughout,

\begin{equation}\tag{26.73} H_{\text{c}}=\int\dd^{3}x\left[ \frac{\pi^{i}_{a}\pi^{i}_{a}}{2\epsilon_{0}} +\frac{B^{i}_{a}B^{i}_{a}}{2\mu_{0}} -c\,A^{a}_{0}\,\mathcal{G}^{a}\right]\ec \end{equation}

with \(B^{i}_{a}=-\tfrac{1}{2}\epsilon^{ijk}F^{a}_{jk}\) the chromomagnetic field, in which \(A^{a}_{0}\) is again a multiplier, so that consistency of Equation (26.70) produces Equation (26.71). The two energy terms are the abelian ones; what the nonlinearity of Equation (102.24) adds is hidden inside \(B^{i}_{a}\), which is cubic and quartic in \(A\), so the free Yang–Mills Hamiltonian already describes a self-interacting field.

For the algebra, smear the constraint with a colour-valued test function, \(G[\varepsilon]:=\int\dd^{3}x\,\varepsilon^{a}\mathcal{G}^{a}\), and take the matter contribution to \(\rho^{a}\) to generate the colour rotation of the matter fields, which it does by construction. Then

\begin{equation}\tag{26.74} \frac{\delta G[\varepsilon]}{\delta A^{b}_{i}} =-\frac{g}{\hbar} f^{abc}\varepsilon^{a}\pi^{i}_{c}\ec\qquad \frac{\delta G[\varepsilon]}{\delta \pi^{i}_{b}} =-\left(D_{i}\varepsilon\right)^{b}\ec \end{equation}

the second by Equation (26.69) and the total antisymmetry of \(f\). Substituting both into Equation (26.47),

\begin{equation}\tag{26.75} \pb{G[\varepsilon]}{G[\eta]} =\frac{g}{\hbar}\int\dd^{3}x\;f^{abc}\pi^{i}_{c} \left[\varepsilon^{a}\left(D_{i}\eta\right)^{b} -\eta^{a}\left(D_{i}\varepsilon\right)^{b}\right]\ep \end{equation}

Integrate the first term by parts. The adjoint derivative is antisymmetric under integration by parts because \(f\) is totally antisymmetric, and it is a derivation, so

\begin{equation}\tag{26.76} \int f^{abc}\pi^{i}_{c}\varepsilon^{a}\left(D_{i}\eta\right)^{b} =-\int\eta^{b}\left[f^{abc}\left(D_{i}\varepsilon\right)^{a}\pi^{i}_{c} +f^{abc}\varepsilon^{a}\mathcal{G}^{c}\right]\ec \end{equation}

where \(\left(D_{i}\pi^{i}\right)^{c}\) has been recognised as \(\mathcal{G}^{c}\) up to the matter term, which the matter bracket supplies. The first bracket on the right cancels the second term of Equation (26.75) after relabelling \(a\leftrightarrow b\) and using \(f^{abc}=-f^{bac}\), leaving

\begin{equation}\tag{26.77} \pb{G[\varepsilon]}{G[\eta]} =-\frac{g}{\hbar}\int\dd^{3}x\;f^{cab}\varepsilon^{a}\eta^{b} \,\mathcal{G}^{c} =-\frac{g}{\hbar}\,G\!\left[\comm{\varepsilon}{\eta}\right]\ec\qquad \comm{\varepsilon}{\eta}^{c}=f^{cab}\varepsilon^{a}\eta^{b}\ec \end{equation}

which unsmeared is Equation (26.72). Since the right-hand side is a combination of the constraints, and \(\pb{\pi^{0}_{a}}{\mathcal{G}^{b}} =0\) because \(\mathcal{G}\) contains no \(A^{a}_{0}\), all \(2N\) constraints are first class by Definition 26.12, and Proposition 26.13 is realised with constant structure functions. The count follows from Equation (26.21) with \(2n=8N\), \(F=2N\), \(S=0\).

Remark 26.38 (What the non-abelian term costs).

Three consequences follow from Equation (26.72) and none of them is available in the abelian case. First, the gauge generator Equation (26.61) becomes \(G[\Lambda]=\int\left[c^{-1}\dot{\Lambda}^{a}\pi^{0}_{a} -\Lambda^{a}\mathcal{G}^{a}\right]\) with \(\Lambda\) colour-valued, and the transformation it generates is Equation (102.23) — so the classical Hamiltonian analysis and the Lagrangian gauge principle Theorem 102.19 agree, which is the check on the whole construction. Second, the Faddeev–Popov determinant of Proposition 26.24 is no longer field independent: for a linear gauge condition \(\psi^{a}=\pp_{i}A^{a}_{i}\) the matrix \(\pb{\mathcal{G}^{a}}{\psi^{b}}\) is \(\pp_{i}D_{i}\) rather than \(\nabla^{2}\), which depends on \(A\) and therefore cannot be dropped as a constant — this is exactly the classical input to [Faddeev:1967] and the reason ghosts are unavoidable in Quantum Chromodynamics and absent from abelian Quantum Electrodynamics and Renormalization. Third, \(\pp_{i}D_{i}\) has zero modes for large enough \(A\), which is Remark 26.27 said in the canonical language.

Remark 26.39 (The count against the evidence).

Two polarizations per generator is a testable statement, and the two places it has been tested are the eight gluons of Quantum Chromodynamics — whose count enters the running of \(\alpha_{s}\) through the coefficient of the beta function, measured over three decades of energy — and the electroweak sector, where the three broken generators acquire mass and their third polarization along with it (Electroweak Unification and the Higgs Boson), while the photon keeps two. Neither count is postulated; both follow from Proposition 26.37 and its massive counterpart Proposition 26.34, and both are what is observed.

General relativity

General relativity is a constrained system of the most extreme kind: its canonical Hamiltonian is a sum of constraints and therefore vanishes weakly, exactly as for the relativistic particle of Proposition 26.29 and for the same reason — the invariance group contains reparametrizations of the evolution parameter. The decomposition that exhibits this is due to Arnowitt, Deser and Misner [Arnowitt:1962].

Definition 26.40 (The $3+1$ decomposition).

Foliate spacetime by spacelike surfaces \(t=\text{const}\) and write, with \(x^{0}=ct\) and in the signature of Minkowski Space and Its Symmetries,

\begin{equation}\tag{26.78} \dd s^{2}=N^{2}\left(\dd x^{0}\right)^{2} -h_{ij}\left(\dd x^{i}+N^{i}\dd x^{0}\right) \left(\dd x^{j}+N^{j}\dd x^{0}\right)\ec \end{equation}

with \(h_{ij}\) the positive-definite Riemannian metric of the surface — so that \(h_{ij}\dd x^{i}\dd x^{j}\) is squared proper distance, and the metric \(g_{ij}\) induced by Equation (26.78) is \(-h_{ij}\) in this signature — and with \(N\), \(N^{i}\) dimensionless. \(N\) is the lapse, measuring proper time per unit \(x^{0}\) along the normal, and \(N^{i}\) the shift, measuring how the spatial coordinates slide from one surface to the next. Neither is a property of the geometry: both encode the choice of foliation and of coordinates on it. The extrinsic curvature of the surface is

\begin{equation}\tag{26.79} K_{ij}=\frac{1}{2N}\left(\pp_{0}h_{ij} -D_{i}N_{j}-D_{j}N_{i}\right)\ec \end{equation}

with \(N_{i}=h_{ij}N^{j}\) and \(D_{i}\) the Levi-Civita connection of \(h_{ij}\); it carries dimension \(/\mathrm{m}\) and is the only place the time derivative of the metric enters. Rests on Definition 43.2 and Notation 26.28.

Proposition 26.41 (The ADM form of the Einstein–Hilbert Lagrangian).

With Equation (26.78) one has \(\sqrt{\abs{g}}=N\sqrt{h}\) and

\begin{equation}\tag{26.80} \sqrt{\abs{g}}\,R =N\sqrt{h}\left({}^{(3)}\!R+K_{ij}K^{ij}-K^{2}\right) +\text{total derivatives}\ec \end{equation}

where \({}^{(3)}\!R\) is the Ricci scalar of \(h_{ij}\) and \(K=h^{ij}K_{ij}\). The Lagrangian of Equation (43.4) therefore contains no time derivative of \(N\) or of \(N^{i}\). Rests on Definition 26.40 and Equation (43.4).

Derives Proposition 26.41.

The route taken there is the Gauss–Codazzi decomposition of the four-dimensional Riemann tensor: its projections onto the leaf and onto the unit normal give the Gauss equation, which relates the wholly tangential projection to the intrinsic curvature and to a quadratic in \(K_{ij}\), and the Codazzi equation for the mixed projection; contracting twice assembles Equation (26.80) together with two divergences, one of which is the term whose subtraction is the Gibbons–Hawking–York boundary action. The point of care specific to this treatise is the sign bookkeeping between the mostly-minus signature of Minkowski Space and Its Symmetries and the positive-definite induced metric of Definition 26.40, which is settled there explicitly rather than inherited from a mostly-plus source. Only the two facts displayed above are used in this chapter: that \(\sqrt{\abs{g}}=N\sqrt{h}\) and that no time derivative of \(N\) or \(N^{i}\) survives.

Theorem 26.42 (The Hamiltonian structure of general relativity).

The momenta conjugate to the lapse and the shift vanish identically,

\begin{equation}\tag{26.81} \pi_{N}\approx0\ec\qquad\pi_{i}\approx0\ec \end{equation}

four primary constraints per point of space, while

\begin{equation}\tag{26.82} \pi^{ij}=\pdv{\Lag}{\left(\pp_{0}h_{ij}\right)} =\frac{\sqrt{h}}{2\kappa}\left(K^{ij}-Kh^{ij}\right)\ec\qquad \kappa=\frac{8\pi G}{c^{4}}\ec \end{equation}

is unconstrained. The canonical Hamiltonian is a sum of constraints,

\begin{equation}\tag{26.83} H_{\text{c}}=\int\dd^{3}x \left(N\,\Ham_{\perp}+N^{i}\,\Ham_{i}\right) +\text{boundary terms}\ec \end{equation}

with

\begin{align} \Ham_{\perp}&=\frac{2\kappa}{\sqrt{h}} \left(\pi_{ij}\pi^{ij}-\tfrac{1}{2}\pi^{2}\right) -\frac{\sqrt{h}}{2\kappa}\,{}^{(3)}\!R\ec \tag{26.84}\\ \Ham_{i}&=-2\,D_{j}\pi^{j}{}_{i}\ec \tag{26.85} \end{align}

so that consistency of Equation (26.81) yields the four secondary constraints \(\Ham_{\perp}\approx0\) and \(\Ham_{i}\approx0\). All eight are first class, the canonical Hamiltonian is weakly zero, and the count per point of space is \(2\times10-2\times8=4\): two physical degrees of freedom. Rests on Proposition 26.41, Proposition 26.10 and Theorem 26.17.

Proof.

Derives Theorem 26.42. By Proposition 26.41 the Lagrangian \(L=\left(2\kappa\right)^{-1}\int\dd^{3}x\,N\sqrt{h} \left({}^{(3)}\!R+K_{ij}K^{ij}-K^{2}\right)\) contains \(\pp_{0}N\) and \(\pp_{0}N^{i}\) nowhere, so their conjugate momenta vanish identically and Equation (26.81) are primary constraints in the sense of Proposition 26.5. The ten components of \(h_{ij}\) and of \(N,N^{i}\) make \(2n=2\times10\) per point.

For \(\pi^{ij}\), only \(K\) carries \(\pp_{0}h_{ij}\), and by Equation (26.79) \(\pp K_{ij}/\pp\left(\pp_{0}h_{kl}\right) =\delta^{(kl)}_{(ij)}/2N\), so

\begin{equation}\tag{26.86} \pi^{kl}=\frac{N\sqrt{h}}{2\kappa}\cdot\frac{1}{N} \left(K^{kl}-Kh^{kl}\right)\ec \end{equation}

which is Equation (26.82). Its trace is \(\pi=-\sqrt{h}\,K/\kappa\), so the relation inverts to \(K^{ij}=2\kappa h^{-1/2}\left(\pi^{ij}-\tfrac{1}{2}\pi h^{ij}\right)\) and

\begin{equation}\tag{26.87} K_{ij}K^{ij}-K^{2} =\frac{4\kappa^{2}}{h} \left(\pi_{ij}\pi^{ij}-\tfrac{1}{2}\pi^{2}\right)\ec \end{equation}

by direct substitution and collection of the two \(\pi^{2}\) terms. Now Legendre transform. Using \(\pp_{0}h_{ij}=2NK_{ij}+D_{i}N_{j}+D_{j}N_{i}\) from Equation (26.79),

\begin{equation}\tag{26.88} \pi^{ij}\pp_{0}h_{ij}-\Lag =\frac{N\sqrt{h}}{2\kappa}\left(K_{ij}K^{ij}-K^{2}\right) -\frac{N\sqrt{h}}{2\kappa}\,{}^{(3)}\!R +2\pi^{ij}D_{i}N_{j}\ec \end{equation}

the first term because \(2N\pi^{ij}K_{ij}\) is twice the corresponding term of \(\Lag\). Substituting Equation (26.87) into the first term gives \(N\Ham_{\perp}\) with \(\Ham_{\perp}\) as in Equation (26.84); integrating the last term by parts gives \(-2N_{i}D_{j}\pi^{ji}\) plus a surface term, that is \(N^{i}\Ham_{i}\) with \(\Ham_{i}\) as in Equation (26.85). That is Equation (26.83).

Since \(N\) and \(N^{i}\) appear in Equation (26.83) linearly and without derivatives, they are multipliers, and consistency of Equation (26.81) reads \(\dot{\pi}_{N}=-\delta H_{\text{c}}/\delta N=-\Ham_{\perp}\approx0\) and \(\dot{\pi}_{i}=-\Ham_{i}\approx0\): four secondary constraints, outcome 2 of Proposition 26.10. That they are first class is Equation (26.89) below, whose right-hand sides are combinations of the constraints; the primary constraints commute with everything because no \(\Ham\) contains \(N\) or \(N^{i}\). Hence \(F=8\), \(S=0\), \(2n=20\), and Equation (26.21) gives \(20-16=4\): two configuration degrees of freedom per point of space. Finally, substituting the secondary constraints into Equation (26.83) gives \(H_{\text{c}}\approx0\) up to the boundary terms of Remark 26.45.

Proposition 26.43 (The hypersurface-deformation algebra).

Smearing \(\Ham_{\perp}\) with a scalar \(f\) and \(\Ham_{i}\) with a vector \(\xi^{i}\), and writing \(H_{\perp}[f]=\int f\Ham_{\perp}\), \(H_{\parallel}[\xi]=\int\xi^{i}\Ham_{i}\), the constraints close as

\begin{align} \pb{H_{\parallel}[\xi]}{H_{\parallel}[\zeta]} &=H_{\parallel}\!\left[\comm{\xi}{\zeta}\right]\ec \tag{26.89}\\ \pb{H_{\parallel}[\xi]}{H_{\perp}[f]} &=H_{\perp}\!\left[\xi^{i}\pp_{i}f\right]\ec \tag{26.90}\\ \pb{H_{\perp}[f]}{H_{\perp}[g]} &=H_{\parallel}\!\left[h^{ij} \left(f\pp_{j}g-g\pp_{j}f\right)\right]\ep \tag{26.91} \end{align}

The first two say that the momentum constraints generate spatial diffeomorphisms and that \(\Ham_{\perp}\) is a scalar density under them. The third is the one that matters: its coefficient is \(h^{ij}\), a function on phase space, so the first-class algebra of general relativity has structure functions and is not a Lie algebra. Rests on Theorem 26.42 and Proposition 26.13.

Derives Proposition 26.43.

Each bracket is a direct computation from the densities Equations (26.84) and (26.85), and the third is long: both the kinetic and the intrinsic-curvature terms contribute, and the curvature contribution has to be integrated by parts twice before the surviving terms assemble into \(\Ham_{i}\). The geometric reading, written there beside the computation, is that these brackets reproduce the composition law of deformations of a spacelike surface embedded in a Lorentzian spacetime; the metric appears in the third because a normal deformation at one point becomes a tangential one at a neighbouring point through the geometry of the surface itself, and that is why \(h^{ij}\) — a function on phase space, not a constant — sits where a structure constant would.

Remark 26.44 (Three consequences, and one that is out of scope).

The constraints are the initial-value equations. The four secondary constraints of Theorem 26.42 involve no time derivative of \(\pi^{ij}\): they are conditions on the data \(\left(h_{ij},\pi^{ij}\right)\) laid down on one surface, and they are precisely the \(G^{0}{}_{\mu}=\kappa T^{0}{}_{\mu}\) components of the Einstein equations. Not every initial datum is admissible, and the constraints say which are — this is the content of Section 44.9 and Section 44.9.1, and the well-posedness of the resulting evolution problem is [ChoquetBruhat:1952], with the constraint-solving method of [York:1972].

The count matches the waves. Two degrees of freedom per point is the same answer Proposition 26.32 gave for light, and it is the two polarizations of Gravitational-Wave Theory, detected as such in Experiment: Gravitational Waves. That a theory with ten metric functions propagates two is not obvious from the field equations and is transparent here.

The Hamiltonian vanishes. \(H_{\text{c}}\approx0\) says, by Remark 26.46, that evolution in the coordinate \(t\) is a gauge transformation — a relabelling of the foliation — and not a physical motion. What is physical is the relation between the geometry and the matter in it, exactly as what was physical for the relativistic particle was the worldline and not its parametrization.

What is not claimed. The statement that this vanishing creates a “problem of time” is a statement about a quantum theory of gravity, in which Equation (26.92) would read \(\hat{\Ham}_{\perp}\ket{\psi}=0\) and leave the state vector with no evolution parameter. No such theory has observational support, so it is outside the evidence base of this treatise and is treated only in Quantum Gravity: The Honest Status. Classically there is no problem of time and nothing to solve.

Remark 26.45 (The energy is a boundary term).

The surface terms discarded in Equation (26.83) cannot in fact be discarded. On a region with a boundary, or in an asymptotically flat spacetime, the constraints do not have well-defined functional derivatives unless they are supplemented by surface integrals; the supplement to \(\Ham_{\perp}\) is the ADM mass and the supplement to \(\Ham_{i}\) the ADM momentum and angular momentum [Arnowitt:1962] [Wald:1984]. So the total energy of an isolated gravitating system is a boundary integral at spatial infinity and not the volume integral of any density: the numerical value of the Hamiltonian is carried entirely by the surface term, the bulk being weakly zero. This is the constrained-dynamics origin of the ADM integrals referred to in The Einstein Field Equations, and it is why the localization of gravitational energy is not a well-posed question.

Remark 26.46 (A vanishing Hamiltonian, again).

Remark 26.31 was stated for the relativistic particle and applies verbatim here, with the reparametrizations of the worldline replaced by the deformations of the spacelike surface. The two systems are the same phenomenon at different sizes, which is why the particle is worth doing first.

Quantizing a constrained system

Postulate 26.47 (Dirac quantization).

Quantize a constrained system by

  1. eliminating the second-class constraints, replacing the Poisson bracket by the Dirac bracket Equation (26.25) and imposing \(\chi_{\alpha}=0\) strongly; then applying the correspondence rule Equation (25.31) to the Dirac bracket, \(\pb{A}{B}_{\text{D}}\mapsto \tfrac{1}{\ii\hbar}\comm{\hat{A}}{\hat{B}}\); and

  2. imposing the first-class constraints as conditions on the states rather than on the operators,

    \begin{equation}\tag{26.92} \hat{\gamma}_{A}\ket{\psi}=0\ec \end{equation}

    the physical Hilbert space being the space of solutions.

Rests on Postulate 25.27, Theorem 26.21 and Theorem 26.14.

The asymmetry between the two clauses is the whole design and is forced by Theorem 26.21. A second-class constraint has vanishing Dirac bracket with everything, so it may be set to zero as an operator identity without contradicting any commutator; a first-class constraint may not, because it generates a nontrivial transformation and setting it to zero as an operator would set that generator to zero too. Imposing it on states instead selects the gauge-invariant sector, which is exactly what Theorem 26.14 says the physical states are.

Proposition 26.48 (Consistency of the state conditions).

Equation (26.92) is consistent only if the quantum constraint algebra closes,

\begin{equation}\tag{26.93} \comm{\hat{\gamma}_{A}}{\hat{\gamma}_{B}} =\ii\hbar\,\hat{f}_{AB}{}^{C}\hat{\gamma}_{C}\ec \end{equation}

with the constraint operators standing to the right. A failure of Equation (26.93) is an anomaly: a symmetry of the classical theory that no quantization preserves. Rests on Postulate 26.47 and Equation (26.16).

Proof.

Derives Proposition 26.48. Let \(\ket{\psi}\) satisfy Equation (26.92) for every \(A\). Applying the commutator to it, \(\comm{\hat{\gamma}_{A}}{\hat{\gamma}_{B}}\ket{\psi} =\hat{\gamma}_{A}\hat{\gamma}_{B}\ket{\psi} -\hat{\gamma}_{B}\hat{\gamma}_{A}\ket{\psi}=0\), so the commutator must annihilate every physical state. The classical bracket is \(f_{AB}{}^{C}\gamma_{C}\) by Equation (26.16), and its operator image annihilates \(\ket{\psi}\) if and only if every \(\hat{\gamma}_{C}\) stands to the right of the corresponding \(\hat{f}_{AB}{}^{C}\): the opposite ordering leaves \(\hat{\gamma}_{C}\hat{f}_{AB}{}^{C}\ket{\psi} =\comm{\hat{\gamma}_{C}}{\hat{f}_{AB}{}^{C}}\ket{\psi}\), which need not vanish. Where the structure functions are constants the question does not arise, which is why Equation (26.72) is easier to quantize than Equation (26.91). If no operator ordering makes Equation (26.93) hold, the conditions Equation (26.92) are mutually inconsistent and only \(\ket{\psi}=0\) satisfies them.

This is the ordering ambiguity of Definition 25.42 seen from the other end. There, the choice among \(\hat{q}\hat{p}\), \(\hat{p}\hat{q}\) and their symmetrizations was underdetermined by the correspondence rule and had to be fixed by experiment (Remark 25.49); here the requirement that a gauge symmetry survive quantization fixes part of the same choice, and Theorem 25.38 guarantees that it cannot always be fixed consistently. An anomaly is precisely the case in which it cannot.

Remark 26.49 (Anomalies are an experimental constraint).

An anomaly in a global symmetry is a prediction, and a correct one: the observed decay of the neutral pion to two photons is the textbook case [Adler:1969] [Bell:1969], and its rate counts the colours (Proposition 102.16). An anomaly in a gauge symmetry is fatal, because Equation (26.92) then has no solutions and the theory has no physical states. That the gauge anomalies of the Standard Model cancel between quarks and leptons, generation by generation, is therefore not an aesthetic observation but a consistency condition that the observed particle content satisfies and that constrains any proposed addition to it (Electroweak Unification and the Higgs Boson and Discrete Symmetries and CPT). The cancellation is not only of the triangle kind: a global obstruction also exists, and it requires the number of left-handed \(\SU(2)\) doublets to be even [Witten:1982] — which the observed generations satisfy, one quark doublet in three colours plus one lepton doublet making four.

Reduced-phase-space quantization

The alternative to Postulate 26.47 is to do all the classical work first: fix the gauge as in Definition 26.23, construct the reduced phase space of dimension \(2n-2F-S\), take its Dirac bracket, and only then apply the correspondence rule. The two routes are often called constrain-then-quantize and quantize-then-constrain, and the honest statement is that they are not equivalent.

They differ in three ways, each of which is a known difficulty rather than a technicality.

There is no principle in this treatise that decides between them. Where both have been carried out and compared against measurement they agree on what has been measured: for quantum electrodynamics the Coulomb-gauge reduced construction and the covariant construction of Section 26.7.2 give the same scattering amplitudes [Weinberg:1995], and the agreement is checked to the precision of the electron anomalous moment (Experiment: The Electron Anomalous Magnetic Moment). That is an experimental fact about one theory, not a theorem about the two procedures, and it should not be read as one.

Weak imposition: the Gupta–Bleuler construction

For quantum electrodynamics in a manifestly covariant gauge the condition Equation (26.92) cannot be imposed as it stands, and the reason is visible already in Proposition 26.32: the primary constraint is \(\pi^{0}\approx0\), so a covariant commutator \(\comm{\hat{A}_{\mu}(\vect{x})}{\hat{\pi}^{\nu}(\vect{y})} =\ii\hbar\,\delta_{\mu}^{\nu}\delta^{3}(\vect{x}-\vect{y})\) contradicts it outright. The covariant route therefore does not impose the constraint at all. Instead it adds to Equation (26.54) a gauge-fixing term

\begin{equation}\tag{26.94} \Lag_{\text{gf}} =-\frac{1}{2\xi\mu_{0}}\left(\pp_{\mu}A^{\mu}\right)^{2}\ec \end{equation}

which makes \(\pi^{0}=-\left(\xi\mu_{0}c\right)^{-1}\pp_{\mu}A^{\mu}\) nonzero. The Lagrangian is then non-singular: there are no constraints, all four components of \(A_{\mu}\) propagate, and the canonical formalism of Hamiltonian Mechanics applies unchanged. The price is paid twice over. The time component has a commutator of the opposite sign to the others, so the Fock space carries an indefinite metric and is not a Hilbert space; and the theory has four polarizations where Proposition 26.32 counted two.

The Gupta–Bleuler construction recovers both [Gupta:1950] [Bleuler:1950]. The Lorenz condition \(\pp_{\mu}\hat{A}^{\mu}\ket{\psi}=0\) is too strong — no state in the Fock space satisfies it — so only its positive-frequency part is required,

\begin{equation}\tag{26.95} \left(\pp_{\mu}\hat{A}^{\mu}\right)^{(+)}\ket{\psi}=0\ec \end{equation}

which is the annihilation half and which does have solutions. The subspace \(\mathcal{V}_{\text{phys}}\) they span has a positive semi-definite norm: the timelike and longitudinal excitations survive only in the combination that has zero norm, because their opposite-sign contributions cancel exactly. Quotienting by that null subspace,

\begin{equation}\tag{26.96} \mathcal{H}_{\text{phys}} =\mathcal{V}_{\text{phys}}/\mathcal{V}_{0}\ec \end{equation}

gives a genuine Hilbert space in which only the two transverse polarizations remain — Proposition 26.32's count, recovered. Here and in Section 26.7.3 the symbol \(\mathcal{H}_{\text{phys}}\) is a Hilbert space and not a Hamiltonian density, the one place in this chapter where the script letter carries its other standard meaning. Zero-norm states do not contribute to any expectation value, so observables are insensitive to the choice of representative, and the gauge parameter \(\xi\) cancels from every physical amplitude, which is the standard check performed in Quantum Electrodynamics and Renormalization. The construction is developed with the field operators in Canonical Quantization of Fields.

The essential limitation is that Equation (26.95) splits the constraint operator into positive- and negative-frequency parts, and that split is available only for a free field — a linear equation of motion. For a non-abelian theory the constraint algebra Equation (26.72) is nonlinear in the fields, no such split exists, and the Gupta–Bleuler condition has no non-abelian generalization. That failure is what the next subsection exists to repair.

BRST

The systematic replacement of the constraint conditions Equation (26.92) is due to Becchi, Rouet and Stora and, independently, to Tyutin [Becchi:1976] [Tyutin:1975]. It trades the \(F\) separate conditions for a single one, at the price of enlarging the phase space.

Definition 26.50 (The BRST charge).

For each first-class constraint \(\gamma_{A}\) adjoin a Grassmann-odd pair \(\left(\eta^{A},\mathcal{P}_{A}\right)\) — the ghost and its momentum — canonically conjugate under the graded bracket, \(\pb{\eta^{A}}{\mathcal{P}_{B}}=-\delta^{A}_{B}\), and carrying ghost numbers \(+1\) and \(-1\). Under quantization the pair becomes an anticommutator, \(\acomm{\hat{\eta}^{A}}{\hat{\mathcal{P}}_{B}} =-\ii\hbar\,\delta^{A}_{B}\); the overall sign is a convention, fixed by the sign chosen for \(\mathcal{P}\), and no statement below depends on it. The BRST charge is the Grassmann-odd function of ghost number \(+1\)

\begin{equation}\tag{26.97} \Omega=\eta^{A}\gamma_{A} -\tfrac{1}{2}\eta^{B}\eta^{A}f_{AB}{}^{C}\mathcal{P}_{C} +\cdots\ec \end{equation}

determined by the requirement that it be nilpotent,

\begin{equation}\tag{26.98} \pb{\Omega}{\Omega}=0\ec\qquad\text{quantum mechanically}\quad \hat{\Omega}^{2}=0\ec \end{equation}

which is a nontrivial condition because the bracket of two odd functions is symmetric. The displayed terms suffice when the structure functions of Equation (26.16) are constants; the omitted terms, of higher order in the ghosts, are needed exactly when they are not. Rests on Equation (26.16) and Postulate 26.47.

That such an \(\Omega\) exists at all, for every first-class set and not merely for the ones whose algebra happens to close on constants, is the theorem that makes BRST a general method rather than a device.

Theorem 26.51 (Existence and uniqueness of the BRST charge).

Let \(\gamma_{A}\) be a regular and irreducible set of first-class constraints in the sense of Remark 26.7, with the algebra Equation (26.16). Then the series Equation (26.97) can be continued, order by order in the ghosts, to a Grassmann-odd function \(\Omega\) of ghost number \(+1\) satisfying Equation (26.98) exactly; and any two such functions with the same leading term are related by a canonical transformation of the extended phase space preserving the ghost number. When the structure functions \(f_{AB}{}^{C}\) are constants the series terminates at the two displayed terms. Rests on Definition 26.50, Equation (26.16) and Remark 26.7.

Derives Theorem 26.51.

The construction there is by induction on the ghost number. Nilpotency at each order is an equation of the form \(\delta\,\Omega_{(k)}= \text{(known, built from lower orders)}\), where \(\delta\) is the Koszul–Tate differential of the constraint surface; the obstruction to solving it is a class in the homology of \(\delta\) at positive degree, which vanishes precisely because the constraints are regular and irreducible — that is what makes the Koszul–Tate complex a resolution. Uniqueness is the same statement one degree down.

The point of Equation (26.98) is that a nilpotent operator has a cohomology, and

\begin{equation}\tag{26.99} \mathcal{H}_{\text{phys}} =\frac{\ker\hat{\Omega}}{\im\hat{\Omega}} \quad\text{at ghost number zero} \end{equation}

reproduces the physical state space of Postulate 26.47: the kernel condition \(\hat{\Omega}\ket{\psi}=0\) contains Equation (26.92) at lowest ghost order, and the quotient by the image removes exactly the zero-norm states that Equation (26.96) removed by hand in the abelian case. Gupta–Bleuler is therefore the abelian shadow of Equation (26.99), and the ghosts that appear here as bookkeeping variables are the same objects that Proposition 26.24 produced as the exponentiation of a Jacobian.

Three things make this the practical form in which the content of this chapter enters a calculation. The gauge-fixed action is \(S_{\text{gf}}=S+\pb{\Omega}{\Psi}\) for a gauge-fixing function \(\Psi\) of ghost number \(-1\), and nilpotency makes physical amplitudes independent of \(\Psi\) — which is why a gauge may be chosen for convenience and why the \(\xi\) of Equation (26.94) cancels. The Ward identities of the abelian theory generalize to the Slavnov–Taylor identities, which is what makes the renormalization of Quantum Chromodynamics preserve gauge invariance order by order. And unitarity of the physical subspace follows from Equation (26.99) rather than being assumed, so the unphysical polarizations and the ghosts cancel in every cut. The path-integral development, where this is done in practice, is Path-Integral Quantization.

Constraints in field theory

Everything above was written for finitely many degrees of freedom. The passage to a field theory is not automatic, and the places where it is not are recorded here so that Parts VII and XI can refer to them rather than rediscover them.