Canonical Quantization of Fields
This chapter quantizes a classical field the way The Poisson Algebra and the Canonical Bridge to Quantum Mechanics quantized a mechanical system: a field is a mechanical system with one degree of freedom at every point of space, so the canonical procedure — conjugate momenta, equal-time brackets promoted to operators — applies unchanged, and the free field turns out to be an infinite collection of the harmonic oscillators of Oscillations and Mechanical Waves. Everything particle-like is then a statement about those oscillators: a particle is a quantum of excitation, the number of particles is an eigenvalue, and processes that change that number are describable at all only because the field, not the particle, is the primitive object. That is the reason this chapter opens the part rather than The Klein–Gordon Equation or The Dirac Equation, whose single-particle relativistic wave equations both fail — negative energies, negative probabilities, the Klein paradox — for exactly the want of a field. The classical input is the Lagrangian field theory of Generalized Classical Field Theory and the Poincaré representation theory of Particles as Poincaré Representations; the output is Fock space, propagators, microcausality and the spin–statistics connection, on which Quantum Electrodynamics and Renormalization and everything after it is built.
The free-field results below are exact and among the very few in this part that are. Interactions are treated perturbatively from Quantum Electrodynamics and Renormalization onward, and the path-integral alternative — indispensable for the non-abelian theories of Quantum Chromodynamics — is developed in Path-Integral Quantization; the axiomatic reformulation that makes the causality statements below precise is Axiomatic Quantum Field Theory. Standard modern treatments are [Peskin:1995] [Weinberg:1995] [Itzykson:1980].
The conventions are those fixed once for the whole part in Section 100.1.1 and are repeated here only in the two places where a scalar field needs them. The metric is \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\); coordinates are \(x^{\mu}=(ct,\vect{x})\), so that \(\dd^{4}x=c\,\dd t\,\dd^{3}x\) carries \(\mathrm{m}^{4}\) and \(\pp_{\mu}=(c^{-1}\pp_{t},\nabla)\) carries \(/\mathrm{m}\); the d'Alembertian is
Wave four-vectors \(k^{\mu}=(\omega/c,\vect{k})\) carry \(/\mathrm{m}\) and four-momenta \(p^{\mu}=\hbar k^{\mu}\) carry \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), with the mass shell \(p^{2}=m^{2}c^{2}\). A mass \(m\) enters every field equation below only through the inverse length
the reciprocal of the reduced Compton wavelength; the mass shell reads \(k^{2}=\kappa^{2}\). The letter \(\kappa\) never means anything else in this chapter, and \(\mu\) never means a mass: it occurs only as a Lorentz index and in \(\mu_{0}\), the magnetic constant. As in Notation 100.1, \(\epsilon\) standing alone is the positive infinitesimal of the Feynman contour, \(\varepsilon_{0}\) the electric constant, and \(\vect{\varepsilon}_{\lambda}\) a polarization vector. An action is \(S=c^{-1}\int\dd^{4}x\,\Lag\) with \(\Lag\) an energy density, \(\mathrm{J}/\mathrm{m}^{3}\); the SI dimension of each field is stated where the field is introduced. Every equation in this chapter, without exception, is written with \(\hbar\) and \(c\) in place.
The literature cited below works in units in which \(\hbar=c=1\), so that \(\kappa\) and \(m\) are the same symbol and a propagator is written \(1/(k^{2}-m^{2})\). The dictionary is the one given in Remark 100.4, applied here to the two objects this chapter introduces: a mass becomes an inverse length by Equation (99.2), and the scalar propagator \(1/(k^{2}-m^{2}+\ii\epsilon)\) of those texts is \(\ii\hbar c/(k^{2}-\kappa^{2}+\ii\epsilon)\) in Equation (99.93) below. No derivation in this chapter is performed in those units; the remark exists so that the reader can check the chapter against its sources, which is a different thing.
Why fields
Quantum field theory is not a generalization undertaken for its own sake. It is forced, and this section is the argument that forces it. The three steps are independent and each is fatal on its own: a relativistic one-particle wave equation has no consistent probability interpretation; a theory with a fixed number of particles contradicts what is seen in every accelerator and every cloud chamber; and the uncertainty principle forbids the localization on which a one-particle position operator depends. What survives all three is a system with infinitely many degrees of freedom whose excitations are counted, not tracked.
Failure of the single-particle relativistic theories
The Schrödinger equation \(\ii\hbar\,\pp_{t}\psi =-(\hbar^{2}/2m)\nabla^{2}\psi\) is first order in time and second order in space, so it cannot be covariant under a transformation that mixes the two. The obvious repair is to quantize the relativistic mass shell \(E^{2}=\vect{p}^{2}c^{2}+m^{2}c^{4}\) instead of the Newtonian \(E=\vect{p}^{2}/2m\), substituting \(E\to\ii\hbar\pp_{t}\) and \(\vect{p}\to-\ii\hbar\nabla\). The result is the Klein–Gordon equation [Gordon:1926],
studied as a wave equation in The Klein–Gordon Equation. It is manifestly covariant, and it is manifestly useless as a one-particle theory, for two reasons that turn out to be one reason.
Negative energies. The plane wave \(\psi=N\exp[-\ii(Et-\vect{p}\cdot\vect{x})/\hbar]\) solves Equation (99.3) whenever \(E=\pm\sqrt{\vect{p}^{2}c^{2}+m^{2}c^{4}}\), and the negative branch cannot be thrown away. Equation (99.3) is second order in time, so the initial data are \(\psi\) and \(\pp_{t}\psi\), freely specifiable; the positive-frequency solutions alone are not a complete set and cannot match arbitrary initial data. Nor is the restriction stable: any interaction that is not diagonal in energy couples the two branches, which is what the Klein paradox [Klein:1929b] exhibits explicitly for a step potential of height greater than \(2mc^{2}\).
Non-positive density. Multiplying Equation (99.3) by \(\psi^{*}\), subtracting the conjugate equation multiplied by \(\psi\), and dividing by \(2\ii mc^{2}/\hbar\) gives a continuity equation \(\pp_{t}\rho+\nabla\cdot\vect{j}=0\) with
Evaluated on the plane wave above, \(\pp_{t}\psi=-(\ii E/\hbar)\psi\) and
which is negative on the negative branch. Born's reading of \(\abs{\psi}^{2}\) as a probability density [Born:1926a] therefore cannot be transferred to Equation (99.4): a probability density is non-negative by definition, and this one is not.
Dirac's response was to insist on an equation first order in time, so that the density would be \(\psi^{\dagger}\psi\ge0\) by construction [Dirac:1928]. That succeeded — The Dirac Equation follows the construction — and it did not remove the negative energies, which are a consequence of the mass shell and not of the order of the equation. The Dirac Hamiltonian has spectrum \(\pm\sqrt{\vect{p}^{2}c^{2}+m^{2}c^{4}}\) and is therefore unbounded below: an atomic electron could cascade down the negative continuum without limit, radiating an unbounded amount of energy, and no state of the theory would be stable.
The hole theory [Dirac:1930a] repairs this by fiat. Declare that in the vacuum every negative-energy level is occupied; the exclusion principle then forbids the cascade, and the absence of an electron from the filled sea behaves as a particle of opposite charge. What that particle's mass is the 1930 paper does not settle: Dirac there identified the hole with the proton [Dirac:1930a], and only in 1931 [Dirac:1931] concluded that it must carry the electron's own mass and so be a new particle unknown to experiment. That prediction, and its confirmation by Anderson [Anderson:1933] (Experiment: The Positron), is one of the great successes of theoretical physics. It is nevertheless untenable as a foundation, for three reasons.
-
It is unavailable for bosons. The sea is held down by the exclusion principle, and a spin-\(0\) particle obeys no such principle; there is no way to fill the negative-energy states of Equation (99.3). Yet the observed spin-\(0\) particles — the charged pions above all — do have antiparticles: \(\pi^{+}\) and \(\pi^{-}\) are listed as a particle–antiparticle pair of equal mass by the Particle Data Group [Navas:2024].
-
The vacuum acquires infinite charge and energy density, removed by a subtraction whose only justification is that it is needed.
-
It is a one-particle theory only in name. A state of the sea plus one electron is a state of infinitely many particles; the formalism has already left the single-particle Hilbert space, without admitting it.
The consistent version of the repair was found by Pauli and Weisskopf [Pauli:1934]: reinterpret Equation (99.3) not as an equation for a wave function but as an equation for a classical field, and quantize the field. Then \(\rho\) of Equation (99.4) is not a probability density at all but a charge density, whose sign is the sign of the charge, and the energy comes out positive for both signs. That construction is Section 99.3 below, and it works for spin \(0\) and spin \(1/2\) alike, with no sea.
The decisive objection is not internal but experimental. A one-particle theory has a Hilbert space \(L^{2}(\R^{3})\) (tensored with a finite-dimensional spin space), on which the number of particles is the constant \(1\); no unitary evolution on that space can produce a two-particle state. Nature produces them routinely.
A photon traversing matter converts into an electron–positron pair in the Coulomb field of a nucleus, provided its energy exceeds
where \(M\) is the nuclear mass. Pair production is the dominant photon-absorption process in lead above about \(5\,\mathrm{MeV}\), and the positrons so produced are the ones Anderson identified [Anderson:1933]. Rests on Definition 40.4 and Theorem 40.5.
Derivation. Derives Phenomenon 99.3. Energy and momentum conservation are the statement that the total four-momentum before and after are equal, so the invariant \(s=(p_{\gamma}+P)^{2}\) is the same before and after. Before: the photon has \(p_{\gamma}^{2}=0\) and the nucleus, at rest, has \(P^{\mu}=(Mc,\vect{0})\), so
After: three particles, and the smallest value \(s\) can take for a given set of masses is the one in which they are all at relative rest, \(s\ge(M+2m_{e})^{2}c^{2}\). Combining,
With the CODATA 2022 electron rest energy \(m_{e}c^{2}=0.51099895069\,\mathrm{MeV}\) [Mohr:2025] the infinite-mass limit is \(1.02199790\,\mathrm{MeV}\). The recoil term \(m_{e}/M\) is why a nucleus must be present at all: for a photon alone \(s=0\), and Equation (99.8) can never be satisfied, so pair creation in empty space is kinematically forbidden however energetic the photon.
∎A theory in which the number of particles is a fixed integer cannot describe Phenomenon 99.3, and no reinterpretation of the wave function repairs that. The state before the conversion has one quantum, the state after has three; the two live in different sectors, and only a formalism whose state space contains both, with operators connecting them, can hold the process at all.
The last objection is that even the one-particle sector is not what a single-particle theory needs it to be, because the position of a relativistic particle is not measurable to arbitrary precision without creating others.
Suppose a particle of mass \(m\) is confined to a region of linear size \(\Delta x\). The uncertainty principle then requires \(\Delta p\gtrsim\hbar/(2\Delta x)\), and for a relativistic particle \(\Delta E\approx c\,\Delta p\gtrsim\hbar c/(2\Delta x)\). The energy uncertainty reaches the threshold \(2mc^{2}\) of Equation (99.6) when
that is, at a fraction of the reduced Compton wavelength \(\kappa^{-1}=\hbar/mc\). For the electron the Compton wavelength is \(\lambda_{C}=h/(m_{e}c)=2.42631023538\times 10^{-12}\,\mathrm{m}\) and its reduced form is
both from CODATA 2022 [Mohr:2025]. Attempting to pin an electron down to a tenth of that supplies enough energy to make several more, and the apparatus that would do the pinning — a potential well of depth \(\gtrsim mc^{2}\) — is precisely a device for creating pairs. There is therefore no operational meaning to “the position of the electron” below \(\kappa_{e}^{-1}\), and any formalism that carries an exact position operator is claiming more than can be measured.
Newton and Wigner [Newton:1949] asked how much of a position operator survives. Demanding only that the localized states at a fixed time transform correctly under rotations, translations and space inversion, and that they be orthogonal for distinct points, they showed that a unique such operator exists on the positive-energy one-particle space — but that its eigenstates are not localized in any Lorentz-invariant sense. A Newton–Wigner state localized at a point in one frame has support spread over a region of order \(\kappa^{-1}\) in another, and its wave function has tails falling as \(\ee^{-\kappa r}\) rather than vanishing outside a bounded set. The one-particle position operator exists; what does not exist is a frame-independent notion of where the particle is, to better than \(\kappa^{-1}\).
Each objection is really the same one seen from a different side. The mass shell has two branches; the second branch cannot be removed covariantly; a theory that keeps it needs states of arbitrarily negative energy unless the number of quanta is allowed to change; and the scale at which the change becomes unavoidable is \(\kappa^{-1}\), which is also the scale at which localization fails. The resolution is to stop treating the wave equation as an equation for a probability amplitude and to treat it as the classical equation of motion of a field, to be quantized. Then the two branches become creation and annihilation of two species of quantum, the energy is bounded below, the conserved quantity of Equation (99.4) is a charge that may take either sign, and states with any number of quanta live in one space. Everything after this section is the execution of that programme.
Field as a mechanical system with continuously many degrees of freedom
A field is not a new kind of object requiring new mechanics. It is an ordinary mechanical system with an infinite number of coordinates labelled by a continuous index, and the cleanest way to see this is to watch the index become continuous.
Take \(N\) equal masses \(m_{0}\) threaded on a light string of tension \(T\) at spacing \(\ell\), with transverse displacements \(q_{n}(t)\) and fixed ends, the discrete system of Oscillations and Mechanical Waves. For small displacements the string segment between neighbours pulls with the transverse component \(T(q_{n+1}-q_{n})/\ell\), so
which follows from the Lagrangian
Now let \(\ell\to0\) at fixed linear density \(\varrho=m_{0}/\ell\) and fixed \(T\). Write \(q_{n}(t)=u(t,x_{n})\) with \(x_{n}=n\ell\); then \((q_{n+1}-q_{n})/\ell\to\pp_{x}u\) and \(\sum_{n}\to \ell^{-1}\int\dd x\), so
whose Euler–Lagrange equation is the wave equation \(\pp_{t}^{2}u=v^{2}\pp_{x}^{2}u\) with \(v=\sqrt{T/\varrho}\). Nothing conceptual happened in the limit. The discrete label \(n\) became the continuous label \(x\); the finitely many coordinates \(q_{n}\) became the one-parameter family \(u(x)\); and the Lagrangian, previously a sum, became an integral of a density. The coordinate of the system is now the whole function \(u(\cdot)\) at one instant, and \(x\) is not a dynamical variable but a name for which coordinate one is talking about.
A (real, scalar) classical field is a map \(\phi:\R^{4}\to\R\), \(x^{\mu}\mapsto\phi(x)\), regarded as the continuously indexed set of coordinates \(\set{\phi(t,\vect{x})}\) labelled by \(\vect{x}\). A local Lagrangian is the integral of a density depending on the field and its first derivatives at one point,
and the momentum density conjugate to \(\phi\) is
The Euler–Lagrange equations follow from stationarity of \(S=c^{-1}\int\dd^{4}x\,\Lag\) under variations vanishing on the boundary of the region of integration. Expanding \(\delta\Lag=(\pp\Lag/\pp\phi)\,\delta\phi +\bigl(\pp\Lag/\pp(\pp_{\mu}\phi)\bigr)\pp_{\mu}\delta\phi\) and integrating the second term by parts leaves, once the boundary term is discarded,
which vanishes for arbitrary \(\delta\phi\) only if the bracket does. For a local density the equations therefore read
The derivative of an integral with respect to one of continuously many coordinates is a functional derivative, defined by
the continuum replacement of \(\pp q_{n}/\pp q_{m}=\delta_{nm}\); the Dirac delta carries \(/\mathrm{m}^{3}\), which is where the volume factors in every formula below come from.
The relativistic field whose equation of motion is Equation (99.3) is obtained from
for which Equation (99.16) gives \(\left(\Box+\kappa^{2}\right)\phi=0\) directly. Since \([\pp_{\mu}]=/\mathrm{m}\) and \([\Lag]=\mathrm{J}/\mathrm{m}^{3}\), the field itself carries
and by Equation (99.15)
so that \([\phi][\pi]=\mathrm{J}\,\mathrm{s}/\mathrm{m}^{3}\) — an action per unit volume, which is what the equal-time commutator of Section 99.2.1 will need it to be. Legendre-transforming,
manifestly non-negative — a first sign that the field reading cures the disease of Section 99.1.1, since the same mass shell that gave energies of both signs gives a Hamiltonian density that is a sum of squares.
Classically the field obeys Poisson brackets that are the continuum form of \(\pb{q_{n}}{p_{m}}=\delta_{nm}\),
obtained by putting the functional derivative Equation (99.17) in place of \(\pp q_{n}/\pp q_{m}=\delta_{nm}\) in the definition of the bracket. Generalized Classical Field Theory is the general classical field theory built on these brackets.
Equation (99.14) contains a physical hypothesis that is easy to overlook because it is built into the notation: the Lagrangian density at \(x\) depends on the field at \(x\) and on its first derivatives there, and on nothing else. A term such as \(\int\dd^{3}y\,K(\vect{x}-\vect{y})\phi(\vect{x})\phi(\vect{y})\) with \(K\) of finite range would be perfectly well defined and would still give a linear equation of motion; it is excluded by assumption.
The assumption is not free of consequences. Locality plus a second-order-in-time equation makes the field equation hyperbolic with characteristic speed \(c\), so classical disturbances propagate on and inside the light cone; after quantization it is what makes the commutator of the field with itself vanish at spacelike separation, the statement verified explicitly in Section 99.7.1. A nonlocal kernel \(K\) of range \(R\) would leave a non-vanishing commutator out to spacelike separations of order \(R\), and thus a signal outside the light cone. The bound on such effects is therefore an experimental bound, and the axiomatic programme of Axiomatic Quantum Field Theory takes locality — not the Lagrangian — as the primitive assumption for exactly this reason.
The historical sequence
The formalism of this chapter was assembled between 1925 and 1932, in a sequence in which almost every step was taken for a reason different from the one that now justifies it.
The first field ever quantized was the free radiation field, in the concluding chapter of the three-author paper of Born, Heisenberg and Jordan of 1926. Their object was not particles: it was to show that the new matrix mechanics reproduced Einstein's 1909 formula for the energy fluctuations of black-body radiation, whose two terms had for fifteen years looked like an unexplained superposition of a wave term and a particle term. Treating each cavity mode as a matrix-mechanical oscillator produced both terms at once. In the same year Born's collision papers [Born:1926a] fixed the probability interpretation of the wave amplitude that Section 99.1.1 showed cannot survive the relativistic generalization — the two ideas were contemporaneous, and it took another eight years to see that they were in tension.
Dirac's radiation theory of 1927 [Dirac:1927] is the paper in which the method became a method. He quantized the electromagnetic field in Coulomb gauge, obtained the ladder operators for each mode, and computed the emission and absorption rates of an atom coupled to it. The matrix element for emission into a mode already containing \(n\) quanta came out proportional to \(\sqrt{n+1}\) and that for absorption to \(\sqrt{n}\); the ratio of the resulting rates reproduces Einstein's \(A\) and \(B\) coefficients, with the \(+1\) — spontaneous emission — coming from the mode's zero-point term and not from any external perturbation. The consequences for quantum optics are Quantum Optics and the Photon.
Jordan and Wigner [Jordan:1928] then showed that the same apparatus describes fermions if the commutators are replaced by anticommutators, deriving the exclusion principle rather than postulating it. Heisenberg and Pauli [Heisenberg:1929] [Heisenberg:1930] gave the general canonical formalism for an arbitrary system of interacting fields; in doing so they met the difficulty that the momentum conjugate to the time component of the electromagnetic potential vanishes identically, the first appearance of the constrained Hamiltonian structure taken up systematically in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism. Fock [Fock:1932] supplied the state space in which all of this acts, and Pauli and Weisskopf [Pauli:1934] closed the argument of Section 99.1.1 by quantizing the scalar field and showing that spin \(0\) has antiparticles without any sea.
| Year | Step | Motivation at the time |
|---|---|---|
| 1926 | Radiation field quantized mode by mode (Born, Heisenberg and Jordan) | Reproduce Einstein's black-body fluctuation formula |
| 1927 | Ladder operators, emission and absorption [Dirac:1927] | Derive the Einstein $A$ and $B$ coefficients |
| 1928 | Anticommuting field [Jordan:1928] | Derive the exclusion principle from an algebra |
| 1928 | First-order wave equation [Dirac:1928] | Obtain a positive probability density |
| 1929 | General canonical formalism [Heisenberg:1929] [Heisenberg:1930] | Quantize interacting fields covariantly |
| 1930 | Hole theory, with the hole taken to be the proton [Dirac:1930a] | Stabilize the negative-energy continuum |
| 1931 | The hole given the electron's own mass [Dirac:1931] | Remove the proton identification; predict a new particle |
| 1932 | Fock space [Fock:1932] | A state space for variable particle number |
| 1934 | Charged scalar field [Pauli:1934] | Show spin $0$ needs no sea |
| 1939 | Irreducible representations of the Poincaré group [Wigner:1939] | Define “particle” invariantly |
| 1948 | Boundary-dependent vacuum energy [Casimir:1948b] | Explain colloidal stability |
| 1950 | Wick's theorem [Wick:1950] | Organize the perturbation series |
Two later entries in Table 99.1 belong to the same story. Wigner's classification of the unitary irreducible representations of the Poincaré group [Wigner:1939] supplied the definition of “elementary particle” that this chapter uses — an irreducible representation labelled by the two invariants of Particles as Poincaré Representations — and Casimir's calculation of the force between conducting plates [Casimir:1948b], undertaken to explain the stability of colloidal suspensions, turned the zero-point energy from an embarrassment into a measurable quantity (Section 99.8).
The free real scalar field
Everything the rest of the chapter does is done first, and in full, for the simplest field there is: one real number at each point of spacetime, obeying Equation (99.3). It has no spin, no charge and no gauge freedom, so nothing distracts from the two operations being performed — imposing canonical commutators, and diagonalizing the Hamiltonian. The result is that the field is a continuous family of harmonic oscillators and that its excitations are counted by integers.
Canonical commutation relations
The classical system is Equation (99.18) with the conjugate momentum Equation (99.20) and the Poisson brackets Equation (99.22). Quantization is the same step taken in The Poisson Algebra and the Canonical Bridge to Quantum Mechanics for a system of finitely many degrees of freedom (Postulate 25.27): replace the classical variables by operators on a Hilbert space and the Poisson bracket by \((\ii\hbar)^{-1}\) times the commutator. The continuous label rides along unchanged.
The field \(\phi\) and its conjugate momentum density \(\pi\) become operator-valued distributions on a Hilbert space, satisfying at equal times
and the operators evolve in the Heisenberg picture, \(\ii\hbar\,\pp_{t}A=\comm{A}{H}\) with \(H\) of Equation (99.21).
Both sides of Equation (99.23) carry \(\mathrm{J}\,\mathrm{s}/\mathrm{m}^{3}\), by Equations (99.19) and (99.20) on the left and because \(\delta^{3}\) carries \(/\mathrm{m}^{3}\) on the right. The \(\hbar\) is not decoration and cannot be scaled away: it fixes the magnitude of the field fluctuations, and the classical field theory of Generalized Classical Field Theory is recovered as \(\hbar\to0\), where the right-hand side vanishes and the field components commute.
Equation (99.24) deserves a comment. It says that the field at two different points of space at the same time commutes. That is not an extra assumption but the continuum version of \(\comm{q_{n}}{q_{m}}=0\): the two are independent coordinates of one system, and specifying both simultaneously is exactly what a configuration is. The nontrivial statement — that \(\comm{\phi(x)}{\phi(y)}\) vanishes for all spacelike separations, not merely equal-time ones — is a theorem, proved in Section 99.7.1.
With \(H\) of Equation (99.21) and the brackets Equations (99.23) and (99.24),
and eliminating \(\pi\) gives \(\left(\Box+\kappa^{2}\right)\phi=0\). Rests on Equations (99.21), (99.23) and (99.24).
Derives Proposition 99.6. Only the term quadratic in \(\pi\) fails to commute with \(\phi\). Using \(\comm{A}{B^{2}}=B\comm{A}{B}+\comm{A}{B}B\) and Equation (99.23),
which is the first of Equation (99.25). For the second, the gradient and mass terms contribute. The mass term gives \(\tfrac12\kappa^{2}\int\dd^{3}y\,\comm{\pi(\vect{x})}{\phi(\vect{y})^{2}} =-\ii\hbar\kappa^{2}\phi(\vect{x})\), and the gradient term gives
the last step an integration by parts, the surface term vanishing for fields falling off at infinity. Adding and dividing by \(\ii\hbar\) gives the second of Equation (99.25); substituting \(\pi=c^{-2}\pp_{t}\phi\) gives \(c^{-2}\pp_{t}^{2}\phi-\nabla^{2}\phi+\kappa^{2}\phi=0\).
∎Proposition 99.6 is the consistency check that licenses everything else: the quantized field obeys, as an operator equation, exactly the classical equation of motion, so the classical solution theory — plane waves, superposition, the mass shell — can be carried over intact and used to diagonalize \(H\).
One technical point must be stated now rather than discovered later. Equation (99.23) contains a Dirac delta, and a delta is a distribution, not a function. Correspondingly \(\phi(x)\) is not an operator: the object \(\phi(x)\ket{0}\) has infinite norm, since Equation (99.91) below gives
which diverges quadratically at large \(k\). What is an operator is the smeared field
with \(f\) a smooth test function of compact support. Physically this is not a technicality but a statement about measurement: a field cannot be measured at a mathematical point, only averaged over a region, which is precisely the conclusion Bohr and Rosenfeld reached from the side of the apparatus (Section 99.7.3). It is because the smeared objects, and not the pointwise ones, are the well-defined quantities that the axiomatic reformulation of Axiomatic Quantum Field Theory takes fields to be operator-valued distributions from the start [Wightman:1956] [Streater:1964].
Mode expansion, creation and annihilation operators
The field equation is linear, so it is solved by superposing plane waves; the equal-time commutators then determine the algebra of the coefficients, and the coefficients turn out to be ladder operators.
Write the general real solution of \(\left(\Box+\kappa^{2}\right)\phi=0\) as a Fourier integral over the mass shell. With \(k\cdot x:=k_{\mu}x^{\mu}=\omega_{\vect{k}}t-\vect{k}\cdot\vect{x}\) and
each Fourier component \(\widetilde{\phi}(t,\vect{k})\) obeys \(\ddot{\widetilde{\phi}}=-\omega_{\vect{k}}^{2}\widetilde{\phi}\): one harmonic oscillator per wave vector, of frequency \(\omega_{\vect{k}}\). Reality of \(\phi\) relates the components at \(\vect{k}\) and \(-\vect{k}\), so the independent data are one complex amplitude per \(\vect{k}\), written \(a(\vect{k})\).
Equation (99.30) is not an independent assumption: it is \(c^{-2}\pp_{t}\) of Equation (99.29). The normalization factor is fixed by the requirement that Equation (99.23) hold, as Proposition 99.8 shows, and it is worth reading before the proof. The amplitude of each mode carries \(\sqrt{\hbar}\): as \(\hbar\to0\) the quantum field flattens to zero and only the classical solution built from a macroscopic occupation survives. That \(\sqrt{\hbar}\) is the entire quantum content of the free theory.
Dimensionally, \([\dd^{3}k]=/\mathrm{m}^{3}\) and \([\sqrt{\hbar c^{2}/2\omega}] =\mathrm{J}^{1/2}\,\mathrm{m}\), so \([a]=\mathrm{m}^{3/2}\) if \([\phi]\) is to be Equation (99.19) — which is exactly the dimension demanded by Equation (99.31) below, since \(\delta^{3}(\vect{k}-\vect{k}')\) carries \(\mathrm{m}^{3}\).
Equations (99.23) and (99.24) hold if and only if
Rests on Equation (99.23), Equation (99.24) and Definition 99.7.
Derives Proposition 99.8. Invert Equations (99.29) and (99.30) at a fixed time. Writing \(\alpha_{\vect{k}}=\sqrt{2\omega_{\vect{k}}/(\hbar c^{2})}\) and \(\beta_{\vect{k}}=\sqrt{2c^{2}/(\hbar\omega_{\vect{k}})}\), the combinations
follow from \(\int\dd^{3}x\,\ee^{-\ii\vect{k}\cdot\vect{x}} \ee^{\pm\ii\vect{k}'\cdot\vect{x}} =(2\pi)^{3}\delta^{3}(\vect{k}\mp\vect{k}')\). Note that \(\alpha_{\vect{k}}\) and \(\beta_{\vect{k}}\) are the reciprocals of the two normalization factors appearing in Equations (99.29) and (99.30), so that the spatial Fourier transforms of those two equations are
whose sum is \(2a(\vect{k})\ee^{-\ii\omega_{\vect{k}}t}\): the \(a^{\dagger}(-\vect{k})\) terms cancel, which is Equation (99.32). Then
the \(\comm{\phi}{\phi}\) and \(\comm{\pi}{\pi}\) terms having dropped by Equation (99.24). The remaining integral is \((2\pi)^{3}\delta^{3}(\vect{k}-\vect{k}')\), which sets \(\vect{k}'=\vect{k}\) in the bracket; there \(\alpha_{\vect{k}}\beta_{\vect{k}} =\sqrt{(2\omega/\hbar c^{2})(2c^{2}/\hbar\omega)}=2/\hbar\), so the bracket is \(4/\hbar\) and the prefactor \(\hbar/4\) cancels it exactly. For \(\comm{a(\vect{k})}{a(\vect{k}')}\) the same expansion gives \(\ii\alpha_{\vect{k}}\beta_{\vect{k}'}\comm{\phi(\vect{x})}{\pi(\vect{y})} +\ii\beta_{\vect{k}}\alpha_{\vect{k}'}\comm{\pi(\vect{x})}{\phi(\vect{y})}\), whose two terms are \(+\ii\hbar\) and \(-\ii\hbar\) times the same delta and therefore cancel; likewise for the daggered pair. The converse is the same algebra run backwards: substituting Equation (99.31) into Equations (99.29) and (99.30) reproduces Equation (99.23), as the computation in Proposition 99.44 does explicitly for a general time difference.
∎Equation (99.31) is the ladder algebra of the quantum harmonic oscillator, one copy for each \(\vect{k}\), with the Kronecker delta of the discrete oscillator replaced by a Dirac delta; it is the algebra Dirac used for the radiation field [Dirac:1927]. The modes it acts on are the continuum limit of the coupled oscillators of Oscillations and Mechanical Waves, which is a classical chapter and carries the normal modes but not this algebra. What the algebra buys is that the oscillator spectrum applies mode by mode, and the next theorem derives that here rather than importing it.
Substituting Equations (99.29) and (99.30) into Equation (99.21),
and the momentum operator \(\vect{P}=-\int\dd^{3}x\,\pi\nabla\phi\) is
Derives Theorem 99.9. Take the three terms of Equation (99.21) in turn. Each is a product of two mode integrals, and \(\int\dd^{3}x\) produces \((2\pi)^{3}\delta^{3}(\vect{k}\pm\vect{k}')\), so every term reduces to a single integral over \(\vect{k}\) in which the partner is either \(\vect{k}\) (for the \(aa^{\dagger}\) and \(a^{\dagger}a\) pieces) or \(-\vect{k}\) (for the \(aa\) and \(a^{\dagger}a^{\dagger}\) pieces).
For the \(aa\) and \(a^{\dagger}a^{\dagger}\) pieces, collect coefficients. From \(\tfrac12c^{2}\pi^{2}\) the coefficient is \(-\tfrac12c^{2}(\hbar\omega/2c^{2})=-\hbar\omega/4\), the minus sign coming from \((-\ii)^{2}\). From \(\tfrac12(\nabla\phi)^{2}\), the gradient acting on \(\ee^{\ii\vect{k}\cdot\vect{x}}\) gives \(\ii\vect{k}\) and on the partner \(\ee^{-\ii\vect{k}\cdot\vect{x}}\) gives \(-\ii\vect{k}\), so the coefficient is \(+\tfrac12(\hbar c^{2}/2\omega)\vect{k}^{2}\); the mass term adds \(+\tfrac12(\hbar c^{2}/2\omega)\kappa^{2}\). Their sum is \(\tfrac12(\hbar c^{2}/2\omega)(\vect{k}^{2}+\kappa^{2}) =\hbar\omega/4\) by Equation (99.28), cancelling the contribution from \(\pi^{2}\) exactly. The time-dependent factors \(\ee^{\mp2\ii\omega t}\) that multiply these pieces therefore never appear, which is the statement that \(H\) is conserved.
For the \(a^{\dagger}a\) and \(aa^{\dagger}\) pieces the relative sign is reversed: \(\pi^{2}\) contributes \(+\hbar\omega/4\) to each and the gradient and mass terms contribute \(+\hbar\omega/4\) to each, giving \(\hbar\omega/2\) for each of \(a^{\dagger}a\) and \(aa^{\dagger}\), which is Equation (99.33). Equation (99.34) is the same computation with one factor of \(\nabla\) in place of one factor of \(\pp_{t}/c^{2}\).
∎Two consequences follow at once from Equation (99.31) and Equation (99.33):
so \(a^{\dagger}(\vect{k})\) raises the energy by \(\hbar\omega_{\vect{k}}\) and \(a(\vect{k})\) lowers it by the same amount; and by Equation (99.34) the same operators raise and lower the momentum by \(\hbar\vect{k}\). The energy and momentum added by one application of \(a^{\dagger}(\vect{k})\) satisfy
the relativistic mass shell of a particle of mass \(m\). This is the central identification of the chapter, and it was not put in: it came out of diagonalizing a field Hamiltonian.
The combination \(\omega_{\vect{k}}\,\delta^{3}(\vect{k}-\vect{k}')\) is invariant under proper orthochronous Lorentz transformations, and
for any \(F\). Consequently the states
have a Lorentz-invariant normalization. Rests on Equations (99.28) and (99.31).
Derives Proposition 99.10. For a boost of rapidity \(\eta\) along the third axis, \(k'^{3}=\gamma(k^{3}-\beta\omega_{\vect{k}}/c)\) with \(\gamma=\cosh\eta\), \(\beta=\tanh\eta\), while \(k^{1},k^{2}\) are unchanged. On the mass shell \(\omega_{\vect{k}}=c\sqrt{\vect{k}^{2}+\kappa^{2}}\), so \(\pp\omega_{\vect{k}}/\pp k^{3}=c^{2}k^{3}/\omega_{\vect{k}}\) and
the last equality being the transformation of \(k^{0}=\omega/c\) itself. Hence \(\dd^{3}k'=(\omega_{\vect{k}'}/\omega_{\vect{k}})\dd^{3}k\), so \(\dd^{3}k/\omega_{\vect{k}}\) is invariant, and a delta transforms with the inverse Jacobian of its argument, so \(\omega_{\vect{k}}\delta^{3}(\vect{k}-\vect{k}')\) is invariant too. For Equation (99.37), write \(k^{2}-\kappa^{2}=(k^{0})^{2}-\vect{k}^{2}-\kappa^{2}\) and use \(\delta(g(k^{0}))=\sum_{i}\delta(k^{0}-k^{0}_{i})/\abs{g'(k^{0}_{i})}\) with the roots \(k^{0}=\pm\omega_{\vect{k}}/c\); the step function keeps the positive one, at which \(\abs{g'}=2\omega_{\vect{k}}/c\).
∎Equation (99.37) is the reason every momentum-space formula in this part carries \(\dd^{3}k/(2\pi)^{3}2\omega_{\vect{k}}\) rather than \(\dd^{3}k\): it is the unique (up to normalization) Lorentz-invariant measure on the mass shell, and using anything else makes covariance invisible. The convention Equation (99.38) is the one Quantum Electrodynamics and Renormalization uses, where the external-line spinors of Proposition 100.14 carry the matching normalization \(\bar{u}u=2m_{e}c\) and \(u^{\dagger}u=2E/c\); Scattering Theory is deliberately nonrelativistic and does not use it.
Fock space
The spectrum of Equation (99.33) is built exactly as for one oscillator: find the state annihilated by every lowering operator, then climb.
The vacuum \(\ket{0}\) is the normalized state with
The \(n\)-quantum states are \(a^{\dagger}(\vect{k}_{1})\cdots a^{\dagger}(\vect{k}_{n})\ket{0}\), their closed span is the \(n\)-particle sector \(\mathcal{F}_{n}\), and Fock space is the completed direct sum
with \(\mathcal{H}_{1}\) the one-quantum space. The number operator is
whose eigenvalue on \(\mathcal{F}_{n}\) is \(n\).
This is Fock's construction [Fock:1932]. Three properties make it the right state space, and each is a direct consequence of Equation (99.31).
The one-quantum states are particles. By Equation (99.35) and its momentum analogue, \(a^{\dagger}(\vect{k})\ket{0}\) is a simultaneous eigenstate of \(H\) and \(\vect{P}\) with eigenvalues \(\hbar\omega_{\vect{k}}\) and \(\hbar\vect{k}\) (measured from the vacuum values, which Section 99.2.4 disposes of), related by Equation (99.36). The space \(\mathcal{H}_{1}\) they span carries a unitary representation of the Poincaré group, and by Proposition 99.25 below that representation is irreducible with \(\hat{P}^{2}=m^{2}c^{2}\) and \(\hat{W}^{2}=0\) — mass \(m\), spin \(0\), in the classification Wigner obtained for the group itself [Wigner:1939] and which Particles as Poincaré Representations carries out in full. That is what entitles one to call the excitation a particle rather than a wave amplitude.
Multi-quantum states are automatically symmetric. Since \(\comm{a^{\dagger}(\vect{k})}{a^{\dagger}(\vect{k}')}=0\),
so the two-quantum state is unchanged under exchange of its labels. The symmetrization postulate of Identical Particles is therefore not a postulate here but a consequence of the choice of commutators, and with it the Bose–Einstein counting that Quantum Statistics builds the statistical mechanics of photons and phonons on. The point deserves emphasis: identical particles are identical in this formalism because they are not particles at all but quanta of one field, and there is nothing to label. What the label \(\vect{k}\) names is a mode, not an individual.
The construction is exhaustive and stable. Every state of the free theory is reached from \(\ket{0}\) by creation operators, because Equation (99.33) is a sum of commuting oscillator Hamiltonians each of which has the familiar nondegenerate ladder spectrum; and \(\comm{N}{H}=0\) by Equations (99.31) and (99.33), so the number of quanta of a free field is conserved. It is exactly the addition of an interaction term to \(H\) — something cubic or quartic in \(\phi\), hence not commuting with \(N\) — that lets the number change, which is what Section 99.1.1 demanded and what Quantum Electrodynamics and Renormalization exploits.
With \(\ket{\vect{k}_{1}\ldots\vect{k}_{n}} =\prod_{i}\sqrt{2\omega_{\vect{k}_{i}}}\, a^{\dagger}(\vect{k}_{i})\ket{0}\),
and the resolution of the identity on \(\mathcal{F}\) is
Derives Proposition 99.12. Move every annihilation operator of the bra to the right through the creation operators of the ket using Equation (99.31). Each annihilator must be contracted with exactly one creator, since any uncontracted annihilator meets \(\ket{0}\) and gives zero; the sum over which creator it meets is the sum over permutations, and each contraction supplies one factor \((2\pi)^{3}\delta^{3}\), while the normalization factors supply \(2\omega\). Equation (99.44) follows because the \(1/n!\) compensates the \(n!\) permutations in Equation (99.43) when the identity is applied to a symmetric state.
∎Definition 99.11 presupposes that a state annihilated by every \(a(\vect{k})\) exists and is unique. For the free field it does. For an interacting field it is exactly what Haag's theorem (Section 99.9.1) denies: the interacting theory's vacuum is not in the free theory's Fock space, and the two representations of Equation (99.23) are unitarily inequivalent. Nothing in this section is wrong, but its scope is the free field, and the reader should not carry the picture of “the” Fock space into the interacting theory without the warning of Section 99.9.1.
Normal ordering and the vacuum energy
Equation (99.33) was left in symmetric form on purpose. Using Equation (99.31) to move the annihilator to the right,
and \((2\pi)^{3}\delta^{3}(\vect{0})=\int\dd^{3}x=V\) is the volume of space. The vacuum energy is therefore \(E_{0}=V\varrho_{0}\) with energy density
which diverges as the fourth power of the upper limit: cutting the integral off at wave number \(\Lambda\gg\kappa\) gives
It is one half-quantum \(\tfrac12\hbar\omega_{\vect{k}}\) per mode — the zero-point energy of a single quantum oscillator, which Phenomenon 78.6 derives from the uncertainty relation and which is measured in the isotope shift of a band spectrum — summed over infinitely many modes.
The normal-ordered product \(:\!X\!:\) of a product of free-field operators is the same product with every creation operator moved to the left of every annihilation operator, the moves performed as if all the operators commuted — that is, discarding the commutators generated.
The prescription is Wick's [Wick:1950], and by construction \(\bra{0}:\!X\!:\ket{0}=0\). Replacing \(H\) by
removes Equation (99.46) and assigns the vacuum zero energy. It is essential to be honest about what this step is and is not.
What it is: a choice of the additive constant in the Hamiltonian. The equations of motion involve only commutators with \(H\), and a \(c\)-number added to \(H\) commutes with everything, so no Heisenberg equation, no scattering amplitude, no energy difference and no transition rate in flat space depends on the choice. Ordering ambiguities of exactly this kind arise whenever a classical product is promoted to an operator product, and the free field is the case in which the ambiguity is a pure constant.
What it is not: a proof that zero-point energy is unphysical. There are two situations in which the constant is not free, and both are experimental questions rather than formal ones.
Boundary-dependent differences. If boundaries restrict which modes exist, \(\varrho_{0}\) depends on the geometry, and the difference between two geometries is finite even though each is infinite. It is measured. That is the Casimir effect, computed and compared with experiment in Section 99.8.
Gravitational coupling. General relativity couples to the total energy–momentum tensor, not to energy differences, so the constant that Equation (99.48) discards is, in principle, gravitationally visible as a cosmological constant. Taking Equation (99.47) at face value with the cutoff at an energy \(E_{c}=\hbar c\Lambda\) gives
against an observed dark-energy density of order \(10^{-9}\,\mathrm{J}/\mathrm{m}^{3}\): a discrepancy of some \(56\) orders of magnitude if one stops at the highest energy at which the Standard Model has been tested, and of some \(120\) if one runs to the Planck energy. This is the cosmological constant problem [Weinberg:1989]. No symmetry of the Standard Model forbids the term, no cancellation among the known fields achieves it, and the problem is stated as open in Evidence-Based Cosmology and What We Observe but Do Not Understand; nothing in this chapter solves it, and the reader should treat any text that quietly normal-orders the problem away as having changed the subject.
Equation (99.49) is a cutoff estimate, and a sharp cutoff in \(\abs{\vect{k}}\) is not Lorentz invariant — it treats the modes as if there were a preferred frame, and a Lorentz-invariant regulator gives an energy density with the equation of state \(p=-\varrho\) appropriate to a cosmological constant rather than the \(p=\varrho/3\) that Equation (99.47) naively suggests. What survives the change of regulator is the magnitude, which is set by the fourth power of the largest energy at which the field description is assumed valid. It is the magnitude, not the tensor structure, that is \(10^{56}\) times too large.
The complex scalar field and antiparticles
A real field describes one species of quantum, which is its own antiquantum and carries no charge. Allowing the field to be complex doubles the classical degrees of freedom, and the doubling shows up in the quantum theory as two species of quantum of equal mass and opposite charge. Antimatter is a prediction of field theory, obtained here without any reference to a filled sea.
Global phase symmetry and conserved charge
Let \(\phi\) be a complex field with Lagrangian density
with \([\phi]\) again given by Equation (99.19); the factor \(\tfrac12\) of Equation (99.18) is absent because \(\phi\) and \(\phi^{\dagger}\) are independent. Both \(\phi\) and \(\phi^{\dagger}\) obey \(\left(\Box+\kappa^{2}\right)\phi=0\). The conjugate momenta are
and the canonical brackets are
all others vanishing.
Equation (99.50) is invariant under the global phase rotation
a representation of the group \(\U(1)\) of unit-modulus complex numbers. That group is isomorphic to the rotation group \(\SO(2,\R)\) of Section 14.2.1 by \(\ee^{\ii\theta}\mapsto R(\theta)\) — the section is written for \(\SO(2,\R)\) and does not mention \(\U(1)\), but its representation theory transfers along the isomorphism unchanged. By Noether's theorem [Noether:1918], in the field form derived in Section 99.4.1 below, the symmetry carries a conserved current.
The current
satisfies \(\pp_{\mu}j^{\mu}=0\) on solutions, and the dimensionless quantity
is conserved. The electric charge of a field of charge quantum \(q\) is \(qQ\). Rests on Equations (99.50), (99.51) and (99.53).
Derives Proposition 99.16. Conservation is direct: \(\pp_{\mu}j^{\mu} =\ii(\phi^{\dagger}\Box\phi-\phi\Box\phi^{\dagger}) =\ii(-\kappa^{2}\phi^{\dagger}\phi+\kappa^{2}\phi\phi^{\dagger})=0\), the cross terms \(\pp_{\mu}\phi^{\dagger}\pp^{\mu}\phi\) cancelling identically and the field equation supplying the rest. That Equation (99.54) is the Noether current of Equation (99.53) is Theorem 99.22 applied with \(\delta\phi=-\ii\phi\), \(\delta\phi^{\dagger}=+\ii\phi^{\dagger}\):
For Equation (99.55), note \(j^{0}=\ii(\phi^{\dagger} c^{-1}\pp_{t}\phi-(c^{-1}\pp_{t}\phi^{\dagger})\phi) =\ii c(\phi^{\dagger}\pi^{\dagger}-\pi\phi)\) by Equation (99.51) — the second factor written to the right of its derivative, as the Noether construction delivers it, since \(\phi\) and \(\pi\) do not commute and the two orderings differ by the \(c\)-number \(\comm{\pi}{\phi}\) that Equation (99.60) normal-orders away. Dividing by \(\hbar c\) makes the integrand an inverse volume: \([j^{0}]=\mathrm{J}/\mathrm{m}^{2}\) and \([\hbar c]=\mathrm{J}\,\mathrm{m}\). Conservation of \(Q\) follows from \(\pp_{\mu}j^{\mu}=0\) by integrating over a spatial slice and discarding the surface term at infinity.
∎This settles the interpretive question left open in Section 99.1.1. The density Equation (99.4) that could not be a probability density is Equation (99.54) up to normalization, and the reason it may take either sign is that it is a charge density. Charge densities of either sign are unremarkable; probability densities of either sign are nonsense. Nothing had to be repaired — the quantity had been misidentified. The same reading is what makes minimal coupling \(\pp_{\mu}\to\pp_{\mu}+\ii qA_{\mu}/\hbar\) legitimate, since \(-j^{\mu}A_{\mu}\) is then the interaction energy of a charge distribution with the potential, exactly as in Theorem 100.5; a probability current has no business coupling to the electromagnetic field.
Notice also what has not been explained. Nothing above forces the eigenvalues of \(Q\) to be integers times one universal quantum, and nothing forces the electron and proton charges to be equal and opposite to the precision at which that is measured. The \(\U(1)\) of Equation (99.53) acts on one field, and nothing fixes the phase a second field of a different charge would pick up: the symmetry actually forced is a group of real phases, whose representations are labelled by a real number. That the observed charges are integer multiples of one quantum is an experimental fact — its current status, including the fractional quark charges never seen free, is Section 63.11.3 — which the free field theory does not derive and which nothing in this chapter explains.
Two species of quantum: particle and antiparticle
The field is complex, so its Fourier coefficients at \(\vect{k}\) and the conjugates at \(-\vect{k}\) are no longer related. Two independent families of ladder operators appear.
with
and every other commutator among \(a,a^{\dagger},b,b^{\dagger}\) vanishing.
That Equation (99.58) is equivalent to Equation (99.52) is the computation of Proposition 99.8 repeated: the \(\comm{\phi}{\pi}\) bracket receives the \(a\) contribution with one sign and the \(b\) contribution with the other, and the two add rather than cancel because \(\comm{b^{\dagger}}{b}=-\comm{b}{b^{\dagger}}\); the normalization \(\sqrt{\hbar c^{2}/2\omega}\) is unchanged.
With normal ordering,
The two species have identical dispersion Equation (99.28), hence identical mass, and charges \(+1\) and \(-1\). Rests on Definition 99.17, Equation (99.55) and Theorem 99.9.
Derives Theorem 99.18. The Hamiltonian density is \(\Ham=c^{2}\pi^{\dagger}\pi+\nabla\phi^{\dagger}\cdot\nabla\phi +\kappa^{2}\phi^{\dagger}\phi\); substituting Equations (99.56) and (99.57) and integrating over \(\vect{x}\) proceeds exactly as in Theorem 99.9, the \(aa\) and \(a^{\dagger}a^{\dagger}\) pieces cancelling between the \(\pi\) term and the gradient-plus-mass terms, and yields \(\int_{\vect{k}}\hbar\omega_{\vect{k}} (a^{\dagger}a+b\,b^{\dagger})\), which normal-ordering turns into Equation (99.59). For the charge, insert Equation (99.51) into Equation (99.55):
and substitute the expansions. The spatial integral pairs each mode either with itself or with its reflection, and the terms pairing two annihilation or two creation operators — \(b\,a\) carrying \(\ee^{-2\ii\omega t}\) and \(a^{\dagger}b^{\dagger}\) carrying \(\ee^{2\ii\omega t}\) — occur in \(\phi^{\dagger}\dot{\phi}\) and in \(\dot{\phi}^{\dagger}\phi\) with the same coefficient and cancel in the difference. The surviving terms occur with opposite coefficients in the two products and therefore add: \(a^{\dagger}a\) with \(-2\ii\omega_{\vect{k}}f_{\vect{k}}^{2}\) and \(b\,b^{\dagger}\) with \(+2\ii\omega_{\vect{k}}f_{\vect{k}}^{2}\), where \(f_{\vect{k}}^{2}=\hbar c^{2}/2\omega_{\vect{k}}\). Multiplying by the prefactor \(\ii/\hbar c^{2}\) turns these into \(+1\) and \(-1\), giving \(Q=\int_{\vect{k}}(a^{\dagger}a-b\,b^{\dagger})\) and Equation (99.60) after normal ordering.
∎Both terms of Equation (99.59) are non-negative: the \(b\) quanta cost positive energy, not negative. The negative-energy solutions of Section 99.1.1 have not been suppressed, filled or forbidden — they have been reinterpreted. The coefficient of the “negative-frequency” exponential \(\ee^{+\ii k\cdot x}\) in Equation (99.56) is not the amplitude to find a particle of negative energy but the operator that creates an antiparticle of positive energy. This is the Stueckelberg–Feynman reading [Stueckelberg:1941] [Feynman:1949a], and Section 99.6.1 shows that the same exponential is what makes a single propagator serve for both species.
Compare with Section 99.1.1. Hole theory required an infinitely occupied vacuum, which is unavailable for bosons, and it predicted antiparticles only for fermions. Here the antiparticle appears for a spin-\(0\) field, with a vacuum containing no quanta at all, and it appears for exactly the reason the hole picture could not accommodate: the field, not the particle, is fundamental, and a complex field has two kinds of excitation because it has two independent real components. That is Pauli and Weisskopf's argument [Pauli:1934], and it is the reason the hole theory [Dirac:1930a] is now of historical interest only, correct predictions notwithstanding.
Two experimental anchors are worth naming precisely, since the claim being tested is not the existence of antimatter but the exact degeneracy that Theorem 99.18 asserts. The first is the positron itself: predicted, then found in cloud-chamber tracks whose curvature in a magnetic field gave a positive particle of about the electron mass [Anderson:1933], the subject of Experiment: The Positron. The second is the modern programme of antiproton and antihydrogen measurements pursued in Experiment: Precision Spectroscopy and Atomic Clocks, which tests the equality of particle and antiparticle masses and charges directly rather than inferring it. Both belong to the wider \(CPT\) programme of Discrete Symmetries and CPT: the degeneracy asserted here follows from \(CPT\), so a measured violation of it would be a violation of one of the assumptions — Lorentz invariance, locality, or the positivity of the energy — on which this whole chapter rests.
Charge conjugation
The doubling of species comes with a discrete symmetry that exchanges them.
\(C\) is the unitary operator with
with \(\abs{\eta_{C}}=1\) a convention-dependent phase, conventionally \(\eta_{C}=1\).
With \(\eta_{C}=1\),
and \(\comm{C}{H}=0\) for the free field. Rests on Equation (99.61), Definition 99.17 and Equation (99.60).
Derives Proposition 99.21. Apply Equation (99.61) term by term to Equation (99.56): the coefficient of \(\ee^{-\ii k\cdot x}\) becomes \(b(\vect{k})\) and that of \(\ee^{\ii k\cdot x}\) becomes \(a^{\dagger}(\vect{k})\), which is Equation (99.57). Then \(j^{\mu}=\ii(\phi^{\dagger}\pp^{\mu}\phi-\phi\pp^{\mu}\phi^{\dagger})\) maps to \(\ii(\phi\pp^{\mu}\phi^{\dagger} -\phi^{\dagger}\pp^{\mu}\phi)=-j^{\mu}\), and integrating gives the statement for \(Q\); equivalently, Equation (99.60) is manifestly odd under \(a\leftrightarrow b\). That Equation (99.59) is even under the same exchange is \(\comm{C}{H}=0\).
∎For a real scalar field \(\phi^{\dagger}=\phi\), so Equation (99.62) forces \(C\phi C^{-1}=\pm\phi\): a real field is its own antifield, has \(Q\equiv0\), and carries an intrinsic \(C\) parity. The electromagnetic field has \(C=-1\), since \(C\) reverses all charges and hence all currents and the potentials they source; a state of \(n\) photons therefore has \(C=(-1)^{n}\), which is why the \(C\)-even neutral pion decays to two photons and not to three.
Beyond this the account must be deferred. \(C\) alone is one of three discrete operations, and the free complex scalar is too simple a system to show what is interesting about them: the weak interaction violates \(C\) and \(P\) separately and maximally (Weak Interactions), the combination \(CP\) is violated too but only slightly, and the combination \(CPT\) is a theorem of any Lorentz-invariant local field theory with a bounded-below Hamiltonian [Luders:1957] [Jost:1957] — the same three hypotheses that appear in the spin–statistics theorem of Section 99.7.2. That is not an accident, and it is the subject of Discrete Symmetries and CPT.
Noether charges and the Poincaré algebra
Two results used repeatedly above were stated as they were needed: that a continuous symmetry of the Lagrangian density gives a conserved current, and that the one-quantum states of a free field carry an irreducible representation of the Poincaré group labelled by a mass and a spin. This section proves both. They are what make the phrase “a particle of mass \(m\) and spin \(s\)” a definition rather than a description.
Noether's theorem in field form
Let \(\Lag(\phi_{a},\pp_{\mu}\phi_{a})\) be a local Lagrangian density for a collection of fields \(\phi_{a}\), and let the infinitesimal transformation \(\phi_{a}\mapsto\phi_{a}+\varepsilon F_{a}[\phi]\) change \(\Lag\) by at most a total divergence, \(\delta\Lag=\varepsilon\,\pp_{\mu}K^{\mu}\). Then
satisfies \(\pp_{\mu}j^{\mu}=0\) on solutions of the field equations, and \(Q=c^{-1}\int\dd^{3}x\,j^{0}\) is constant in time whenever the fields fall off fast enough at spatial infinity. Rests on Definition 99.4 and Equation (99.16).
Derives Theorem 99.22. Compute \(\delta\Lag\) two ways. Directly from the chain rule, with summation over \(a\),
On a solution the first bracket vanishes by Equation (99.16). By hypothesis the same \(\delta\Lag\) equals \(\varepsilon\pp_{\mu}K^{\mu}\), so \(\pp_{\mu}\bigl(\pp\Lag/\pp(\pp_{\mu}\phi_{a})\,F_{a} -K^{\mu}\bigr)=0\). Integrating \(\pp_{\mu}j^{\mu}=0\) over a spatial slice, \(\pp_{t}\int\dd^{3}x\,j^{0}/c=-\int\dd^{3}x\,\nabla\cdot\vect{j}\), which is a surface term at infinity.
∎In the quantum theory the conserved charge does more than sit there: it generates the symmetry it came from.
For the \(\U(1)\) charge of Equation (99.55),
which is Equation (99.53). Rests on Equations (99.53), (99.58) and (99.60).
Derives Proposition 99.23. From Equation (99.60) and Equation (99.58), \(\comm{:\!Q\!:}{a(\vect{k})}=-a(\vect{k})\) and \(\comm{:\!Q\!:}{b^{\dagger}(\vect{k})}=-b^{\dagger}(\vect{k})\); both coefficients in Equation (99.56) are therefore lowered by one unit, giving the first equation. The second follows from the identity \(\ee^{X}Y\ee^{-X}=Y+\comm{X}{Y}+\tfrac{1}{2!}\comm{X}{\comm{X}{Y}} +\cdots\) with \(X=\ii\theta Q\), which here terminates into an exponential series because \(\comm{Q}{\cdot}\) acts as multiplication by \(-1\) on \(\phi\).
∎Applying Theorem 99.22 to spacetime translations \(\phi_{a}(x)\mapsto\phi_{a}(x+\varepsilon e)\), for which \(\delta\Lag=\varepsilon\,e^{\nu}\pp_{\nu}\Lag\) is a total divergence with \(K^{\mu}=e^{\mu}\Lag\), gives four conserved currents assembled into the canonical energy–momentum tensor
of which \(\hat{P}^{0}=H/c\) and \(\hat{P}^{i}=\vect{P}\) recover Equations (99.33) and (99.34): for the real scalar, \(T^{00}=\Ham\) of Equation (99.21) and \(T^{0i}=c^{-1}\pp_{t}\phi\,\pp^{i}\phi =-c\pi\pp_{i}\phi\). Lorentz transformations \(\delta\phi=\tfrac12\omega_{\rho\sigma} (x^{\rho}\pp^{\sigma}-x^{\sigma}\pp^{\rho})\phi\) give six more currents,
which for fields carrying spin acquire an extra term from the rotation of the field components themselves — the term whose separation from the orbital piece is the Belinfante–Rosenfeld problem [Belinfante:1940] treated in Generalized Classical Field Theory.
The Poincaré algebra and the labels of a particle
The ten charges \(\hat{P}^{\mu},\hat{J}^{\rho\sigma}\) of Equations (99.65) and (99.66) are Hermitian operators on Fock space, and they close on the Poincaré algebra Equations (94.5), (94.6) and (94.7) of Proposition 94.8, with an \(\ii\hbar\) in every bracket because the generators are the physical momentum and angular momentum and not the dimensionless differential operators of Section 14.3.2. That the same algebra appears here as an algebra of field integrals, and in Particles as Poincaré Representations as an algebra of abstract generators, is the whole content of the phrase “the field theory is Poincaré invariant”.
\(\comm{\hat{P}^{\mu}}{\phi(x)}=-\ii\hbar\,\pp^{\mu}\phi(x)\), so that \(\ee^{\ii a\cdot\hat{P}/\hbar}\phi(x)\ee^{-\ii a\cdot\hat{P}/\hbar} =\phi(x+a)\), and \(\comm{\hat{P}^{\mu}}{\hat{P}^{\nu}}=0\). Rests on Proposition 99.6, Equation (99.34) and Equation (99.31).
Derives Proposition 99.24. For \(\mu=0\), where \(\hat{P}^{0}=H/c\), this is the Heisenberg equation \(\ii\hbar\,\pp_{t}\phi=\comm{\phi}{H}\) of Proposition 99.6. For the spatial components use Equation (99.34) and Equation (99.31): \(\comm{\vect{P}}{a(\vect{k})}=-\hbar\vect{k}\,a(\vect{k})\) and \(\comm{\vect{P}}{a^{\dagger}(\vect{k})} =+\hbar\vect{k}\,a^{\dagger}(\vect{k})\), so in Equation (99.29) each term acquires the factor that \(-\ii\hbar\nabla\) would produce on its exponential. The two components commute because both are diagonal in \(\vect{k}\).
∎On \(\mathcal{H}_{1}=\operatorname{Sym}^{1}(\mathcal{H}_{1})\) of the free real scalar field,
with \(\hat{W}_{\mu}=\tfrac12\epsilon_{\mu\nu\rho\sigma} \hat{J}^{\nu\rho}\hat{P}^{\sigma}\) the Pauli–Lubanski vector of Definition 94.15, and the representation is irreducible: the free real scalar field describes one species of particle of mass \(m\) and spin \(0\). Rests on Proposition 99.24, Theorem 94.18 and Definition 94.15.
Derives Proposition 99.25. By Proposition 99.24 the state \(a^{\dagger}(\vect{k})\ket{0}\) is a simultaneous eigenstate of \(\hat{P}^{\mu}\) with eigenvalue \(\hbar k^{\mu}=(\hbar\omega_{\vect{k}}/c,\hbar\vect{k})\), so \(\hat{P}^{2}\) has eigenvalue \(\hbar^{2}k^{2}=\hbar^{2}\kappa^{2}=m^{2}c^{2}\) by Equations (99.2) and (99.28) — which is Equation (94.14), the definition of mass. For \(\hat{W}^{2}\), evaluate in the rest frame, legitimate because \(\hat{W}^{2}\) is a Casimir (Theorem 94.18) and therefore frame independent: there \(\hat{P}^{\mu}=(mc,\vect{0})\) and \(\hat{W}_{i}=-mc\,\hat{J}_{i}\), \(\hat{W}_{0}=0\), so \(\hat{W}^{2}=-m^{2}c^{2}\hat{\vect{J}}^{2}\). For a scalar field \(\hat{J}^{\rho\sigma}\) is purely orbital by Equation (99.66), so \(\hat{\vect{J}}\) annihilates a zero-momentum one-quantum state and \(\hat{W}^{2}=0\): comparing with \(\hat{W}^{2}=-m^{2}c^{2}\hbar^{2}s(s+1)\) of Equation (94.25) gives \(s=0\). Irreducibility: the states \(a^{\dagger}(\vect{k})\ket{0}\) for \(\vect{k}\) on the mass shell form a single orbit of the proper orthochronous Lorentz group (Proposition 94.21), and the little group of a massive momentum is \(\SU(2)\) (Proposition 94.24), acting here on a one-dimensional space, so there is no invariant subspace.
∎Wigner's theorem [Wigner:1939] makes “particle” a group-theoretic notion: a species is a unitary irreducible representation of the Poincaré group, labelled by the two invariants of Theorem 94.18, and the free fields of this chapter are devices for realizing those representations on a Fock space where quanta can be created and destroyed. The classification is complete for free particles and is silent about interacting ones: an interacting theory's asymptotic states are again irreducible representations (that is what makes scattering theory possible, Scattering Theory), but a confined quark, which never appears asymptotically, is not classified by it at all (Quantum Chromodynamics). The massless case is different in kind and not merely in the value of \(m\): there \(\hat{P}^{2}=\hat{W}^{2}=0\) forces \(\hat{W}^{\mu}\propto\hat{P}^{\mu}\), the proportionality constant is the helicity, and a massless particle of nonzero spin has two states rather than \(2s+1\) (Proposition 94.32) — which is why the photon of Section 99.5.2 has two polarizations and the massive vector of Section 99.5.3 has three.
Fermions and gauge fields
The scalar field is the whole method in miniature, but it is not any field that occurs in Nature. The matter fields of the Standard Model carry spin \(1/2\), and the force fields carry spin \(1\) and a gauge redundancy. Each brings a genuinely new feature, and in both cases the new feature is forced rather than chosen.
The Dirac field and anticommutation
The classical Dirac field is a four-component object \(\psi_{a}(x)\) with Lagrangian density
with \(\acomm{\gamma^{\mu}}{\gamma^{\nu}}=2\eta^{\mu\nu}\identity\), \((\gamma^{0})^{\dagger}=\gamma^{0}\) and \((\gamma^{i})^{\dagger}=-\gamma^{i}\), exactly as constructed in The Dirac Equation and used in Equation (100.5). The field carries \([\psi]=\mathrm{m}^{-3/2}\) (Definition 100.2), so that \(\psi^{\dagger}\psi\) is a number density and both terms of Equation (99.68) are energy densities. The Euler–Lagrange equation is \((\ii\hbar\gamma^{\mu}\pp_{\mu}-mc)\psi=0\), and every component separately satisfies \((\Box+\kappa^{2})\psi_{a}=0\), obtained by applying \((\ii\hbar\gamma^{\nu}\pp_{\nu}+mc)\) and using the Clifford algebra.
The conjugate momentum is peculiar. Since \(\Lag\supset\ii\hbar c\,\bar{\psi}\gamma^{0}c^{-1}\pp_{t}\psi =\ii\hbar\,\psi^{\dagger}\pp_{t}\psi\),
that is, the momentum is the field's own conjugate rather than something new: Equation (99.68) is first order in time, so \(\psi\) and \(\psi^{\dagger}\) are not independent initial data but together constitute the phase space. This is a constrained system in the sense of Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism, and the canonical bracket that survives the constraint analysis is the one between \(\psi\) and \(\psi^{\dagger}\) alone.
The question this subsection answers is whether that bracket should be a commutator or an anticommutator. The classical theory says nothing: the canonical procedure of Hamiltonian Mechanics produces a Poisson bracket, and both promotions reduce to it as \(\hbar\to0\). Nature decides, through the requirement that the energy be bounded below.
Write \(p^{\mu}=\hbar k^{\mu}\) and let \(u_{s}(p),v_{s}(p)\), \(s=1,2\), be the positive- and negative-frequency spinor solutions of The Dirac Equation, normalized as in Proposition 100.14 by
with \(E_{p}=\hbar\omega_{\vect{k}}\). The negative-frequency spinors are not normalized by copying these: the Lorentz scalar changes sign while the norm, being a density, does not, so
The two are consistent: tracing the completeness relations gives \(\sum_{s}\bar{u}_{s}u_{s}=+4mc\) and \(\sum_{s}\bar{v}_{s}v_{s}=-4mc\). Then
The normalization is fixed by dimensions: \([\sqrt{c/2\hbar\omega}] =\mathrm{kg}^{-1/2}\,\mathrm{m}^{-1/2}\,\mathrm{s}^{1/2}\) and the spinors carry the inverse of that by Equation (99.70), leaving \([\psi]=\mathrm{m}^{-3/2}\) once \([\dd^{3}k]\) and \([b]=\mathrm{m}^{3/2}\) are counted.
Let the mode operators of Equation (99.72) obey either commutation or anticommutation relations of the canonical form. Then the Hamiltonian is
which is unbounded below if \(\comm{d_{s}(\vect{k})}{d^{\dagger}_{s'}(\vect{k}')} =\delta_{ss'}(2\pi)^{3}\delta^{3}(\vect{k}-\vect{k}')\), and equal to a non-negative operator plus a constant if
all other brackets vanishing. The corresponding equal-time relation for the field is
Derives Theorem 99.28. From Equations (99.68) and (99.69) the Hamiltonian density is \(\Ham=\pi\dot{\psi}-\Lag =\psi^{\dagger}\left(-\ii\hbar c\,\gamma^{0}\gamma^{i}\pp_{i} +mc^{2}\gamma^{0}\right)\psi\), and the operator in brackets acting on a solution \(\ee^{\mp\ii k\cdot x}\) returns \(\pm\hbar\omega_{\vect{k}}\) — this is just the Dirac equation solved for \(\ii\hbar\pp_{t}\psi\). Substituting Equation (99.72) and integrating over \(\vect{x}\), the \(\ee^{\pm2\ii\omega t}\) cross terms vanish by \(u^{\dagger}_{s}(p)v_{s'}(\tilde{p})=0\) with \(\tilde{p}=(p^{0},-\vect{p})\), and the diagonal terms are weighted by \(u^{\dagger}u=v^{\dagger}v=2E_{p}/c\), which cancels the \(c/2\hbar\omega\) of the normalization and leaves Equation (99.73), the minus sign on the \(d\) term coming from the sign of the eigenvalue on the negative-frequency solutions.
Now compare the two options. With commutators, \(-d\,d^{\dagger}=-d^{\dagger}d-(2\pi)^{3}\delta^{3}(\vect{0})\), so
and \(d^{\dagger}\) applied repeatedly lowers the energy without limit: the spectrum is unbounded below, the vacuum is not the ground state, and no state is stable. With anticommutators, \(-d\,d^{\dagger}=+d^{\dagger}d-(2\pi)^{3}\delta^{3}(\vect{0})\) by Equation (99.74), so
a sum of non-negative operators, since \(\acomm{d}{d^{\dagger}}\ge0\) makes \(d^{\dagger}d\) positive semi-definite. Equation (99.75) follows by substituting Equation (99.72) and using the completeness relation
which cancels the normalization factor and leaves \(\int\dd^{3}k\,(2\pi)^{-3}\ee^{\ii\vect{k}\cdot(\vect{x}-\vect{y})} \delta_{ab}\). That Equation (99.75) is the correct canonical relation is confirmed by Equation (99.69): the naive prescription \(\acomm{\psi}{\pi}=\ii\hbar\delta^{3}\) with \(\pi=\ii\hbar\psi^{\dagger}\) gives exactly \(\acomm{\psi}{\psi^{\dagger}}=\delta^{3}\), without any \(\hbar\) surviving — the field is dimensionless in the appropriate sense, and the \(\hbar\) has been absorbed into the constraint.
∎The choice is therefore not a convention. Given Equation (99.68), commutators produce a theory with no ground state, which is not a theory of anything; anticommutators produce a positive Hamiltonian. This is one half of the spin–statistics connection; the other half — that a scalar quantized with anticommutators is equally impossible, for a different reason — is Section 99.7.2.
No two quanta of a Dirac field occupy the same mode: a state containing two electrons of identical momentum and spin does not exist. The consequences are the periodic table (Atoms and Molecules) and the degeneracy pressure that supports white dwarfs and neutron stars (Compact Stars and Relativistic Astrophysics). Rests on Theorem 99.28 and Equation (99.74).
Derivation. Derives Phenomenon 99.29. Setting \(s=s'\) and \(\vect{k}'=\vect{k}\) in the second family of Equation (99.74) gives \(\acomm{b^{\dagger}_{s}(\vect{k})}{b^{\dagger}_{s}(\vect{k})} =2\left(b^{\dagger}_{s}(\vect{k})\right)^{2}=0\), so
identically as an operator. Any state containing two quanta in the same mode would be \(b^{\dagger}_{s}(\vect{k})b^{\dagger}_{s}(\vect{k}) \ket{\cdots}\), which is the zero vector, not a state. Equivalently, the number operator for one mode, \(n=b^{\dagger}b\), satisfies \(n^{2}=b^{\dagger}bb^{\dagger}b =b^{\dagger}(1-b^{\dagger}b)b=n\) by Equations (99.74) and (99.77), so its eigenvalues satisfy \(n^{2}=n\) and are \(0\) or \(1\). The antisymmetry of the multi-quantum wave function follows from the same relation with distinct labels: \(b^{\dagger}_{s}(\vect{k})b^{\dagger}_{s'}(\vect{k}')\ket{0} =-b^{\dagger}_{s'}(\vect{k}')b^{\dagger}_{s}(\vect{k})\ket{0}\).
∎Phenomenon 99.29 is worth pausing on. In Identical Particles the antisymmetry of the fermionic wave function is a postulate, added to quantum mechanics because it is observed. Here it is a theorem, following from Equation (99.74), which in turn follows from requiring the energy to be bounded below. Nothing was assumed about identical particles at all — there are no particles in Equation (99.68), only a field, and the quanta are identical because they are excitations of one field.
Equation (99.76) contains two non-negative terms and a vacuum defined by \(b_{s}(\vect{k})\ket{0}=d_{s}(\vect{k})\ket{0}=0\): no level is occupied, and the positron is a \(d\) quantum, not a hole. The hole picture of Section 99.1.1 survives only as a mnemonic. It is not merely superfluous: the constant discarded in passing from Equation (99.73) to Equation (99.76) is \(-\sum_{s}\int_{\vect{k}} \hbar\omega_{\vect{k}}(2\pi)^{3}\delta^{3}(\vect{0})\), negative where the bosonic zero-point energy Equation (99.46) was positive, so a fermion field's vacuum energy has the opposite sign to a boson's. That sign is a genuine structural fact — it is what makes the cancellation of Equation (99.49) conceivable in principle — and it is invisible in the sea picture, where the negative energy is attributed to occupied states.
The electromagnetic field and gauge redundancy
The electromagnetic field is the first case in which the canonical procedure does not simply run. The Lagrangian density is
with \([A^{\mu}]=\mathrm{V}\,\mathrm{s}/\mathrm{m}\) and \([F_{\mu\nu}]=\mathrm{T}\) as in Definition 100.2, and \(\vect{E}=-\nabla A^{0}c-\pp_{t}\vect{A}\), \(\vect{B}=\nabla\times\vect{A}\) in terms of \(A^{\mu}=(A^{0},\vect{A})\). The conjugate momenta are
The sign is the index placement and nothing else: differentiating Equation (99.78) with respect to \(\dot{A}^{i}\), the contravariant component appearing in \(\vect{E}\), gives \(-\varepsilon_{0}E^{i}\), and \(A_{i}=-A^{i}\) then flips it. Both indices are read the same way in every commutator below, so that \(\vect{E}\) appears there through its covariant components \(E_{i}=-E^{i}\).
The vanishing of \(\pi^{0}\) is not an accident of a bad variable choice and cannot be repaired by one. \(F_{\mu\nu}\) is antisymmetric, so \(\dot{A}_{0}\) appears nowhere in Equation (99.78), and \(\pi^{0}=0\) holds identically rather than dynamically: it is a primary constraint, the prototype of the structures analysed by Dirac's algorithm in Constrained Hamiltonian Systems: the Dirac–Bergmann Formalism and the difficulty Heisenberg and Pauli met first [Heisenberg:1929] [Heisenberg:1930]. Consistency, \(\pp_{t}\pi^{0}=0\), then generates the secondary constraint \(\nabla\cdot\vect{\pi}=0\), that is Gauss's law \(\nabla\cdot\vect{E}=0\) in vacuum. The canonical commutator \(\comm{A_{0}}{\pi^{0}}=\ii\hbar\delta^{3}\) is inconsistent with \(\pi^{0}=0\), so Postulate 99.5 cannot be imposed on all four components. The underlying reason is gauge invariance: \(A_{\mu}\) and \(A_{\mu}+\pp_{\mu}\Lambda\) describe the same \(F_{\mu\nu}\) and the same physics, so \(A_{\mu}\) carries more components than there are degrees of freedom, and the extra ones have no dynamics to quantize.
Two treatments are standard, and they differ in which of covariance and positivity is given up first.
Coulomb gauge. Impose \(\nabla\cdot\vect{A}=0\). Gauss's law then makes \(A^{0}\) obey \(\nabla^{2}A^{0}=0\) in vacuum, so \(A^{0}=0\) with sensible boundary conditions: no dynamics, and two independent components of \(\vect{A}\) remain. This is what Born, Heisenberg and Jordan did in 1926 and what Dirac [Dirac:1927] systematized. The expansion is
with \(\vect{\varepsilon}_{\lambda}\cdot\vect{k}=0\) enforcing the gauge condition, \(\comm{a_{\lambda}(\vect{k})} {a^{\dagger}_{\lambda'}(\vect{k}')} =\delta_{\lambda\lambda'}(2\pi)^{3}\delta^{3}(\vect{k}-\vect{k}')\), and the factor \(\mu_{0}\) making the dimensions come out: \([\sqrt{\hbar\mu_{0}c^{2}/2\omega}] =\mathrm{V}\,\mathrm{s}\,\mathrm{m}^{1/2}\), which with \([\dd^{3}k]\) and \([a]\) gives \(\mathrm{V}\,\mathrm{s}/\mathrm{m}\). The price is that the gauge condition is not Lorentz invariant, and that the equal-time commutator is modified: since \(\nabla\cdot\vect{A}=0\) holds as an operator equation, \(\comm{A_{i}}{\pi_{j}}\) cannot be proportional to \(\delta_{ij}\delta^{3}\), whose divergence does not vanish. The correct relation is
the transverse delta, which is the projector onto divergence-free vector fields and satisfies \(\pp_{i}\delta^{\mathrm{T}}_{ij}=0\). Its long-range second term is the Coulomb interaction in disguise: in this gauge the instantaneous Coulomb field is not a dynamical variable but part of the constraint, which is why the gauge is convenient for bound states and awkward for scattering.
Derivation. Derives Equation (99.81). Only the transverse part of \(\vect{A}\) is a dynamical variable, so the canonical bracket can only be the transverse part of \(\ii\hbar\,\delta_{ij}\delta^{3}\); the whole content of Equation (99.81) is which distribution that is. Work in momentum space, writing \(\delta^{\mathrm{T}}_{ij}(\vect{r})=\int\dd^{3}k\,(2\pi)^{-3}\, \widetilde{\delta}^{\mathrm{T}}_{ij}(\vect{k}) \ee^{\ii\vect{k}\cdot\vect{r}}\). The gauge condition \(\nabla\cdot\vect{A}=0\) is \(\vect{k}\cdot\widetilde{\vect{A}}=0\), so \(\widetilde{\delta}^{\mathrm{T}}_{ij}\) must annihilate every longitudinal vector and act as the identity on every transverse one. A symmetric tensor built from \(\delta_{ij}\) and \(\vect{k}\) with those two properties is unique:
which indeed obeys \(k_{i}\widetilde{\delta}^{\mathrm{T}}_{ij}=0\) and \(\widetilde{\delta}^{\mathrm{T}}_{ik} \widetilde{\delta}^{\mathrm{T}}_{kj} =\widetilde{\delta}^{\mathrm{T}}_{ij}\), and whose trace is \(2\): the two transverse polarizations of Equation (99.80).
Now transform back term by term. The \(\delta_{ij}\) gives \(\delta_{ij}\delta^{3}(\vect{r})\). For the second term use \(\int\dd^{3}k\,(2\pi)^{-3}\vect{k}^{-2} \ee^{\ii\vect{k}\cdot\vect{r}}=1/(4\pi\abs{\vect{r}})\), which is the Coulomb Green function, together with the fact that \(\pp_{i}\pp_{j}\) acting on \(\ee^{\ii\vect{k}\cdot\vect{r}}\) brings down \((\ii k_{i})(\ii k_{j})=-k_{i}k_{j}\). Hence \(-k_{i}k_{j}/\vect{k}^{2}\) transforms to \(+\pp_{i}\pp_{j}(4\pi\abs{\vect{r}})^{-1}\), and the sign in Equation (99.81) is a plus.
The position-space statement confirms it independently: \(\pp_{i}\delta^{\mathrm{T}}_{ij} =\pp_{j}\delta^{3}(\vect{r}) +\pp_{j}\nabla^{2}(4\pi\abs{\vect{r}})^{-1} =\pp_{j}\delta^{3}(\vect{r})-\pp_{j}\delta^{3}(\vect{r})=0\), whereas a minus sign would leave \(2\pp_{j}\delta^{3}(\vect{r})\). Carrying out the derivatives gives the explicit form \(\delta^{\mathrm{T}}_{ij}(\vect{r}) =\tfrac{2}{3}\delta_{ij}\delta^{3}(\vect{r}) +(3\hat{r}_{i}\hat{r}_{j}-\delta_{ij})/(4\pi r^{3})\), whose second piece has the angular structure of a point dipole's field. Finally, the prefactor: the canonical relation is \(\comm{A_{i}}{\pi^{j}}=\ii\hbar\,\delta^{\mathrm{T}}_{ij}\) with \(\pi^{j}=\varepsilon_{0}E^{j}\) by Equation (99.79), and \(E_{j}=-E^{j}\) supplies the minus sign shown.
∎Covariant (Gupta–Bleuler) quantization. Add to Equation (99.78) the gauge-fixing term \(-(2\mu_{0}\xi)^{-1}(\pp_{\mu}A^{\mu})^{2}\), which makes \(\dot{A}_{0}\) appear and removes the constraint by hand (Equation (100.16)). All four polarizations are then quantized covariantly, at the cost that the timelike one has the wrong sign:
so that \(\bra{0}a_{0}a^{\dagger}_{0}\ket{0}<0\) and the state space has an indefinite metric — there are states of negative norm, which cannot be probabilities.
The indefinite-metric mode algebra of covariant quantization, displayed just above: that the gauge-fixing term makes all four components of the potential dynamical, that the covariant equal-time bracket \(\comm{A_{\mu}}{\pi^{\nu}}\) then delivers \(-\eta_{\lambda\lambda'}\) where the Coulomb-gauge algebra has \(\delta_{\lambda\lambda'}\), and hence that the timelike mode has negative norm. It needs a basis of four polarization vectors orthonormal in the Minkowski metric, which this chapter does not set up, so it belongs in the long-proof appendix.
Gupta and Bleuler [Gupta:1950] [Bleuler:1950] showed that the theory is nevertheless consistent if the physical states are restricted by the subsidiary condition
the superscript denoting the positive-frequency (annihilation) part. On the subspace so defined the norm is non-negative, the timelike and longitudinal quanta enter only in combinations of zero norm, and their contributions to every observable cancel, leaving the two transverse polarizations of Equation (99.80).
Consistency of the Gupta–Bleuler subsidiary condition displayed just above: that the subspace it selects carries a non-negative norm, that timelike and longitudinal quanta occur in it only in zero-norm combinations, and that those combinations contribute nothing to any observable, so that the counting reduces to the two transverse polarizations. It runs through the mode decomposition of the gauge-fixed field and belongs in the long-proof appendix, next to the indefinite-metric algebra above.
The gauge parameter \(\xi\) survives in the propagator (Equation (100.17)) and must cancel from every observable, a requirement that is one of the strongest internal checks on a QED calculation (Quantum Electrodynamics and Renormalization).
That the photon has two states and not three is not a choice of gauge: it is Remark 99.26. A massless particle is classified by helicity, and the two helicities \(\pm1\) are the whole content of the representation; the third polarization of a massive spin-\(1\) particle has nowhere to go as \(m\to0\). Gauge redundancy is the field-theoretic expression of that representation-theoretic fact, and this is the sense in which the difficulty of the constraints is not a technical nuisance but a consequence of masslessness.
What has been suppressed here, and belongs to Quantum Optics and the Photon, is the phenomenology of the photon number operator: the coherent states that are eigenstates of \(a_{\lambda}\) and approximate a classical wave, the photon-counting statistics that distinguish them from thermal and number states, and the experiments that measure those statistics directly.
Massive vector fields and the general spin case
Adding a mass term to Equation (99.78) gives Proca's theory [Proca:1936],
whose field equation is \(\pp_{\mu}F^{\mu\nu}+\kappa^{2}A^{\nu}=0\). The mass term breaks gauge invariance, and the constraint structure changes character.
For \(\kappa\neq0\) the Proca field satisfies \(\pp_{\mu}A^{\mu}=0\) as a consequence of its equation of motion, not as a gauge choice; each component then obeys \((\Box+\kappa^{2})A^{\nu}=0\), and the physical polarizations number three. Rests on Equation (99.85) and Proposition 94.24.
Derives Proposition 99.31. Take \(\pp_{\nu}\) of the field equation. The first term vanishes identically, \(\pp_{\nu}\pp_{\mu}F^{\mu\nu}=0\), being a symmetric derivative contracted with an antisymmetric tensor, so \(\kappa^{2}\pp_{\nu}A^{\nu}=0\), and for \(\kappa\neq0\) this forces \(\pp_{\nu}A^{\nu}=0\). Substituting back, \(\pp_{\mu}(\pp^{\mu}A^{\nu}-\pp^{\nu}A^{\mu}) =\Box A^{\nu}\), so \((\Box+\kappa^{2})A^{\nu}=0\). Four components subject to one covariant condition leave three; equivalently, the little group of a massive momentum is \(\SU(2)\) (Proposition 94.24) and the spin-\(1\) representation is three-dimensional.
∎Since there is no gauge freedom, \(\pi^{0}=0\) is still a constraint but it now determines \(A^{0}\) in terms of the other variables rather than generating a redundancy, and the canonical quantization of the three remaining polarizations proceeds as for three scalars. The consequence that matters downstream is the propagator. Inverting the quadratic form of Equation (99.85) in the same way as Proposition 100.8 does for the photon gives
and the numerator's second term does not fall off at large \(k\): the propagator tends to a constant \(\propto k^{\mu}k^{\nu}/(\kappa^{2}k^{2})\) rather than to zero. In a theory in which massive vector bosons are exchanged between currents that are not conserved, amplitudes then grow with energy without limit, violating the unitarity bound at a calculable scale. That growth is the argument — not an aesthetic preference for symmetry — that forces the weak interaction to be a spontaneously broken gauge theory rather than a Proca theory with arbitrary couplings, and it is carried out quantitatively in Weak Interactions.
For spin greater than \(1\) the pattern continues and worsens. A free field of spin \(s\) can be constructed as a symmetric traceless tensor subject to auxiliary conditions, and Fierz and Pauli [Fierz:1939] showed that the conditions can be imposed consistently for the free field; but coupling such a field to an external electromagnetic field generically destroys the constraint algebra, producing either fewer propagating modes than the spin requires or propagation faster than light. The observed spectrum contains no elementary particle of spin greater than \(1\) — the hadronic resonances of higher spin are composite — and the obstructions are summarized in Higher-Spin Wave Equations.
Propagators
A propagator is the vacuum expectation value of a time-ordered product of two fields. It is the object that perturbation theory is built from, it is a Green function of the classical field equation, and the particular Green function it is — which of the possible contours — is fixed by the requirement that positive frequencies propagate forward in time and negative frequencies backward. This section derives it, and derives it by contour integration, using the residue calculus of Section 8.7 rather than rebuilding it.
Time ordering and the Feynman propagator
For bosonic fields,
and the Feynman propagator, Wightman function and Pauli–Jordan function of the free real scalar are
of SI dimension \(\mathrm{J}/\mathrm{m}\); it solves \(\left(\Box+\kappa^{2}\right)D^{+}=0\) and is manifestly Lorentz invariant. Rests on Equations (99.29), (99.31) and (99.37).
Derives Lemma 99.33. Substitute Equation (99.29) twice. Of the four terms only \(a(\vect{k})a^{\dagger}(\vect{k}')\) survives between vacua, because \(a\ket{0}=0=\bra{0}a^{\dagger}\); using \(\bra{0}a(\vect{k})a^{\dagger}(\vect{k}')\ket{0} =(2\pi)^{3}\delta^{3}(\vect{k}-\vect{k}')\) from Equation (99.31) collapses the double integral to the first form. The second form is Equation (99.37). Both exponents lie on the mass shell, so \(\left(\Box+\kappa^{2}\right)\) annihilates the integrand: \(\Box\ee^{-\ii k\cdot x}=-k^{2}\ee^{-\ii k\cdot x}\) and \(k^{2}=\kappa^{2}\) on the support of the delta. The dimension is \([\hbar c^{2}/\omega]/\mathrm{m}^{3} =\mathrm{J}\,\mathrm{m}^{2}\cdot/\mathrm{m}^{3}\).
∎where \(\delta^{4}(x)=\delta(x^{0})\delta^{3}(\vect{x})\) carries \(\mathrm{m}^{-4}\); both sides carry \(\mathrm{J}/\mathrm{m}^{3}\). Rests on Equations (99.23), (99.24) and (99.87).
Derives Theorem 99.34. Differentiate Equation (99.87) with respect to \(t=x^{0}/c\), using \(\pp_{t}\theta(x^{0}-y^{0})=\delta(t_{x}-t_{y})\):
and the first term vanishes by Equation (99.24). Differentiating again, the delta now multiplies the equal-time commutator of \(\pp_{t}\phi=c^{2}\pi\) with \(\phi\):
using Equation (99.23). The \(\theta\)-terms are the time-ordered product of \(\pp_{t}^{2}\phi\) with \(\phi\), and the spatial and mass parts of \(\left(\Box+\kappa^{2}\right)\) pass through the \(\theta\) functions untouched, so
whose second term vanishes by the field equation. Finally \(\delta(t_{x}-t_{y})=c\,\delta(x^{0}-y^{0})\), giving Equation (99.92).
∎with \(\epsilon\to0^{+}\), and this — and no other choice of contour — reproduces Equation (99.88). In terms of four-momenta \(p=\hbar k\) with the measure \(\dd^{4}p/(2\pi\hbar)^{4}\) the same statement reads
Derives Theorem 99.35. That Equation (99.93) solves Equation (99.92) is immediate: \(\Box\to-k^{2}\) turns the integrand into \(\ii\hbar c(\kappa^{2}-k^{2})/(k^{2}-\kappa^{2}) =-\ii\hbar c\), whose Fourier transform is \(-\ii\hbar c\,\delta^{4}(x-y)\). The content of the theorem is that the \(\ii\epsilon\) is the right prescription, and that is decided by doing the \(k^{0}\) integral.
Write \(k^{2}-\kappa^{2}=(k^{0})^{2}-\vect{k}^{2}-\kappa^{2} =(k^{0})^{2}-(\omega_{\vect{k}}/c)^{2}\) by Equation (99.28). The \(\ii\epsilon\) displaces the two poles of the integrand off the real \(k^{0}\) axis:
with \(\epsilon'=c\epsilon/2\omega_{\vect{k}}>0\): the positive-frequency pole sits below the real axis and the negative-frequency pole above it. Since \(x^{0}=ct\), the exponential is \(\ee^{-\ii k^{0}(x^{0}-y^{0})}\), which decays exponentially as \(\Im k^{0}\to-\infty\) when \(x^{0}>y^{0}\) and as \(\Im k^{0}\to+\infty\) when \(x^{0}<y^{0}\). The integrand also falls as \(\abs{k^{0}}^{-2}\), so a large semicircle in the half plane where the exponential decays contributes nothing as its radius grows and the contour may be closed there; the residue theorem of Section 8.7 then evaluates the integral.
For \(x^{0}>y^{0}\), closing below encircles the pole at \(k^{0}=\omega_{\vect{k}}/c\) clockwise, contributing \(-2\pi\ii\) times the residue:
which, reinserted into the remaining \(\dd^{3}k\) integral, is exactly \(D^{+}(x-y)\) of Equation (99.91). For \(x^{0}<y^{0}\), closing above encircles \(k^{0}=-\omega_{\vect{k}}/c\) counterclockwise and gives \((\hbar c^{2}/2\omega_{\vect{k}}) \ee^{+\ii\omega_{\vect{k}}(t_{x}-t_{y})}\), which after \(\vect{k}\to-\vect{k}\) in the spatial integral is \(D^{+}(y-x)\). Together these are Equation (99.88).
Any other placement of the poles gives a different function. Displacing both poles downward gives zero for \(x^{0}<y^{0}\) — the retarded function of Section 99.6.2 — and both upward gives the advanced one. The Feynman choice is the unique one that propagates the positive-frequency part forward and the negative-frequency part backward in time, which is the analytic statement of the Stueckelberg–Feynman reading of Section 99.3.2 [Stueckelberg:1941] [Feynman:1949a] [Feynman:1949b]. Finally, \(k=p/\hbar\) turns \(k^{2}-\kappa^{2}\) into \((p^{2}-m^{2}c^{2})/\hbar^{2}\) and \(\dd^{4}k/(2\pi)^{4}\) into \(\dd^{4}p/(2\pi\hbar)^{4}\), giving Equation (99.94); its dimension follows from \([\hbar^{3}c]=\mathrm{J}^{3}\,\mathrm{s}^{3}\,\mathrm{m}/\mathrm{s}\) divided by \([p^{2}]\).
∎Equation (99.94) is the scalar member of the family whose fermionic and photonic partners are Equations (100.17) and (100.23); the \(\ii\epsilon\) there is the one fixed here, which is why Proposition 100.8 refers the prescription back to this chapter. Note that the pole is at \(p^{2}=m^{2}c^{2}\), the mass shell Equation (100.1): a propagator carries information about a particle precisely through the position of its pole, which is what makes the pole a definition of mass in the interacting theory (Quantum Electrodynamics and Renormalization) and what the Källén–Lehmann representation [Kallen:1952] [Lehmann:1954] generalizes.
At spacelike separation, evaluated in the frame where the two points are simultaneous and a distance \(r\) apart,
with \(K_{1}\) the modified Bessel function of the second kind: a solution of the modified Bessel equation \(x^{2}y''+xy'-(x^{2}+\nu^{2})y=0\), which is the Bessel equation of Section 9.8 continued to imaginary argument, and which that section does not itself treat. It is small but not zero. Rests on Equations (99.88) and (99.91).
Derives Proposition 99.36. At equal times Equation (99.91) reduces to \(D^{+}=\tfrac12\hbar c\int\dd^{3}k\,(2\pi)^{-3} \ee^{\ii\vect{k}\cdot\vect{r}}(\vect{k}^{2}+\kappa^{2})^{-1/2}\). Doing the angular integral, \(\int\dd^{3}k\,(2\pi)^{-3}f(\abs{\vect{k}}) \ee^{\ii\vect{k}\cdot\vect{r}} =(2\pi^{2}r)^{-1}\int_{0}^{\infty}\dd k\,k\sin(kr)f(k)\), so
the last step a standard integral representation of \(K_{1}\); the check \(\kappa\to0\) returns \(\int_{0}^{\infty}\sin(kr)\dd k=1/r\) and \(D^{+}\to\hbar c/(4\pi^{2}r^{2})\), the massless Coulomb-like falloff. The asymptotic form is \(K_{1}(z)\sim\sqrt{\pi/2z}\,\ee^{-z}\). That \(D_{F}=D^{+}\) here is because at spacelike separation the two orderings give the same answer, which is Proposition 99.44.
∎The exponential in Equation (99.96) has range \(\kappa^{-1}=\hbar/mc\): the reduced Compton wavelength of Equation (99.10) appears once more, now as the distance over which a virtual quantum of mass \(m\) can carry an influence. The same function, sourced by a static charge, is the Yukawa potential of The Klein–Gordon Equation. That \(D_{F}\) does not vanish outside the light cone is often reported as a violation of causality. It is not, and Section 99.7.1 explains exactly why.
The family of Green functions
Theorem 99.35 obtained one Green function of the Klein–Gordon operator by one choice of contour. The other choices are equally legitimate solutions of Equation (99.92) — they differ by solutions of the homogeneous equation — and each answers a different physical question. All are built from the same integrand, so the family is best displayed as one formula with four contours.
with both poles pushed into the lower half plane, and \(D_{A}\) with \(\ii\epsilon\to-\ii\epsilon\), pushing both into the upper half plane.
Closing the contour as in Theorem 99.35 shows at once that \(D_{R}(x)=0\) for \(x^{0}<0\) (the contour closes in the upper half plane, which is now empty of poles) and \(D_{A}(x)=0\) for \(x^{0}>0\). \(D_{R}\) is therefore the classical response function: it gives the field produced after a source acts, and it is the object behind the retarded potentials of Radiation and Scattering of Electromagnetic Waves, whose mathematics — the retarded and advanced Green functions of the wave operator, and the causality condition that selects the retarded one — is Partial Differential Equations. It is not the object that appears in perturbation theory, because a scattering amplitude is not a response to a classical source but a matrix element of a time-ordered product, and Equation (99.87) is what the expansion of \(\exp(\ii S/\hbar)\) generates.
With \(\Delta\) of Equation (99.90),
and \(\Delta\) is a \(c\)-number, not an operator. Rests on Equation (99.90), Definition 99.37 and Equation (99.91).
Derives Proposition 99.38. Equation (99.98) is the definition of the commutator written out, and Equation (99.99) is Equation (99.87) between vacua. For Equation (99.100), subtract the two contours of Definition 99.37: the difference is a closed contour encircling both poles clockwise, so it is \(-2\pi\ii\) times the sum of the residues, which is \(\int\dd^{3}k(2\pi)^{-3}(\hbar c^{2}/2\omega_{\vect{k}}) (\ee^{-\ii k\cdot x}-\ee^{\ii k\cdot x}) =D^{+}(x)-D^{+}(-x)\) by Equation (99.91), and that is \(\ii\hbar\,\Delta\) by Equation (99.98). The \(\hbar\) and no \(c\) is what the dimensions require: \(D_{R}\), \(D_{A}\) and \(D^{+}\) all carry \(\mathrm{J}/\mathrm{m}\) while \([\Delta]=/\mathrm{m}/\mathrm{s}\), so only \(\hbar\Delta\) matches. That \(\Delta\) is a \(c\)-number is Proposition 99.44: the commutator of two free fields is proportional to the identity, so its vacuum expectation value determines it completely.
∎| Function | Contour in $k^{0}$ | Question it answers |
|---|---|---|
| $D_{F}$ | poles at $+\omega/c$ and $-\omega/c$ displaced into the lower and upper half plane respectively | What does a time-ordered product between vacua give? The building block of the perturbation series |
| $D_{R}$ | both poles below the axis; vanishes for $x^{0}<0$ | What field does a source produce afterwards? Classical radiation |
| $D_{A}$ | both poles above the axis; vanishes for $x^{0}>0$ | What source would have produced a given field? |
| $D^{+}$ | positive-frequency mass shell only | What is the vacuum correlation of two field measurements, in a definite order? |
| $\Delta$ | both signs of frequency, no $\theta$ | Do measurements at $x$ and $y$ interfere? Zero outside the light cone |
The same construction applies to the other fields, with two differences.
With the fermionic time ordering \(T\{\psi(x)\bar{\psi}(y)\} =\theta(x^{0}-y^{0})\psi(x)\bar{\psi}(y) -\theta(y^{0}-x^{0})\bar{\psi}(y)\psi(x)\), the propagator \(S_{F}(x-y)=\bra{0}T\{\psi(x)\bar{\psi}(y)\}\ket{0}\) satisfies
whose momentum-space solution is Equation (100.23),
Derives Proposition 99.39. The minus sign in the fermionic \(T\) product is not a convention but what makes the derivative of the \(\theta\) functions produce an anticommutator: differentiating with respect to \(x^{0}\),
using \(\acomm{\psi_{a}}{\bar{\psi}_{b}} =\acomm{\psi_{a}}{\psi^{\dagger}_{c}}\gamma^{0}_{cb} =\gamma^{0}_{ab}\delta^{3}(\vect{x}-\vect{y})\) from Equation (99.75) and \((\gamma^{0})^{2}=\identity\). The remaining terms are the \(T\) product of the Dirac equation with \(\bar{\psi}\) and vanish. In momentum space \(\ii\hbar\gamma^{\mu}\pp_{\mu}\to\gamma^{\mu}p_{\mu}\), and the second equality in Equation (99.102) follows from \((\gamma^{\mu}p_{\mu}+mc)(\gamma^{\nu}p_{\nu}-mc) =p^{2}-m^{2}c^{2}\), itself a consequence of \(\acomm{\gamma^{\mu}}{\gamma^{\nu}}=2\eta^{\mu\nu}\). Had the \(T\) product been defined with a plus sign, the derivative would have produced a commutator, which by Equation (99.75) is not a \(c\)-number, and no Green function would have resulted.
∎The photon propagator is different again, because there is no unique answer: the field \(A_{\mu}\) is not gauge invariant, so its two-point function depends on the gauge condition. In Coulomb gauge it is non-covariant and contains the instantaneous Coulomb term visible in Equation (99.81); in the covariant gauges of Section 99.5.2 it is Equation (100.17), carrying the free parameter \(\xi\). The resolution is that the gauge dependence cancels between diagrams whenever the external currents are conserved, so that every measurable quantity is gauge independent — a statement proved, along with the Ward–Takahashi identities that organize it, in Quantum Electrodynamics and Renormalization. That such a cancellation must happen is guaranteed in advance by the fact that the physical states of Equation (99.84) are gauge invariant; that it does happen diagram by diagram is a nontrivial calculation.
Wick's theorem
Perturbation theory expands \(\exp(\ii S_{\text{int}}/\hbar)\) and produces vacuum expectation values of time-ordered products of many fields. Wick's theorem [Wick:1950] reduces every such object to propagators, and in doing so turns an operator calculation into combinatorics.
Split the free field into its annihilation and creation parts, \(\phi=\phi^{+}+\phi^{-}\), where \(\phi^{+}\) collects the \(a(\vect{k})\) terms of Equation (99.29) and \(\phi^{-}\) the \(a^{\dagger}(\vect{k})\) terms. The contraction of two fields is the \(c\)-number
That the contraction is the Feynman propagator is the two-field case of the theorem and is worth doing explicitly. For \(x^{0}>y^{0}\),
since normal ordering differs from the plain product only by the one commutator generated in moving \(\phi^{+}(x)\) past \(\phi^{-}(y)\); and \(\comm{\phi^{+}(x)}{\phi^{-}(y)} =\bra{0}\phi(x)\phi(y)\ket{0}=D^{+}(x-y)\), because the normal-ordered term has zero vacuum expectation value. For \(y^{0}>x^{0}\) the same with \(x\leftrightarrow y\). Comparing with Equation (99.99) gives Equation (99.103).
For free fields,
the sum running over all ways of contracting zero, one, two, … pairs, the uncontracted fields left inside the normal ordering. Consequently
with \((2r-1)!!\) terms for \(n=2r\). Rests on Definition 99.14 and Equation (99.103).
Derives Theorem 99.41. By induction on \(n\). The case \(n=2\) is Equation (99.103). Suppose Equation (99.104) holds for \(n-1\) fields and let the fields be labelled so that \(x_{1}^{0}>x_{2}^{0}>\cdots>x_{n}^{0}\), which is no loss since both sides are symmetric under relabelling. Then \(T\{\phi_{1}\cdots\phi_{n}\} =\phi_{1}\,T\{\phi_{2}\cdots\phi_{n}\}\), and by hypothesis the second factor is a sum of normal-ordered products of subsets of \(\phi_{2},\ldots,\phi_{n}\) times contractions. It therefore suffices to show that
where the hat omits a factor. Write \(\phi_{1}=\phi_{1}^{+}+\phi_{1}^{-}\). The creation part \(\phi_{1}^{-}\) is already to the left of everything and contributes the first term. The annihilation part must be moved through the creation operators inside the normal product; each such move generates one commutator \(\comm{\phi_{1}^{+}}{\phi_{j}^{-}}\), and since \(x_{1}^{0}\) is the latest time, \(\comm{\phi_{1}^{+}}{\phi_{j}^{-}}=D^{+}(x_{1}-x_{j}) =D_{F}(x_{1}-x_{j})=\overline{\phi_{1}\phi_{j}}\). The commutators are \(c\)-numbers, so they factor out of the normal ordering, leaving the sum shown. Equation (99.105) follows because \(\bra{0}:\!X\!:\ket{0}=0\) unless \(X\) is empty, so only the terms in which every field is contracted survive; the number of ways of pairing \(2r\) objects is \((2r-1)!!\).
∎Once an interaction is switched on, Equation (99.105) is the Feynman rules: each contraction is an internal line carrying Equation (99.94), each interaction vertex is one factor from the expansion of \(\exp(\ii S_{\text{int}}/\hbar)\), and the sum over pairings is the sum over diagrams, with the symmetry factors counting the pairings that give the same diagram. That translation is carried out in Proposition 100.14.
For anticommuting fields the same expansion holds with two changes: the contraction is \(\overline{\psi(x)\bar{\psi}(y)}=S_{F}(x-y)\) of Equation (99.102), and each term carries the sign of the permutation needed to bring the contracted pairs adjacent. A closed loop of \(n\) fermion propagators therefore contributes an extra factor \(-1\). Rests on Theorem 99.41 and Equation (99.102).
Derives Proposition 99.42. The induction of Theorem 99.41 goes through verbatim with commutators replaced by anticommutators, provided every interchange of two fermion operators is accompanied by a minus sign; that is the origin of the permutation signature. For the loop, consider \(n\) vertices joined in a cycle, so that the contractions are \(\overline{\psi(x_{1})\bar{\psi}(x_{2})}, \overline{\psi(x_{2})\bar{\psi}(x_{3})},\ldots, \overline{\psi(x_{n})\bar{\psi}(x_{1})}\). Bringing each \(\bar{\psi}(x_{i+1})\) next to its partner \(\psi(x_{i})\) can be done with an even number of transpositions for the first \(n-1\) contractions, but the last one must move \(\bar{\psi}(x_{1})\) past all the remaining \(2n-2\) operators and past the closing of the cycle, which is one odd permutation — a cyclic permutation of an even number of anticommuting objects. The net factor is \(-1\), independent of \(n\), and the resulting expression is the Dirac trace around the loop.
∎The extra minus sign of Proposition 99.42 is not a bookkeeping detail. It is what makes the vacuum polarization of Quantum Electrodynamics and Renormalization screen rather than antiscreen the charge, and hence what makes the fine-structure constant increase with energy while the strong coupling of Quantum Chromodynamics decreases: the sign of the leading term in the beta function of The Renormalization Group is decided by the competition between fermion loops carrying this \(-1\) and gauge-boson loops which do not.
Causality and spin–statistics
Special relativity forbids influence outside the light cone. Quantum mechanics makes influence a statement about operators, not about trajectories, so the relativistic requirement has to be restated as an algebraic condition on the fields. The condition is microcausality, it is satisfied by the free fields constructed above, and — this is the remarkable part — it is satisfied only if bosons are quantized with commutators and fermions with anticommutators.
Microcausality
A field theory satisfies microcausality if every pair of local observables commutes at spacelike separation:
For the free real scalar field,
a \(c\)-number times the identity, Lorentz invariant, and equal to zero for every spacelike separation \((x-y)^{2}<0\). Rests on Equations (99.29), (99.31) and (99.91).
Derives Proposition 99.44. Substituting Equation (99.29) twice and using Equation (99.31),
which is a \(c\)-number because Equation (99.31) is: the operators have cancelled. Writing each Wightman function in the covariant form of Equation (99.91) and sending \(k\to-k\) in the second gives Equation (99.107), since \(\theta(k^{0})-\theta(-k^{0})=\sgn(k^{0})\).
Invariance: \(\dd^{4}k\) and \(\delta(k^{2}-\kappa^{2})\) are invariant under the full Lorentz group, and \(\sgn(k^{0})\) is invariant under the orthochronous subgroup, because a transformation that preserves the direction of time cannot move a four-vector from the forward mass hyperboloid to the backward one — the two sheets \(k^{0}>0\) and \(k^{0}<0\) of \(k^{2}=\kappa^{2}\) are disconnected and each is preserved (Proposition 94.21).
Vanishing: take first the case of equal times, \(x^{0}=y^{0}\), and write \(\vect{r}=\vect{x}-\vect{y}\). Then
because \(\omega_{\vect{k}}\) depends only on \(\abs{\vect{k}}\), so substituting \(\vect{k}\to-\vect{k}\) in the second term turns it into the first and the two cancel. Now let \((x-y)\) be any spacelike vector. There exists a proper orthochronous Lorentz transformation taking it to a purely spatial vector — that is the definition of spacelike — and the left side is invariant under such a transformation by the paragraph above. Hence it vanishes there too.
Note that the same argument fails for timelike separation, where no such transformation exists, and indeed it must fail: two timelike separated measurements can influence one another, and a formalism in which they could not would be describing nothing.
∎Let \(\mathcal{O}_{1}\) be an observable localized in a region \(R_{1}\) and let a projective measurement of it be performed. Then the expectation value of any observable \(\mathcal{O}_{2}\) localized in a region \(R_{2}\) spacelike to \(R_{1}\) is unchanged. Rests on Equation (99.106) and Proposition 99.44.
Derives Corollary 99.45. Let \(\set{P_{i}}\) be the spectral projectors of \(\mathcal{O}_{1}\); they are functions of \(\mathcal{O}_{1}\) and hence localized in \(R_{1}\), so \(\comm{P_{i}}{\mathcal{O}_{2}}=0\) by Equation (99.106). A measurement whose outcome is not recorded takes the state \(\rho\) to \(\sum_{i}P_{i}\rho P_{i}\), so
using cyclicity, then \(\comm{P_{i}}{\mathcal{O}_{2}}=0\), then \(\sum_{i}P_{i}=\identity\).
∎Corollary 99.45 is what microcausality buys, and it is exactly the right amount. It does not say that spacelike separated measurements are uncorrelated — they are correlated, as Bell-type experiments show and as the Reeh–Schlieder theorem [Reeh:1961] makes vivid by showing that the vacuum itself is entangled across any spacelike cut. It says that no local operation can change what a spacelike observer will see, which is the only thing relativity requires.
This resolves the apparent paradox of Proposition 99.36. The propagator \(D_{F}(x-y)=\bra{0}T\phi(x)\phi(y)\ket{0}\) is not zero outside the light cone; it falls only as \(\ee^{-\kappa r}\). But \(D_{F}\) is not an observable, and it is not a probability: it is one term in a perturbative expansion, and a matrix element of an ordered product. The observable statement is the commutator, and there the two orderings enter with opposite signs and cancel exactly outside the cone. The particle and antiparticle contributions — the two exponentials in Equation (99.107) — destructively interfere. That cancellation is another way of saying that antiparticles are not optional: a theory with a particle and no antiparticle of equal mass would have a nonvanishing spacelike commutator and would signal faster than light.
The axiomatic programme of Axiomatic Quantum Field Theory takes Equation (99.106) as one of the Wightman axioms [Wightman:1956] [Streater:1964] rather than deriving it from a Lagrangian, which is the right order of business: it is the physical requirement, and the Lagrangian is only one way of meeting it.
The spin–statistics theorem
Two facts have now been established separately. A spin-\(1/2\) field quantized with commutators has an energy unbounded below (Theorem 99.28). A scalar field quantized with commutators satisfies microcausality (Proposition 99.44). Pauli [Pauli:1940] showed that these are two instances of one theorem.
Let a relativistic field theory satisfy
-
invariance under the proper orthochronous Poincaré group;
-
a Hamiltonian bounded below, with a Poincaré-invariant state of lowest energy;
-
microcausality, Equation (99.106).
Then fields of integer spin must be quantized with commutators and fields of half-integer spin with anticommutators. Dropping either choice forces the abandonment of one of the three hypotheses. Rests on Equation (99.106).
The theorem in this generality requires the analytic continuation of the Wightman functions to complex Lorentz transformations, and belongs in the axiomatic setting where the hypotheses can be stated without reference to any Lagrangian. The axiomatic proofs are Burgoyne's [Burgoyne:1958] and, independently, Lüders and Zumino's [Lueders:1958]; both rest on the analyticity domain Jost established in his work on the \(CPT\) theorem [Jost:1957], which is a different theorem. The standard textbook account is [Streater:1964], and Axiomatic Quantum Field Theory returns to it.
The spin–statistics theorem for arbitrary spin, by analytic continuation of the Wightman functions to the Jost points: this is a long argument and belongs in the long-proof appendix. The two cases relevant to the Standard Model — spin \(0\) and spin \(1/2\) — are proved in full in the text below.
What can be done here in full is the two cases that matter, and they already display the mechanism: each wrong choice destroys one hypothesis, and a different one in each case.
Commutators for the Dirac field of Equation (99.68) contradict hypothesis (ii). Rests on Equation (99.68) and Theorem 99.28.
Derives Proposition 99.47. This is Theorem 99.28: with commutators, Equation (99.73) becomes \(\sum_{s}\int_{\vect{k}}\hbar\omega_{\vect{k}} (b^{\dagger}b-d^{\dagger}d)\) up to a constant, whose spectrum extends to \(-\infty\) because \(d^{\dagger}d\) has arbitrarily large eigenvalues. There is then no state of lowest energy, so no vacuum in the sense of hypothesis (ii).
∎Anticommutators for the real scalar field of Equation (99.18) contradict hypothesis (iii), and indeed destroy the theory entirely. Rests on Equation (99.18), Proposition 99.44 and Equation (99.96).
Derives Proposition 99.48. Suppose \(\acomm{a(\vect{k})}{a^{\dagger}(\vect{k}')} =(2\pi)^{3}\delta^{3}(\vect{k}-\vect{k}')\) with all other anticommutators vanishing, and expand \(\phi\) as in Equation (99.29). Repeating the computation of Proposition 99.44 with anticommutators in place of commutators changes one sign, giving
in which the two terms now add instead of cancelling. At spacelike separation both equal Equation (99.96), which is strictly positive, so Equation (99.108) does not vanish outside the light cone.
Worse, put \(y=x\). Then \(\acomm{\phi(x)}{\phi(x)}=2\phi(x)^{2}\) equals the \(c\)-number \(2D^{+}(0)\), so \(\phi(x)^{2}\) is a multiple of the identity and
identically. Every even local polynomial in \(\phi\) is therefore a \(c\)-number and every odd one is linear in \(\phi\): the mass term in Equation (99.21) is a constant, an interaction \(\lambda\phi^{4}\) is a constant, and the theory has no nontrivial local observable at all. The would-be anticommuting scalar is not a theory with the wrong causal structure; it is not a theory.
∎It is worth being exact about the logic, because Theorem 99.46 is often quoted as though it derived statistics from spin alone. It does not. It says that the three hypotheses are jointly incompatible with the wrong pairing. One can build a theory of anticommuting scalars — Faddeev–Popov ghosts are exactly that, and they appear in Quantum Chromodynamics and Path-Integral Quantization — but such fields are not observables, do not appear as asymptotic states, and exist only to cancel unphysical polarizations inside loops. They violate the spin–statistics theorem precisely by not being subject to its hypotheses.
The theorem has consequences at every scale, and they are among the most thoroughly verified facts in physics.
Atomic structure. Exclusion applied to atomic electrons is what produces shells, the periodic table and chemistry (Atoms and Molecules); without it every electron would occupy the \(1s\) orbital and all elements would be chemically alike.
Stellar remnants. The pressure of a degenerate electron gas supports white dwarfs and that of a degenerate neutron gas supports neutron stars, with the Chandrasekhar limit following from the competition between degeneracy pressure and gravity (Compact Stars and Relativistic Astrophysics). Degeneracy pressure is Equation (99.77) expressed thermodynamically.
Direct tests. The exclusion principle is not merely assumed but bounded. Ramberg and Snow [Ramberg:1990] passed a large current through a thin copper strip and searched for the X-rays that would be emitted if a fresh conduction electron made a transition into an already full \(1s\) shell — a transition whose energy differs from the ordinary \(K_{\alpha}\) line by a calculable amount, so that the search is background-free in a narrow window. None were seen, bounding the probability of a Pauli-violating transition at \(\beta^{2}/2<1.7\times 10^{-26}\). The VIP-2 experiment [Napolitano:2022] repeats the measurement in the low-background environment of an underground laboratory with a silicon drift detector of much better energy resolution, improving on that limit. A violation at any level would falsify one of the three hypotheses of Theorem 99.46, which is why the searches are worth the effort even though nobody expects a signal.
The same three hypotheses appear again in the \(CPT\) theorem [Luders:1957] [Jost:1957], and that is not a coincidence: both theorems are proved from the analyticity that Lorentz invariance plus positivity of the energy forces on the Wightman functions. The connection is drawn in Discrete Symmetries and CPT.
Measurability of the field operators
Equation (99.27) said that the field at a point is not an operator and must be smeared. That is a mathematical statement. Bohr and Rosenfeld [Bohr:1933] asked the corresponding physical question — what can a real apparatus actually measure of a quantized electromagnetic field? — and found that the two answers agree.
Their apparatus is a charged test body of finite extent and finite mass, whose displacement under the field is read off. Such a body cannot measure \(\vect{E}\) at a point: it necessarily averages over its own volume and over the time it is exposed. Nor can it be made arbitrarily small, because the momentum transfer needed to read a smaller displacement grows and the body's own radiation reacts back on the field being measured. What it measures is
an average over a spacetime region, which is Equation (99.27) with the test function equal to the indicator of that region.
Bohr and Rosenfeld's conclusion was that the uncertainties unavoidably introduced by such an apparatus are exactly those demanded by the commutators of the averaged field strengths, no more and no less: the formalism claims neither more measurability than an apparatus can deliver, nor less. The structural part of that conclusion is a corollary of what has already been proved here.
Let \(R_{1}\) and \(R_{2}\) be spacetime regions every point of which is spacelike separated from every point of the other. Then the field strengths averaged over them, in the sense of Equation (99.109), are compatible observables: the averages commute, \(\comm{\overline{E}_{i}(R_{1})}{\overline{B}_{j}(R_{2})}=0\), and the two averages may be measured simultaneously with no mutual disturbance. Rests on Equation (99.109), Equation (99.107) and Proposition 99.44.
Derives Corollary 99.50. \(\vect{E}\) and \(\vect{B}\) are built from \(F_{\mu\nu}\) and hence from derivatives of the free field, whose commutator is the \(c\)-number Equation (99.107) (with \(\kappa=0\)) differentiated. That \(c\)-number vanishes at spacelike separation by Proposition 99.44, and so do its derivatives, since a function vanishing on the whole spacelike region vanishes with all its derivatives there. The averages are integrals of the fields over \(R_{1}\) and \(R_{2}\), so their commutator is the double integral of a function that is zero throughout the domain of integration.
∎The commutator is nonzero when the regions are in light contact, which is when one measurement can influence the other, and this is what makes the operational reading of the formalism defensible in the sense discussed in Epistemology and the Scientific Method: an operator formalism whose commutators did not match the disturbances of a possible apparatus would be claiming a structure that no experiment could probe. Bohr and Rosenfeld's analysis is the check that the two agree. Note also what the analysis licenses: measurement by a classical test body of a quantized field, which is the hybrid description every real experiment uses.
Vacuum energy made observable: the Casimir effect
Section 99.2.4 argued that the zero-point energy Equation (99.46) is unobservable in flat empty space because it is an additive constant, but that differences between geometries are not. This section computes such a difference, gives its magnitude in SI units, and reports the measurements.
The prediction
Two parallel, plane, perfectly conducting plates of area \(A\) at separation \(a\ll\sqrt{A}\) in vacuum attract, with energy per unit area and force per unit area
independent of the charge of the electron and of every property of the plates except that they are perfect conductors. At \(a=1\,\mu\mathrm{m}\) the pressure is \(1.30\times 10^{-3}\,\mathrm{Pa}\), about \(10^{-8}\) of atmospheric. Rests on Equation (99.46), Equation (99.80) and Theorem 99.9.
Derivation. Derives Phenomenon 99.51. Between perfectly conducting plates the tangential electric field vanishes on both surfaces, so the allowed wave vectors normal to the plates are quantized, \(k_{z}=n\pi/a\) with \(n=0,1,2,\ldots\); the wave vector parallel to the plates, \(\vect{k}_{\parallel}\), is continuous. For \(n\ge1\) there are two polarizations, for \(n=0\) only one. The zero-point energy per unit area is therefore
where the prime means the \(n=0\) term carries weight \(\tfrac12\) (compensating its single polarization) and
the middle expression being the analytic continuation of the convergent integral \(\int\dd^{2}k_{\parallel}(2\pi)^{-2} (\vect{k}_{\parallel}^{2}+b^{2})^{-s} =b^{2-2s}/[4\pi(s-1)]\), valid for \(\Re s>1\), to \(s=-\tfrac12\).
The individual terms of Equation (99.111) are infinite and so is their sum; what is finite is the difference between the sum and the same quantity with the plates absent, that is with the \(k_{z}\) spectrum continuous. That difference is what Euler–Maclaurin computes. With \(F(n)=f(n\pi/a)\),
the omitted terms involving \(F^{(5)}\) and higher. The integral on the left is precisely the free-space zero-point energy of the enclosed volume, since \(\dd n=(a/\pi)\dd k_{z}\) converts it into \(\int\dd^{3}k\) times \(a\): it is proportional to the volume, and it is cancelled by the identical contribution from the unbounded space outside when the plates are moved. It therefore exerts no net force, which is the honest version of “discard the divergence”. By Equation (99.112), \(F(n)=-(\pi/a)^{3}n^{3}/6\pi\), so \(F(0)=0\), \(F'(0)=0\), \(F'''(0)=-\pi^{2}/a^{3}\), and all higher derivatives vanish identically because \(F\) is a cubic. Hence
and \(F/A=-\pp(E/A)/\pp a=-\pi^{2}\hbar c/(240a^{4})\), negative and therefore attractive. Numerically \(\pi^{2}\hbar c/240=1.3001\times 10^{-27}\,\mathrm{J}\,\mathrm{m}\), which at \(a=10^{-6}\,\mathrm{m}\) gives \(1.300\times 10^{-3}\,\mathrm{J}/\mathrm{m}^{3} =1.300\times 10^{-3}\,\mathrm{Pa}\).
The same answer follows from \(\zeta\)-function regularization in one line — Equation (99.111) becomes \(-(\hbar c\pi^{2}/6a^{3})\zeta(-3)\) with \(\zeta(-3)=1/120\) — but the Euler–Maclaurin route is preferable here because it exhibits the subtraction explicitly instead of hiding it in an analytic continuation.
∎The result contains no coupling constant: \(e\), \(\varepsilon_{0}\) and the properties of the metal have all disappeared, leaving only \(\hbar\), \(c\) and the geometry. That is the signature of a pure vacuum effect, and it is also the signature of the idealization: the plates entered only through the boundary condition “perfect conductor at all frequencies”, which no real metal is.
| Separation $a$ | $\abs{F/A}$ | Comparison |
|---|---|---|
| \(100\,\mathrm{nm}\) | \(13.0\,\mathrm{Pa}\) | the weight of \(0.13\,\mathrm{g}\) on \(1\,\mathrm{cm}^{2}\) |
| \(200\,\mathrm{nm}\) | \(0.813\,\mathrm{Pa}\) | |
| \(500\,\mathrm{nm}\) | \(2.08\times 10^{-2}\,\mathrm{Pa}\) | |
| \(1\,\mu\mathrm{m}\) | \(1.300\times 10^{-3}\,\mathrm{Pa}\) | $10^{-8}$ atmospheres |
| \(6\,\mu\mathrm{m}\) | \(1.00\times 10^{-6}\,\mathrm{Pa}\) |
Three corrections separate Equation (99.110) from what an experiment sees, and all three are larger than the experimental precision.
Temperature. Equation (99.111) counted zero-point energy only. At temperature \(T\) the thermally excited modes contribute too, and the scale that decides when they matter is the one at which \(k_{B}T\) becomes comparable with \(\hbar c/a\), the thermal wavelength
using \(k_{B}=1.380649\times 10^{-23}\,\mathrm{J}/\mathrm{K}\) exactly by the SI definition of the kelvin: the thermal correction sets in for \(a\gtrsim\lambda_{T}\). Note that this is not the same as demanding that \(k_{B}T\) reach the energy \(\pi\hbar c/a\) of the lowest confined mode, which would put the crossover at \(\pi\lambda_{T}=24\,\mu\mathrm{m}\); the modes that matter are the continuum of parallel wave vectors, not the lowest normal mode alone. Room-temperature experiments at separations of a few micrometres are therefore already approaching the crossover region and must model it.
Finite conductivity. A real metal reflects poorly at frequencies above its plasma frequency, so modes with \(k\gtrsim\omega_{p}/c\) are not confined at all. The relevant scale is the plasma wavelength, of order \(100\,\mathrm{nm}\) for gold and aluminium, and the correction reduces the force by tens of percent at separations near that scale — precisely the range where the force is largest and easiest to measure.
Surface roughness and patch potentials. Deviations from flatness of order \(\delta\) change the local separation and, because the force goes as \(a^{-4}\), do not average out; and spatial variation of the work function across a real surface produces electrostatic patch potentials whose force must be measured and subtracted. These are the dominant systematics of the modern experiments.
A closely related effect is the retarded van der Waals force between a neutral atom and a conducting wall, computed by Casimir and Polder [Casimir:1948a] in the same year and measured directly by Sukenik and collaborators [Sukenik:1993] by deflecting a sodium beam between gold plates. The existence of that derivation matters for interpretation: the same force between plates can be obtained by summing retarded intermolecular forces between the constituent atoms, a calculation in which no zero-point mode sum ever appears. The two routes agree, which means that the Casimir force confirms the relativistic quantized electromagnetic field, and does not by itself establish that the vacuum carries an absolute energy density. Section 99.2.4 is not overturned by Section 99.8; it is illustrated by it.
Measurements
Casimir's prediction stood for a decade before anyone tried to see it, and for four decades before anyone saw it convincingly. The difficulty is Table 99.3: forces of nanonewtons at separations of micrometres, between surfaces that must be parallel to a small fraction of that separation.
| Work | Method | Range | Outcome |
|---|---|---|---|
| Sparnaay 1958 [Sparnaay:1958] | parallel plates on a spring balance | \(0.5\text{–}2\,\mu\mathrm{m}\) | consistent with theory; uncertainty of order the effect |
| Lamoreaux 1997 [Lamoreaux:1997] | torsion pendulum, spherical lens against a flat plate | \(0.6\text{–}6\,\mu\mathrm{m}\) | agreement at the \(5\,\mathrm{\%}\) level |
| Mohideen and Roy 1998 [Mohideen:1998] | atomic force microscope, metallized sphere against a plate | \(0.1\text{–}0.9\,\mu\mathrm{m}\) | agreement at the \(1\,\mathrm{\%}\) level after conductivity and roughness corrections |
Sparnaay [Sparnaay:1958] used the geometry of the calculation: two flat plates, one on a spring balance, with the separation measured capacitively. He found an attraction of the right sign and order of magnitude, and stated the result honestly as not contradicting Casimir's prediction — the uncertainties, dominated by residual electrostatic charges and by the difficulty of keeping the plates parallel, were comparable to the effect itself. It was a null result in the useful sense: it excluded the absence of the force, and nothing more.
The modern experiments abandoned the parallel-plate geometry. Lamoreaux [Lamoreaux:1997] suspended a spherical lens near a flat plate from a torsion pendulum and measured the torque, sweeping the separation from \(0.6\,\mu\mathrm{m}\) to \(6\,\mu\mathrm{m}\); the electrostatic force at larger separations, where the Casimir force is negligible, calibrated the apparatus and measured the residual contact potential that had to be compensated. Mohideen and Roy [Mohideen:1998] replaced the pendulum with an atomic force microscope, gluing a metallized sphere to the cantilever, which pushed the accessible range down to \(0.1\,\mu\mathrm{m}\) where the force is four orders of magnitude larger.
For a sphere of radius \(R\) at closest separation \(a\ll R\) from a plane, the proximity-force approximation gives
Rests on Equation (99.110).
Derives Proposition 99.52. Divide the sphere into annuli at transverse radius \(\rho\), whose local separation from the plane is \(h(\rho)=a+\rho^{2}/2R+O(\rho^{4}/R^{3})\). Treating each annulus as a piece of parallel plate — legitimate when \(a\ll R\), since the surfaces are then locally parallel over the region that contributes — the energy is
using \(\rho\,\dd\rho=R\,\dd h\). Differentiating with respect to \(a\) gives \(F=-\pp E/\pp a=2\pi R\,E(a)/A\), and substituting Equation (99.110) gives Equation (99.115).
∎Two things should be said about what the measurements establish.
They establish that the mode structure of the quantized electromagnetic field is changed by boundaries in the way Equation (99.111) says, to the percent level over an order of magnitude in separation, with a functional form (\(a^{-3}\) for the sphere, \(a^{-4}\) for the pressure between plates) that no adjustable parameter was used to fit. That is a genuine confirmation of the quantized field, and of Section 99.2.4's claim that differences of zero-point energy are physical.
They do not establish the absolute normalization of the vacuum energy, for the reason given above: the same force follows from summing retarded van der Waals interactions [Casimir:1948a], in which no zero-point sum appears. Nor do they bear on the cosmological constant problem, whose difficulty is precisely that gravity couples to the absolute value that the Casimir experiment never sees. A text that offers the Casimir effect as evidence for the vacuum energy of Equation (99.49) has overstated its case by more than fifty orders of magnitude.
Limits of the canonical approach
Everything above is exact and everything above is free. The transition to interacting fields is made in the following chapters by perturbation theory, and it is worth ending with the reason that transition is less well founded than its success suggests.
Haag's theorem
The interaction picture, on which the whole of perturbation theory is built, supposes that the interacting field at a fixed time is related to a free field by a unitary transformation — so that the two share a Hilbert space, and the interacting dynamics can be written as a unitary evolution of free-field states. Haag [Haag:1955] proved that this cannot be.
Let \(\phi_{1}\) and \(\phi_{2}\) be two scalar fields at a fixed time, each irreducible, each covariant under spatial translations and rotations with a unique invariant vacuum, and suppose there is a unitary \(V\) with \(\phi_{2}=V\phi_{1}V^{-1}\) and \(\pi_{2}=V\pi_{1}V^{-1}\) at that time. If \(\phi_{1}\) is a free field of mass \(m\), then \(\phi_{2}\) is also free, of the same mass. Consequently the interaction picture exists only for a theory that is not interacting. Rests on Postulate 99.5 and Equation (99.23).
Haag's theorem: the proof runs through the equality of all equal-time vacuum expectation values under the unitary, the reconstruction of the theory from them, and the identification of the mass from the two-point function. It is a page or two of distribution-theoretic argument and belongs in the long-proof appendix; the statement and its consequences are used here, and the axiomatic setting in which the hypotheses are stated is the axiomatic quantum field theory chapter.
The mechanism is easy to state even without the proof. A unitary \(V\) preserves all vacuum expectation values, so the two theories have the same equal-time correlation functions; the Wightman reconstruction theorem [Wightman:1956] [Streater:1964] says that those functions determine the theory; and the two-point function of a free field determines the mass. There is therefore no room for \(\phi_{2}\) to be anything but a free field of the same mass. The sharpest form of the obstruction is that the interacting vacuum is not merely a different vector in the free theory's Fock space — it is not in that space at all, and the two representations of the canonical commutation relations Equation (99.23) are unitarily inequivalent. This is the warning already entered in Remark 99.13, and it is the starting point of the algebraic approach of [Haag:1992] and Axiomatic Quantum Field Theory.
Why, then, does perturbative quantum electrodynamics work? Three things must be said, and the third is uncomfortable.
First, every actual calculation is done in a regularized theory in which Haag's hypotheses fail. A momentum cutoff breaks Lorentz invariance; a finite volume breaks translation invariance; dimensional regularization changes the number of dimensions. The theorem does not apply to any of them, and the perturbative series is generated inside the regularization and only then continued.
Second, what perturbation theory computes is not a unitary transformation between Hilbert spaces but a formal power series in the coupling. That series is believed to be asymptotic rather than convergent — Dyson's argument that a theory with negative \(\alpha\) would be unstable indicates a zero radius of convergence [Dyson:1952] — so its first several terms can be an excellent approximation while the series as a whole defines nothing.
Third, and this is the uncomfortable part: the justification of the procedure is empirical. The anomalous magnetic moment of the electron, computed to five loops and measured to comparable precision, agrees at the level of parts in \(10^{12}\), and the comparison is the subject of Experiment: The Electron Anomalous Magnetic Moment. That agreement is the best evidence in physics that the perturbative expansion is computing something real. It is not a proof that the object it is expanding around exists. The honest position is that the formalism of this chapter is known to be inconsistent with the interaction picture built on it, that the inconsistency has never produced a wrong answer where the answer could be checked, and that constructing an interacting four-dimensional field theory that satisfies the axioms remains open (Axiomatic Quantum Field Theory).
What comes next
The free field is solved. Everything that makes the Standard Model a theory of anything — scattering, decay, binding, running couplings, confinement — lies in what happens when the Lagrangian acquires a term that is not quadratic. Four routes lead out, and each is a following chapter.
Perturbation theory and renormalization. Expand \(\exp(\ii S_{\text{int}}/\hbar)\), apply Theorem 99.41, and read off diagrams. The loop integrals diverge, and the divergences are absorbed into a redefinition of the mass, the charge and the field normalization; the result is the most accurately tested theory known, and the machinery is Quantum Electrodynamics and Renormalization. The dependence of the resulting couplings on the energy scale — itself a consequence of the renormalization procedure rather than an extra assumption — is The Renormalization Group.
The path integral. The canonical method fixes a time slice and therefore hides Lorentz invariance, and for non-abelian gauge fields its constraint structure becomes intractable. The alternative quantization, in which the amplitude is a sum over field histories weighted by \(\exp(\ii S/\hbar)\), keeps covariance manifest and handles gauge fixing systematically; it is Path-Integral Quantization, and it is the only practical route to the theory of Quantum Chromodynamics. The \(\hbar\) in the phase is the same \(\hbar\) that appears in Equation (99.23): the two quantizations are equivalent for the systems where both apply.
Scattering theory. Field operators are not what a detector records. Converting the one into the other — asymptotic states, the \(S\) matrix, cross-sections and decay rates in SI units — is Scattering Theory, and the reduction formula that expresses an \(S\)-matrix element in terms of the time-ordered products of this chapter is stated in Axiomatic Quantum Field Theory.
Axiomatics. Finally, the results of Sections 99.7.1, 99.7.2 and 99.9.1 are statements that do not depend on any Lagrangian, and they are better stated in a framework that does not begin with one. That framework is Axiomatic Quantum Field Theory, and its principal use in this book is to say precisely which of the results above are theorems and which are properties of a particular model.