Quantum Chromodynamics

Contents
  1. Hadron spectroscopy and the quark hypothesis
  2. Colour
  3. The QCD Lagrangian
  4. Asymptotic freedom
  5. Deep inelastic scattering and the parton model
  6. Jets and the gluon
  7. Confinement
  8. Chiral symmetry and its breaking
  9. QCD under extreme conditions
  10. The strong CP problem

Quantum chromodynamics is the gauge theory of colour, and it is the one part of the Standard Model whose defining property — that its elementary fields have never been seen in isolation — makes the evidence for it indirect at every step. This chapter therefore follows the evidence rather than the formalism: the regularities of the hadron spectrum that produced the quark hypothesis, the spin-statistics paradox of the \(\Delta^{++}\) that forced a threefold colour degree of freedom, the deep inelastic scattering that showed pointlike constituents inside the proton, the three-jet events that made the gluon visible, and the lattice computations that now reproduce the hadron masses from the Lagrangian alone. Only then is the failure to observe a free quark discussed for what it is — a strong empirical regularity supported by a well-tested numerical demonstration, not a theorem.

QCD is the second gauge theory of the part and reuses the entire apparatus of Quantum Electrodynamics and Renormalization, but two features have no QED analogue and organize everything below. The gluons carry the charge they mediate, so the beta function changes sign: the coupling is weak at short distance — asymptotic freedom, which is why perturbation theory works at all here — and strong at hadronic scales, where it does not, and Path-Integral Quantization together with lattice methods must take over. And the near-masslessness of the lightest quarks gives an approximate chiral symmetry whose spontaneous breaking, not the Higgs mechanism of Electroweak Unification and the Higgs Boson, supplies about 99 percent of the mass of ordinary matter. Its experimental anchor is Experiment: Deep Inelastic Scattering; its downstream users are Nuclear Forces and Nuclear Structure, Stellar Structure and Nucleosynthesis and Compact Stars and Relativistic Astrophysics. Standard treatments are [Weinberg:1996] [Halzen:1984]; data throughout are from [Navas:2024].

Units and conventions are those fixed once for the whole part in Section 100.1.1: the metric \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\) of Notation 100.1, the mass shell \(p^{2}=m^{2}c^{2}\) of Equation (100.1), the action Equation (100.2) carrying an explicit \(c^{-1}\), and \(\hbar\) and \(c\) written out in every formula. Nothing below is computed in units where they are set to one; Remark 100.4 gives the dictionary for reading this chapter against the literature, which does.

Hadron spectroscopy and the quark hypothesis

The eightfold way

Between 1947 and 1960 the list of strongly interacting particles grew from a handful to several dozen. The cosmic-ray and accelerator experiments of the period produced the pion, the kaons, the \(\Lambda\), the \(\Sigma\) and \(\Xi\) hyperons and, once the bubble chamber allowed short-lived states to be reconstructed from their decay products, a succession of resonances — the \(\Delta(1232)\), the \(\Sigma(1385)\), the \(\Xi(1530)\), the \(\rho\), the \(K^{*}\). None of them could be called elementary with a straight face. The response was not to propose constituents but to look for a symmetry, and the symmetry that worked is the one whose group theory is established, as mathematics and with no physics attached, in Section 14.2.2.

Two additive quantum numbers organize the data. The first is isospin: the near-equality of the proton and neutron masses, and of the three pion masses, is described by an \(\SU(2)\) under which \((p,n)\) is a doublet, and its representation theory is that of Linear Algebra and Representation Theory. The second is strangeness \(S\), introduced by Gell-Mann [GellMann:1953b] and independently by Nakano and Nishijima [Nakano:1953] to account for the fact that the new particles are produced copiously (strong interaction) but decay slowly (weak interaction): \(S\) is conserved by the strong interaction and violated by the weak one, so a strange particle can only be produced in pairs and must then wait for a weak decay. Writing the hypercharge \(Y=B+S\) with \(B\) the baryon number, the observed electric charges of every hadron obey

\begin{equation}\tag{102.1} \frac{Q}{e}=I_{3}+\frac{Y}{2}\ec \end{equation}

with \(I_{3}\) the third isospin component. Equation (102.1) is a fact about hundreds of measured charges, not a definition.

The content of the eightfold way, proposed independently by Gell-Mann [GellMann:1961] and Ne'eman [Neeman:1961], is that \(I_{3}\) and \(Y\) are the two commuting quantum numbers of a rank-two group, and that the group is \(\SU(3)\). Proposition 14.50 gives the rank as \(\ell=2\) and the dimension as \(8\); the two Cartan generators of Definition 14.51 are \(T_{3}\) and \(T_{8}\), and the identification with the measured quantum numbers is

\begin{equation}\tag{102.2} I_{3}=T_{3}\ec\qquad Y=\frac{2}{\sqrt{3}}\,T_{8}\ec \end{equation}

the factor \(2/\sqrt3\) being fixed once and for all by Equation (14.65). A hadron multiplet is then an irreducible representation, and the weight diagram of that representation — the set of simultaneous eigenvalues of \((T_{3},T_{8})\), computed as pure mathematics in Equation (14.79) and Equation (14.80) — is a prediction of which charges and strangenesses occur together and how many states there are.

The prediction is sharp and it is met. The eight lightest spin-\(\tfrac{1}{2}\) baryons (\(p\), \(n\), \(\Lambda\), \(\Sigma^{\pm,0}\), \(\Xi^{0,-}\)) fill the weight diagram of the adjoint representation \(\vect{8}\) exactly: six states on a hexagon and two at the centre, which is the count \(6+2=8\) of Remark 14.56. The eight lightest pseudoscalar mesons (\(\pi^{\pm,0}\), \(K^{\pm}\), \(K^{0}\), \(\bar{K}^{0}\), \(\eta\)) fill a second copy of the same diagram. The nine lightest spin-\(\tfrac{3}{2}\) baryons known in 1961 fill nine of the ten places of a decuplet, and the tenth place is the subject of Section 102.1.2.

Phenomenon 102.1 (Hadron masses obey an octet mass formula).

Within a multiplet the masses are not equal — \(\SU(3)\) is broken by the strange quark mass — but the breaking is not arbitrary. For the \(J^{P}=\tfrac{1}{2}^{+}\) baryon octet the measured masses satisfy the Gell-Mann–Okubo relation [GellMann:1961] [Okubo:1962]

\begin{equation}\tag{102.3} 2\left(M_{N}+M_{\Xi}\right)=3M_{\Lambda}+M_{\Sigma} \end{equation}

to better than one percent, and for the \(J^{P}=\tfrac{3}{2}^{+}\) decuplet the masses are equally spaced in hypercharge to within a few percent. Rests on Theorem 5.158, Equation (14.70) and Equation (14.69).

Derivation. Derives Phenomenon 102.1. Write the strong Hamiltonian as \(H=H_{0}+H_{1}\), with \(H_{0}\) \(\SU(3)\)-symmetric and \(H_{1}\) the symmetry-breaking piece. The breaking comes from the quark mass term of Section 102.3.3, which for \(m_{u}=m_{d}\neq m_{s}\) splits into a piece proportional to the identity — absorbed into \(H_{0}\) — and a piece proportional to \(\lambda_{8}\) of Equation (14.65). So \(H_{1}\) transforms as the eighth component of an octet. Treat it in first-order perturbation theory (Approximation Methods), so that the mass shift of a state \(\ket{B}\) is \(\ev{H_{1}}{B}\).

By the Wigner–Eckart theorem for \(\SU(3)\) — the same statement as for \(\SU(2)\), and proved in the same way from Schur's lemma Theorem 5.158 — the diagonal matrix elements of an octet operator inside an octet are determined by two reduced matrix elements, because \(\vect{8}\otimes\vect{8}\) contains \(\vect{8}\) twice (once symmetrically, through \(d_{abc}\) of Equation (14.70), and once antisymmetrically, through \(f_{abc}\) of Equation (14.69)). Two reduced matrix elements mean two constants, and the only \(\SU(3)\) scalars that can be built from \(T_{8}\) and the state's own quantum numbers, to first order, are \(Y\) and the isospin Casimir \(I(I+1)\) corrected by the trace term \(Y^{2}/4\). Hence

\begin{equation}\tag{102.4} M=M_{0}+aY+b\left[I(I+1)-\frac{Y^{2}}{4}\right]\ec \end{equation}

with \(M_{0}\), \(a\) and \(b\) the same for the whole multiplet.

For the octet, \((Y,I)\) is \((1,\tfrac{1}{2})\) for \(N\), \((0,0)\) for \(\Lambda\), \((0,1)\) for \(\Sigma\) and \((-1,\tfrac{1}{2})\) for \(\Xi\), so that \(I(I+1)-Y^{2}/4\) equals \(\tfrac{1}{2}\), \(0\), \(2\), \(\tfrac{1}{2}\) respectively, and

\begin{align} M_{N}&=M_{0}+a+\tfrac{1}{2}b\ec & M_{\Lambda}&=M_{0}\ec\nn\\ M_{\Sigma}&=M_{0}+2b\ec & M_{\Xi}&=M_{0}-a+\tfrac{1}{2}b\ep \tag{102.5} \end{align}

The three unknowns are eliminated from the four masses in one way: \(M_{N}+M_{\Xi}=2M_{0}+b\) and \(3M_{\Lambda}+M_{\Sigma}=4M_{0}+2b\), whence Equation (102.3). Numerically, with the isospin averages \(M_{N}c^{2}=938.92\,\mathrm{MeV}\), \(M_{\Lambda}c^{2}=1115.68\,\mathrm{MeV}\), \(M_{\Sigma}c^{2}=1193.15\,\mathrm{MeV}\) and \(M_{\Xi}c^{2}=1318.29\,\mathrm{MeV}\) [Navas:2024],

\begin{equation}\tag{102.6} 2\left(M_{N}+M_{\Xi}\right)c^{2}=4514.4\,\mathrm{MeV}\ec\qquad \left(3M_{\Lambda}+M_{\Sigma}\right)c^{2}=4540.2\,\mathrm{MeV}\ec \end{equation}

a discrepancy of \(25.8\,\mathrm{MeV}\), that is \(0.57\,\mathrm{\%}\) of either side, against a spread of nearly \(400\,\mathrm{MeV}\) among the masses being related. For the decuplet the states satisfy \(I=1+Y/2\), so that \(I(I+1)-Y^{2}/4=2+\tfrac{3}{2}Y\) and Equation (102.4) collapses to a linear function of \(Y\): consecutive isospin multiplets are equally spaced in mass. That is what Section 102.1.2 exploits.

Remark 102.2 (What the eightfold way did and did not assert).

Neither [GellMann:1961] nor [Neeman:1961] proposed constituents. The claim was that an approximate \(\SU(3)\) organizes the spectrum, and the evidence for it was the multiplet structure and Equation (102.3). That the observed multiplets are \(\vect{1}\), \(\vect{8}\) and \(\vect{10}\) — and never the \(\vect{3}\) that generates them — was at the time an unexplained selection rule. It is explained in Section 102.2.1, and its explanation is colour.

The $\Omega^{-}$ prediction and its confirmation

By 1962 the decuplet contained the four \(\Delta\) states (\(Y=1\), \(I=\tfrac{3}{2}\)), the three \(\Sigma(1385)\) states (\(Y=0\), \(I=1\)) and the two \(\Xi(1530)\) states (\(Y=-1\), \(I=\tfrac{1}{2}\)): nine of ten. The tenth place, \(Y=-2\), \(I=0\), was empty. What Equation (102.4) supplies is not merely the existence of a tenth state but every one of its quantum numbers and, through the equal-spacing rule, its mass.

Phenomenon 102.3 (A particle predicted in full and then found).

The equal spacing of the decuplet requires a tenth baryon with strangeness \(S=-3\), isospin \(I=0\), charge \(-e\), spin-parity \(\tfrac{3}{2}^{+}\) and a mass about \(145\,\mathrm{MeV}\) above the \(\Xi(1530)\), that is close to \(1680\,\mathrm{MeV}\)\(/c^{2}\); being below the threshold for any strangeness-conserving strong decay it must decay weakly, and therefore live long enough to leave a macroscopic track. A single bubble-chamber event at Brookhaven displayed exactly such a particle, with mass \(1686(12)\,\mathrm{MeV}/c^{2}\) [Barnes:1964]. The current world average is \(1672.45(29)\,\mathrm{MeV}/c^{2}\) [Navas:2024], within \(0.5\,\mathrm{\%}\) of the prediction. Rests on Equations (102.1) and (102.4).

Derivation. Derives Phenomenon 102.3. Apply Equation (102.4) to the decuplet, where it reduces to \(M=\text{const}+\left(a+\tfrac{3}{2}b\right)Y\) as shown above. The measured spacings between consecutive isospin multiplets of the decuplet are, with modern isospin-averaged masses [Navas:2024],

\begin{equation}\tag{102.7} M_{\Sigma^{*}}-M_{\Delta}=153\,\mathrm{MeV}/c^{2}\ec\quad M_{\Xi^{*}}-M_{\Sigma^{*}}=149\,\mathrm{MeV}/c^{2}\ec\quad M_{\Omega}-M_{\Xi^{*}}=139\,\mathrm{MeV}/c^{2}\ec \end{equation}

equal to within \(5\,\mathrm{\%}\), which is the accuracy first-order perturbation theory in a breaking of this size deserves. With the masses available in 1962 the first two spacings were about \(145\,\mathrm{MeV}\)\(/c^{2}\) each, and extrapolating one further step placed the missing state near \(1680\,\mathrm{MeV}\)\(/c^{2}\).

The quantum numbers follow with no freedom at all. The tenth weight of the decuplet sits at \(Y=-2\), \(I=I_{3}=0\), so Equation (102.1) gives \(Q=-e\), and \(Y=B+S\) with \(B=1\) gives \(S=-3\). The lifetime argument is kinematic: the lightest \(S=-3\) combination reachable by a strong decay, which cannot change strangeness, is \(\Xi K\), whose threshold is \(\left(1315+494\right)\mathrm{MeV}/c^{2} =1809\,\mathrm{MeV}/c^{2}\) [Navas:2024] — above the predicted mass. Every strong channel is therefore closed, the state must decay through the weak interaction of Weak Interactions with \(\Delta S=1\), and its lifetime must be of the order of \(10^{-10}\,\mathrm{s}\) rather than the \(10^{-23}\,\mathrm{s}\) of its nine partners. That is the difference between a resonance visible only as a bump in a mass distribution and a particle that travels centimetres and is photographed.

Remark 102.4 (Why this counts for more than a fit).

The eightfold way was fitted to the nine known decuplet members; the tenth was not in the data. What was published in advance was a mass, a charge, a strangeness, a spin, a decay mode and an order of magnitude for a lifetime, and all six were confirmed by one event [Barnes:1964]. A scheme that reproduces \(n\) known numbers with \(n\) parameters says nothing; a scheme that predicts an unobserved object's entire specification and is then vindicated has been exposed to refutation and has survived. This is the asymmetry between accommodation and prediction argued in Epistemology and the Scientific Method, and the \(\Omega^{-}\) is the cleanest instance of it in hadron physics.

Quarks

The representations that occur in Nature are \(\vect{1}\), \(\vect{8}\) and \(\vect{10}\). None of them is the fundamental \(\vect{3}\), but all of them are built from it: this is the observation of Gell-Mann [GellMann:1964] and, in a more concretely constituent reading, of Zweig [Zweig:1964], whose “aces” are the same objects.

Definition 102.5 (Quarks).

A quark is a spin-\(\tfrac{1}{2}\) fermion of baryon number \(B=\tfrac{1}{3}\) transforming in the fundamental representation \(\vect{3}\) of flavour \(\SU(3)\). Reading the weights Equation (14.79) through Equation (102.2) assigns to the three states \((u,d,s)\) the quantum numbers

\begin{equation}\tag{102.8} \begin{aligned} u:&\quad I_{3}=+\tfrac{1}{2}\ec & Y&=+\tfrac{1}{3}\ec & Q&=+\tfrac{2}{3}e\ec\\ d:&\quad I_{3}=-\tfrac{1}{2}\ec & Y&=+\tfrac{1}{3}\ec & Q&=-\tfrac{1}{3}e\ec\\ s:&\quad I_{3}=0\ec & Y&=-\tfrac{2}{3}\ec & Q&=-\tfrac{1}{3}e\ec \end{aligned} \end{equation}

the charges following from Equation (102.1) and being, for the first time in physics, fractions of the elementary charge. Antiquarks carry the conjugate representation \(\bar{\vect{3}}\) and the opposite signs.

Proposition 102.6 (Mesons are $q\bar{q}$, baryons are $qqq$).

The multiplets observed in Section 102.1.1 are exactly those obtained by combining quarks as

\begin{equation}\tag{102.9} \vect{3}\otimes\bar{\vect{3}}=\vect{1}\oplus\vect{8}\ec\qquad \vect{3}\otimes\vect{3}\otimes\vect{3} =\vect{1}\oplus\vect{8}\oplus\vect{8}\oplus\vect{10}\ep \end{equation}

Rests on Definition 102.5, Theorem 14.57 and Equation (14.81).

Proof.

Derives Proposition 102.6. The first decomposition is Theorem 14.57, proved there both by splitting \(\operatorname{End}(\C^{3})\) into its trace and traceless parts and by counting weights. For the second, tensor Equation (14.81) once more with \(\vect{3}\), or count directly: \(\C^{3}\otimes\C^{3}\otimes\C^{3}\) has dimension \(27\) and decomposes under the symmetric group \(S_{3}\), which commutes with the \(\SU(3)\) action, into a totally symmetric part of dimension \(\binom{5}{3}=10\), a totally antisymmetric part of dimension \(\binom{3}{3}=1\), and two copies of a mixed-symmetry part of dimension \(8\) each; \(10+1+8+8=27\). Each symmetry class is \(\SU(3)\)-invariant, and the four pieces are the representations named. Baryon number is additive with \(B=\tfrac{1}{3}\) per quark, so \(q\bar{q}\) has \(B=0\) and \(qqq\) has \(B=1\), matching mesons and baryons respectively, and Equation (102.1) is reproduced because \(Q\) and \(Y\) are additive by construction.

The decuplet is the totally symmetric combination, which is why it contains \(uuu\), \(ddd\) and \(sss\) — the \(\Delta^{++}\), the \(\Delta^{-}\) and the \(\Omega^{-}\). That fact is about to cause trouble.

Phenomenon 102.7 (The nucleon magnetic moments come out right).

Treating the nucleon as three quarks in a spatially symmetric ground state, with each quark contributing a Dirac moment inversely proportional to its mass, predicts

\begin{equation}\tag{102.10} \frac{\mu_{p}}{\mu_{n}}=-\frac{3}{2}\ec \end{equation}

with no free parameter whatever. The measured ratio is \(2.792847/(-1.913043)=-1.45990\) [Mohr:2025], low by \(2.7\,\mathrm{\%}\). Rests on Definition 102.5, Proposition 102.6 and Equation (93.2).

Derivation. Derives Phenomenon 102.7. A Dirac particle of charge \(Q_{q}e\) and mass \(m_{q}\), with \(Q_{q}\) the dimensionless charge in units of the elementary charge, has magnetic moment \(\mu_{q}=Q_{q}e\hbar/(2m_{q})\), of SI dimension \(\mathrm{J}/\mathrm{T}\); this is the \(g=2\) result of the Dirac equation (The Dirac Equation), whose small radiative correction is the subject of Section 100.8.2 and is irrelevant at the accuracy in question here. Write \(\hat{\mu}:=e\hbar/(2m_{q})\) for the common scale, with \(m_{u}=m_{d}=:m_{q}\) by isospin symmetry, so that \(\mu_{u}=+\tfrac{2}{3}\hat{\mu}\) and \(\mu_{d}=-\tfrac{1}{3}\hat{\mu}\).

The proton is \(uud\) with total spin \(\tfrac{1}{2}\) and no orbital angular momentum. The spin-flavour wavefunction that is symmetric under exchange of all three quarks — required by Section 102.2.1, and there is exactly one such state with \(J=\tfrac{1}{2}\) — gives, on projecting the magnetic moment operator \(\sum_{i}\mu_{i}\sigma_{z}^{(i)}\) onto it,

\begin{equation}\tag{102.11} \mu_{p}=\frac{4\mu_{u}-\mu_{d}}{3}\ec\qquad \mu_{n}=\frac{4\mu_{d}-\mu_{u}}{3}\ep \end{equation}

The combinatorial factors \(4\) and \(-1\) are the statement that in the symmetric \(J=\tfrac{1}{2}\) state the like pair (\(uu\) in the proton) is coupled to spin \(1\) with weight \(\tfrac{2}{3}\) and the odd quark carries the remaining spin with weight \(-\tfrac{1}{3}\). Substituting, \(\mu_{p}=\hat{\mu}\) and \(\mu_{n}=-\tfrac{2}{3}\hat{\mu}\), whence Equation (102.10). The unknown \(m_{q}\) has cancelled, which is why the ratio is a prediction and the individual moments are not.

Feeding the measured \(\mu_{p}=2.792847\,\mu_{N}\) [Mohr:2025], with the nuclear magneton \(\mu_{N}=5.0507837393\times 10^{-27}\,\mathrm{J}/\mathrm{T}\) [Mohr:2025], back into \(\mu_{p}=e\hbar/(2m_{q})\) gives

\begin{equation}\tag{102.12} m_{q}=\frac{m_{p}}{2.792847}=5.99\times 10^{-28}\,\mathrm{kg} =336\,\mathrm{MeV}/c^{2}\ec \end{equation}

almost exactly one third of the nucleon mass. A model in which the nucleon is three objects of mass \(m_{N}/3\) therefore reproduces both the moment ratio and the moment scale. That the number \(336\,\mathrm{MeV}\)\(/c^{2}\) has nothing to do with the \(u\) and \(d\) masses in the Lagrangian, which are \(2.16\,\mathrm{MeV}\)\(/c^{2}\) and \(4.70\,\mathrm{MeV}\)\(/c^{2}\) [Navas:2024], is the first appearance in this chapter of the fact developed in Section 102.8.2: the mass of ordinary matter is not the mass of its quarks.

Remark 102.8 (The status of the model in 1964).

Gell-Mann's paper is explicit that the quarks might be a mathematical device: fractional charge had never been seen, and searches for it were already under way. Zweig, writing at CERN in the same year, took the constituents literally and could not get his paper published in a journal [Zweig:1964]. The successes above — the multiplet structure, Equation (102.3), Equation (102.10) — are all statements about symmetry, and a symmetry can be realized without its generating representation being populated by particles. What converted the scheme into a theory of constituents was the deep inelastic scattering of Section 102.5, which saw the objects directly. The searches for free fractional charge, meanwhile, have never found anything: Section 102.7.3.

Colour

The statistics paradox

The quark model of Section 102.1.3 works, and it contains a contradiction with a theorem.

Phenomenon 102.9 (The $\Delta^{++}$ violates the spin-statistics theorem).

The \(\Delta^{++}(1232)\) is the lightest baryon of charge \(+2e\), with \(J^{P}=\tfrac{3}{2}^{+}\), mass \(1232.0(20)\,\mathrm{MeV}/c^{2}\) and width \(117(3)\,\mathrm{MeV}\) [Navas:2024]. In the quark model its content is \(uuu\): three identical fermions. Being the lightest state of its quantum numbers it has no orbital excitation, so its spatial wavefunction is symmetric; being \(J=\tfrac{3}{2}\) with three spin-\(\tfrac{1}{2}\) constituents, its spin wavefunction in the \(J_{z}=+\tfrac{3}{2}\) state is \(\uparrow\uparrow\uparrow\), also symmetric; and its flavour wavefunction \(uuu\) is symmetric. The total wavefunction is therefore symmetric under exchange of any two of three identical fermions, in direct contradiction with the statistics asserted by the spin-statistics theorem, whose proof belongs to the axiomatic chapter. Rests on Definition 102.5 and Proposition 102.6.

The contradiction is not a technicality that a more careful model might avoid. It is repeated by the \(\Delta^{-}\) (\(ddd\)) and the \(\Omega^{-}\) (\(sss\)), and the same symmetric combination is what produced the successful magnetic moments Equation (102.11): the model's failure and its success come from the same wavefunction. Either the spin-statistics theorem fails — and it is a theorem, whose proof belongs to the axiomatic chapter — or a quark carries a quantum number that has not yet been written down.

Greenberg proposed the first escape, replacing Fermi statistics for quarks by parastatistics of order three, in which up to three identical quarks may occupy the same state [Greenberg:1964]. Han and Nambu proposed the second, and the one that survived: quarks come in three varieties distinguished by a new exactly conserved quantum number, on which acts a further, exact \(\SU(3)\) that has nothing to do with the flavour \(\SU(3)\) of Section 102.1.1 [Han:1965]. The two proposals are related — parastatistics of order three is what an antisymmetrized threefold degeneracy looks like if one refuses to name the degeneracy — but only the second can be gauged, and gauging it is QCD.

Definition 102.10 (Colour).

Each quark flavour \(f\) carries an additional index \(i=1,2,3\) (“red, green, blue”), transforming in the fundamental representation \(\vect{3}\) of a group \(\SU(3)_{c}\), distinct from and commuting with flavour \(\SU(3)\) and with the Poincare group. Antiquarks transform in \(\bar{\vect{3}}\). The generators are the \(T_{a}=\tfrac{1}{2}\lambda_{a}\) of Definition 14.51, with the structure constants Equation (14.69) and the normalization \(\tr(T_{a}T_{b})=\tfrac{1}{2}\delta_{ab}\) of Equation (14.66). Colour is dimensionless and is carried by no lepton.

With a colour index available, the \(\Delta^{++}\) wavefunction is

\begin{equation}\tag{102.13} \ket{\Delta^{++},J_{z}=\tfrac{3}{2}} =\frac{1}{\sqrt{6}}\,\epsilon_{ijk}\, \ket{u^{i}\!\uparrow}\ket{u^{j}\!\uparrow}\ket{u^{k}\!\uparrow}\ec \end{equation}

totally antisymmetric because \(\epsilon_{ijk}\) is, and the theorem is satisfied. Nothing has been bought cheaply: the same index must now explain why the states it permits are not seen.

Postulate 102.11 (The colour-singlet rule).

Every observable asymptotic state is a singlet of \(\SU(3)_{c}\).

Theorem 102.12 (Which quark combinations can be colour singlets).

A state built from \(m\) quarks and \(n\) antiquarks contains a colour singlet only if \(m-n\) is divisible by three. In particular \(q\bar{q}\) and \(qqq\) each contain exactly one singlet, while \(qq\), \(qqq\bar{q}\) and any single quark contain none. Rests on Definition 102.10, Theorem 5.158 and Proposition 14.55.

Proof.

Derives Theorem 102.12. The centre of \(\SU(3)\) is the group of the three matrices \(\omega^{r}\identity\), \(r=0,1,2\), with \(\omega=\ee^{2\pi\ii/3}\); these are unitary, unimodular (\(\det(\omega\identity)=\omega^{3}=1\)) and commute with everything, and no other multiple of the identity is unimodular. On the fundamental \(\vect{3}\) the element \(\omega\identity\) acts as multiplication by \(\omega\), and on \(\bar{\vect{3}}\) as multiplication by \(\omega^{-1}\). On a tensor product of \(m\) copies of \(\vect{3}\) and \(n\) of \(\bar{\vect{3}}\) it therefore acts as \(\omega^{m-n}\), and on the trivial representation it must act as \(1\). Hence \(\omega^{m-n}=1\), i.e. \(3\mid m-n\). This number, \(m-n \bmod 3\), is the triality, and it is a representation-theoretic invariant: no equivalence can change it, so the obstruction is absolute and not a matter of which coupling scheme is used.

For the multiplicities: the number of singlets in \(V\otimes W\) is \(\dim\operatorname{Hom}_{G}(V^{*},W)\), which by Schur's lemma Theorem 5.158 is \(1\) if \(W\) is equivalent to \(V^{*}\) and \(0\) otherwise. Since \(\bar{\vect{3}}\) is inequivalent to \(\vect{3}\) (Proposition 14.55), \(\vect{3}\otimes\vect{3}\) has no singlet, while \(\vect{3}\otimes\bar{\vect{3}}\) has exactly one — as Theorem 14.57 exhibits explicitly, the trace part of a \(3\times3\) matrix. For three quarks, the decomposition Equation (102.9) contains \(\vect{1}\) once, and the invariant realizing it is \(\epsilon_{ijk}\): the unique (up to scale) totally antisymmetric invariant tensor of \(\SL(3,\C)\), hence of \(\SU(3)\).

Corollary 102.13 (The selection rule of the eightfold way).

Combining Postulate 102.11 with Theorem 102.12 and Equation (102.9), the lightest observable hadrons are exactly the flavour multiplets \(\vect{1}\oplus\vect{8}\) from \(q\bar{q}\) and \(\vect{1}\oplus\vect{8}\oplus\vect{8}\oplus\vect{10}\) from \(qqq\): mesons and baryons, and nothing of triality one or two. This is the rule left unexplained in Remark 102.2. Rests on Postulate 102.11, Theorem 102.12 and Equation (102.9).

Remark 102.14 (The rule is a postulate here and a consequence later).

Postulate 102.11 is stated as a postulate because that is its honest status in this chapter: it is inferred from the observed spectrum and it is not derived from the Lagrangian of Section 102.3.3. The linear potential of Section 102.7.2 makes it plausible — separating a non-singlet pair costs an energy growing without bound — and the lattice computations of Section 102.7.1 exhibit it, but a proof from the continuum theory does not exist. Section 102.7 says so at length rather than disguising the gap.

Counting the colours

Colour was introduced to fix a statistics problem, and a quantum number introduced to fix one problem is worth little until it is measured somewhere else. It is measured in three independent places, all of which count three. The counts are independent in the strong sense: they involve different initial states, different interactions and different theoretical inputs, and each would fail by a large factor if \(N_{c}\) were \(1\).

Phenomenon 102.15 (Three colours, counted in an annihilation ratio).

The ratio

\begin{equation}\tag{102.14} R:=\frac{\sigma(e^{+}e^{-}\to\text{hadrons})} {\sigma(e^{+}e^{-}\to\mu^{+}\mu^{-})} \end{equation}

is, away from resonances, nearly independent of energy over wide intervals and steps upward at each new quark threshold. The measured plateaux lie close to \(2\) below the charm threshold and close to \(10/3\) between the charm and bottom thresholds [Navas:2024]. Both values require three colour states per quark flavour: with one colour they would be \(2/3\) and \(10/9\). Rests on Definition 102.10, Equation (100.32) and Theorem 100.58.

Derivation. Derives Phenomenon 102.15. At leading order both processes proceed through a single virtual photon, so they differ only in the charge of the produced pair and in how many distinct pairs can be produced. The denominator is Equation (100.32), \(\sigma(e^{+}e^{-}\to\mu^{+}\mu^{-}) =4\pi\alpha^{2}(\hbar c)^{2}/3E_{\mathrm{cm}}^{2}\), computed in Section 100.3.1. For a pointlike spin-\(\tfrac{1}{2}\) particle of electric charge \(Q_{f}e\) — here and throughout, \(Q_{f}\) is the dimensionless charge of flavour \(f\) in units of \(e\) — created well above its threshold, the same computation applies with the single change \(e^{2}\to Q_{f}e^{2}\) at the production vertex, so the cross-section is that of Equation (100.32) times \(Q_{f}^{2}\); and the rates for distinguishable final states add incoherently, so a quark flavour existing in \(N_{c}\) colours contributes \(N_{c}\) times, the electromagnetic vertex being blind to the colour label. Dividing by the muon channel, for which \(Q=-1\) and there is one internal state,

\begin{equation}\tag{102.15} R=N_{c}\sum_{q}Q_{q}^{2}\ec \end{equation}

the sum running over the flavours light enough to be produced. With \(Q_{u}=Q_{c}=+\tfrac{2}{3}\) and \(Q_{d}=Q_{s}=Q_{b}=-\tfrac{1}{3}\) this gives \(R=\tfrac{2}{3}N_{c}\) for \(u,d,s\), \(R=\tfrac{10}{9}N_{c}\) once charm opens and \(R=\tfrac{11}{9}N_{c}\) once bottom opens — that is \(2\), \(10/3\) and \(11/3\) for \(N_{c}=3\), which is what is measured.

Confinement does not spoil the count. The quarks produced at the photon vertex cannot emerge as free particles, but whatever they do afterwards they end as hadrons with unit probability, so the hadronic total is settled at the moment of pair creation and the long-distance physics cancels in the ratio. This is the same inclusiveness argument that makes Theorem 100.58 work, applied to colour instead of to soft photons.

The residual excess of the data over the integers above is calculable. Gluon emission from the produced pair multiplies Equation (102.15) by \(1+\alpha_{s}(E_{\mathrm{cm}})/\pi+\dots\), the first term being the same \(\mathcal{O}(\alpha_{s})\) real-plus-virtual sum that gives the three-jet rate of Section 102.6.2; at \(E_{\mathrm{cm}}=5\,\mathrm{GeV}\), where \(\alpha_{s}\approx0.2\), this is a \(6\,\mathrm{\%}\) excess, and it is seen. So \(R\) is simultaneously a colour counter at the \(10\,\mathrm{\%}\) level and, once the count is granted, one of the determinations of \(\alpha_{s}\) collected in Section 102.4.2.

The second count is sharper by a factor of ten in precision, and it is carried out in full in Section 106.7.2 rather than repeated here.

Proposition 102.16 (The neutral pion width counts colours squared).

The two-photon width of the \(\pi^{0}\) is fixed by the axial anomaly (Theorem 106.87) to be Equation (106.103),

\begin{equation}\tag{102.16} \Gamma(\pi^{0}\to\gamma\gamma) =\left(\frac{N_{c}}{3}\right)^{2} \frac{\alpha^{2}\left(m_{\pi}c^{2}\right)^{3}}{64\pi^{3}F_{\pi}^{2}}\ec \end{equation}

proportional to \(N_{c}^{2}\). For \(N_{c}=3\) the prediction is \(7.79(14)\,\mathrm{eV}\) [Navas:2024] [Mohr:2025] against the measured \(7.802(117)\,\mathrm{eV}\) [Larin:2020]: agreement to \(0.2\,\mathrm{\%}\). For \(N_{c}=1\) it would be nine times smaller. Rests on Theorem 106.87, Equation (106.103) and Theorem 106.89.

Proof.

Derives Proposition 102.16. The whole content is the colour factor of the triangle graph. The axial current whose divergence carries the anomaly is the third isospin component, \(j^{\mu5}_{3} =\tfrac{1}{2}(\bar{u}\gamma^{\mu}\gamma^{5}u -\bar{d}\gamma^{\mu}\gamma^{5}d)\), and each quark in the loop is weighted by the square of its electric charge and summed over colours, the electromagnetic vertices being colour-blind. The sum is

\begin{equation}\tag{102.17} N_{c}\left[Q_{u}^{2}-Q_{d}^{2}\right] =N_{c}\left(\frac{4}{9}-\frac{1}{9}\right)=\frac{N_{c}}{3}\ec \end{equation}

the relative minus sign coming from the isospin structure of the current. The amplitude is proportional to Equation (102.17) and the width to its square, which is Equation (102.16); the remaining factors — the anomaly coefficient itself, the reduction of the amplitude to the pion decay constant through the partial conservation of the axial current, and the two-body phase space — are derived in Section 106.7.2. The reason this count is sharper than \(R\) is that it is a low-energy theorem: the anomaly receives no corrections at higher order in \(\alpha_{s}\) (Theorem 106.89), so the only uncertainty is the measured \(F_{\pi}\), entering squared.

Remark 102.17 (Drell–Yan as a third, weaker count).

Drell and Yan proposed the production of a massive lepton pair in a hadron–hadron collision through the annihilation of a quark from one hadron with an antiquark from the other [Drell:1970]. Colour enters with the opposite sign to \(R\): the annihilating pair must be in a colour-singlet configuration to make a virtual photon, and if the quark and the antiquark are drawn independently from their parents the probability of matching colours is \(1/N_{c}\). The cross-section therefore carries a factor \(1/N_{c}\) multiplying parton densities that, being measured in the deep inelastic scattering of Section 102.5, already contain a sum over \(N_{c}\) colours. The absolute normalization is thus a count. It is a weaker one, because the leading-order prediction is corrected by an \(\mathcal{O}(\alpha_{s})\) factor of about \(1.5\) to \(2\) — the so-called \(K\) factor, which is not small — so the measurement tests \(N_{c}=3\) at the tens-of-percent level rather than the percent level. It is reported here as a consistency check with an independent sign structure, which is what it is, and not as a precision determination.

The QCD Lagrangian

Yang–Mills theory

Colour, as introduced in Definition 102.10, is a global label. Theorem 100.5 showed what happens when a global phase symmetry is required to hold point by point: an interaction is not chosen but forced. Yang and Mills asked the same question of a non-abelian group [Yang:1954], taking isospin as their example, and the answer differs from the abelian one in exactly one respect — the mediating field is charged under the symmetry it mediates.

Notation 102.18 (Colour fields and their SI dimensions).

The quark field \(q_{f}^{i}\) carries a flavour label \(f\) and a colour index \(i=1,2,3\), and has the dimension \(\mathrm{m}^{-3/2}\) of any Dirac field (Definition 100.2). The eight gauge potentials \(A^{a}_{\mu}\), \(a=1,\dots,8\), carry the dimension \(\mathrm{V}\,\mathrm{s}/\mathrm{m}\) of the electromagnetic potential, and the colour charge \(g_{s}\) carries the dimension \(\mathrm{C}\) of an electric charge. The combination that appears everywhere is

\begin{equation}\tag{102.18} \mathcal{A}_{\mu}:=\frac{g_{s}}{\hbar}A^{a}_{\mu}T_{a}\ec\qquad \left[\mathcal{A}_{\mu}\right]=/\mathrm{m}\ec \end{equation}

which is Equation (106.63). The dimensionless measure of the coupling is

\begin{equation}\tag{102.19} \alpha_{s}:=\frac{g_{s}^{2}}{4\pi\varepsilon_{0}\hbar c} =\frac{\mu_{0}c\,g_{s}^{2}}{4\pi\hbar}\ec \end{equation}

the exact analogue of Equation (100.3) and equal to \(\hat{g}^{2}/4\pi\) in the notation of Equation (106.66). That \(\alpha_{s}\) is a pure number is, as for \(\alpha\), a fact and not a convention: \(g_{s}^{2}/\varepsilon_{0}\) carries \(\mathrm{J}\,\mathrm{m}\), exactly as \(\hbar c\) does.

Theorem 102.19 (The non-abelian gauge principle).

Require the free quark Lagrangian \(\Lag_{0}=\bar{q}(\ii\hbar c\,\gamma^{\mu}\pp_{\mu}-mc^{2})q\), summed over colours, to be invariant under

\begin{equation}\tag{102.20} q(x)\longmapsto U(x)\,q(x)\ec\qquad U(x)=\exp\!\left[\ii\theta^{a}(x)T_{a}\right]\in\SU(3)_{c}\ec \end{equation}

with \(\theta^{a}\) arbitrary real functions. Then invariance holds if and only if eight vector fields \(A^{a}_{\mu}\) are introduced, the derivative is replaced by

\begin{equation}\tag{102.21} D_{\mu}:=\pp_{\mu}+\frac{\ii g_{s}}{\hbar}A^{a}_{\mu}T_{a} =\pp_{\mu}+\ii\mathcal{A}_{\mu}\ec \end{equation}

and the potentials transform as

\begin{equation}\tag{102.22} \mathcal{A}_{\mu}\longmapsto U\mathcal{A}_{\mu}U^{\dagger} +\ii\left(\pp_{\mu}U\right)U^{\dagger}\ec \end{equation}

which infinitesimally reads

\begin{equation}\tag{102.23} \delta A^{a}_{\mu}=-\frac{\hbar}{g_{s}}\,\pp_{\mu}\theta^{a} -f^{abc}\theta^{b}A^{c}_{\mu}\ep \end{equation}

Rests on Definition 102.10, Equation (14.67) and Equation (100.8).

Proof.

Derives Theorem 102.19. Under Equation (102.20) the derivative picks up an inhomogeneous term, \(\pp_{\mu}q\mapsto U\pp_{\mu}q +(\pp_{\mu}U)q\), so \(\Lag_{0}\) is not invariant. Postulate a matrix field \(\mathcal{A}_{\mu}\) valued in the Lie algebra and define \(D_{\mu}\) by Equation (102.21). Demanding \(D_{\mu}q\mapsto U(D_{\mu}q)\) gives

\[ \left(\pp_{\mu}+\ii\mathcal{A}'_{\mu}\right)Uq =\left(\pp_{\mu}U\right)q+U\pp_{\mu}q +\ii\mathcal{A}'_{\mu}Uq \stackrel{!}{=}U\pp_{\mu}q+\ii U\mathcal{A}_{\mu}q\ec \]

for all \(q\), whence \(\ii\mathcal{A}'_{\mu}U=\ii U\mathcal{A}_{\mu}-(\pp_{\mu}U)\) and, multiplying on the right by \(U^{\dagger}\), Equation (102.22). Conversely Equation (102.22) makes \(\bar{q}\gamma^{\mu}D_{\mu}q\) invariant by the same computation read backwards, so the condition is necessary as well as sufficient.

For the infinitesimal form put \(U=\identity+\ii\theta^{a}T_{a}\) and keep first order. The first term of Equation (102.22) gives \(\mathcal{A}_{\mu}+\ii\theta^{b}\comm{T_{b}}{\mathcal{A}_{\mu}}\), and with \(\comm{T_{b}}{T_{c}}=\ii f_{bca}T_{a}\) of Equation (14.67) together with the total antisymmetry of \(f\) this is \(\mathcal{A}_{\mu}-(g_{s}/\hbar)f^{abc}\theta^{b}A^{c}_{\mu}T_{a}\). The second term gives \(\ii(\ii\pp_{\mu}\theta^{a}T_{a})=-\pp_{\mu}\theta^{a}T_{a}\). Dividing by \(g_{s}/\hbar\) yields Equation (102.23). For one abelian generator the second term of Equation (102.23) is absent and the first reproduces Equation (100.8) with \(\Lambda=-\hbar\theta/g_{s}\).

The whole difference between QED and QCD sits in the second term of Equation (102.23): the gluon potential does not merely shift under a gauge transformation, it rotates in colour space. It carries colour, and a field that carries the charge it mediates interacts with itself.

Definition 102.20 (The gluon field strength).
\begin{equation}\tag{102.24} F^{a}_{\mu\nu} :=\pp_{\mu}A^{a}_{\nu}-\pp_{\nu}A^{a}_{\mu} -\frac{g_{s}}{\hbar}f^{abc}A^{b}_{\mu}A^{c}_{\nu}\ec \end{equation}

of dimension \(\mathrm{T}\), equivalently \(\mathcal{F}_{\mu\nu}=(g_{s}/\hbar)F^{a}_{\mu\nu}T_{a} =-\ii\comm{D_{\mu}}{D_{\nu}}\), which is Equation (106.64).

Proposition 102.21 (Curvature and its transformation law).

\(\mathcal{F}_{\mu\nu}=-\ii\comm{D_{\mu}}{D_{\nu}}\) holds with Equation (102.24), and under Equation (102.22)

\begin{equation}\tag{102.25} \mathcal{F}_{\mu\nu}\longmapsto U\mathcal{F}_{\mu\nu}U^{\dagger}\ec \end{equation}

so that \(F^{a}_{\mu\nu}\) transforms in the adjoint representation \(\vect{8}\) and \(F^{a}_{\mu\nu}F^{a\,\mu\nu}\) is gauge invariant. Unlike Equation (100.11), \(F^{a}_{\mu\nu}\) is itself not invariant. Rests on Equations (14.66), (102.22) and (102.24).

Proof.

Derives Proposition 102.21. Compute the commutator on a quark field:

\begin{align} \comm{D_{\mu}}{D_{\nu}} &=\ii\left(\pp_{\mu}\mathcal{A}_{\nu} -\pp_{\nu}\mathcal{A}_{\mu}\right) -\comm{\mathcal{A}_{\mu}}{\mathcal{A}_{\nu}}\nn\\ &=\ii\frac{g_{s}}{\hbar} \left(\pp_{\mu}A^{a}_{\nu}-\pp_{\nu}A^{a}_{\mu}\right)T_{a} -\left(\frac{g_{s}}{\hbar}\right)^{2}A^{b}_{\mu}A^{c}_{\nu} \,\ii f_{bca}T_{a}\ec \tag{102.26} \end{align}

the derivative terms acting on the quark cancelling between the two orderings. Factoring \(\ii g_{s}/\hbar\) and using \(f_{bca}=f_{abc}\) gives Equation (102.24). Since \(D_{\mu}\mapsto UD_{\mu}U^{\dagger}\) by construction — Equation (102.22) is exactly the statement that \(D_{\mu}q\) transforms like \(q\) — the commutator of two such objects transforms as \(U\comm{D_{\mu}}{D_{\nu}}U^{\dagger}\), which is Equation (102.25). Finally \(F^{a}F^{a}\propto\tr(\mathcal{F}_{\mu\nu}\mathcal{F}^{\mu\nu})\) by Equation (14.66), and a trace of a product of two adjoint objects is invariant by cyclicity — it is the Killing form of Proposition 14.53, which is what makes this the only quadratic invariant available.

Remark 102.22 (A sign convention, and how to convert).

Many texts define \(D_{\mu}=\pp_{\mu}-\ii\mathcal{A}_{\mu}\) and obtain \(F^{a}_{\mu\nu}\) with \(+g_{s}f^{abc}A^{b}A^{c}/\hbar\). The two conventions are related by \(g_{s}\to-g_{s}\), under which every physical quantity — which depends on \(g_{s}\) only through \(\alpha_{s}\propto g_{s}^{2}\), or through an even number of vertices — is unchanged. The convention adopted here is the one that makes Equation (102.21) the literal non-abelian copy of Equation (100.9) and agrees with Notation 106.50, so that formulae may be carried between the three chapters without a sign audit.

Proposition 102.23 (The gluon self-couplings).

The Yang–Mills Lagrangian density

\begin{equation}\tag{102.27} \Lag_{\mathrm{YM}}=-\frac{1}{4\mu_{0}} F^{a}_{\mu\nu}F^{a\,\mu\nu}\ec \end{equation}

of dimension \(\mathrm{J}/\mathrm{m}^{3}\), contains — besides the free quadratic term — a cubic and a quartic self-interaction of the gauge field,

\begin{align} \Lag_{\mathrm{YM}}= &-\frac{1}{4\mu_{0}} \left(\pp_{\mu}A^{a}_{\nu}-\pp_{\nu}A^{a}_{\mu}\right) \left(\pp^{\mu}A^{a\nu}-\pp^{\nu}A^{a\mu}\right)\nn\\ &+\frac{g_{s}}{2\mu_{0}\hbar}f^{abc} \left(\pp_{\mu}A^{a}_{\nu}-\pp_{\nu}A^{a}_{\mu}\right) A^{b\,\mu}A^{c\,\nu}\nn\\ &-\frac{g_{s}^{2}}{4\mu_{0}\hbar^{2}} f^{abc}f^{ade}A^{b}_{\mu}A^{c}_{\nu} A^{d\,\mu}A^{e\,\nu}\ec \tag{102.28} \end{align}

both proportional to the structure constants Equation (14.69) and therefore absent in an abelian theory. Rests on Equations (14.69), (100.12) and (102.24).

Proof.

Derives Proposition 102.23. Write \(F^{a}_{\mu\nu}=G^{a}_{\mu\nu} -(g_{s}/\hbar)f^{abc}A^{b}_{\mu}A^{c}_{\nu}\) with \(G^{a}_{\mu\nu}=\pp_{\mu}A^{a}_{\nu}-\pp_{\nu}A^{a}_{\mu}\) and square, the cross term appearing twice by the symmetry of the contraction. Multiplying by \(-1/4\mu_{0}\) gives Equation (102.28). For an abelian group all \(f^{abc}\) vanish and only the first line survives, which is Equation (100.12); the normalization \(1/4\mu_{0}\) is fixed by that limit. Dimensionally \(F^{2}/\mu_{0}\) carries \(\mathrm{T}^{2}\,\mathrm{A}^{2}/\mathrm{N} =\mathrm{J}/\mathrm{m}^{3}\), as in Equation (100.12).

Remark 102.24 (Yang and Mills's mass problem).

A Proca term \(\left(2\mu_{0}\right)^{-1}\left(m_{g}c/\hbar\right)^{2} A^{a}_{\mu}A^{a\,\mu}\), which is the only gluon mass term available and carries \(\mathrm{J}/\mathrm{m}^{3}\) as it must, is not invariant under Equation (102.23) — the inhomogeneous piece \(\pp_{\mu}\theta^{a}\) does not cancel — so gauge invariance forbids a gluon mass, exactly as it forbids a photon mass in Phenomenon 100.7. This was the objection to [Yang:1954] at the time: a massless charged vector particle would have been conspicuous, and none was seen. The two known resolutions are different from each other. For the weak interaction the symmetry is spontaneously broken and the gauge bosons acquire mass through the mechanism of Electroweak Unification and the Higgs Boson. For the strong interaction the gluon really is massless, and the interaction is nevertheless short-ranged because the theory confines (Section 102.7) — a resolution that requires the non-perturbative behaviour of the very self-couplings Equation (102.28) that caused the problem. It took nineteen years.

Gauge fixing, ghosts and the Slavnov–Taylor identities

The quadratic form of Equation (102.28) is singular for the same reason as Equation (100.15): the action does not depend on the gauge direction, so it cannot be inverted there and there is no propagator. In the abelian case the cure of Section 100.1.3 — add \(-(\pp\cdot A)^{2}/2\mu_{0}\xi\) and proceed — is complete, because the Jacobian accompanying the change of variables is a field-independent constant. In the non-abelian case it is not, and the difference is a new field.

Theorem 102.25 (Gauge fixing needs a determinant).

Fixing the gauge by a condition \(G^{a}(\mathcal{A})=0\) inside the functional integral of Path-Integral Quantization requires the insertion of the Faddeev–Popov determinant \(\det\left(\delta G^{a}/\delta\theta^{b}\right)\) [Faddeev:1967]. For the covariant condition \(G^{a}=\pp^{\mu}A^{a}_{\mu}\) this operator is \(\pp^{\mu}D_{\mu}\), which depends on the gauge field whenever the group is non-abelian; the determinant is therefore not a constant and is represented by a functional integral over a pair of anticommuting scalar fields \(c^{a}\), \(\bar{c}^{a}\) in the adjoint representation, the ghosts. Rests on Theorem 106.52, Equation (102.23) and Corollary 106.53.

Proof.

Derives Theorem 102.25. This is Theorem 106.52, proved there; only the group-theoretic input is quoted here. The variation of the gauge condition under Equation (102.23) is \(\delta G^{a}=-(\hbar/g_{s})\pp^{\mu} \left(\delta^{ab}\pp_{\mu}+ (g_{s}/\hbar) f^{acb}A^{c}_{\mu}\right) \theta^{b}\), whose kernel is the adjoint covariant derivative. The \(f^{acb}\) term is absent for an abelian group, which is Corollary 106.53 and the reason the photon needed no ghosts in Section 100.1.3.

The ghosts are not particles: they are anticommuting scalars, in violation of the spin-statistics theorem, and they may therefore never appear in an asymptotic state. Their function is negative — they cancel the contributions of the unphysical gluon polarizations. A massless vector field has four components and two physical (transverse) polarizations; in a covariant gauge all four propagate, and the timelike and longitudinal ones would spoil the unitarity of the theory. That the ghost loop cancels them exactly is Proposition 106.55, and it is the reason the practical quantization of QCD is the functional one of Path-Integral Quantization rather than the canonical route of Canonical Quantization of Fields.

What survives gauge fixing is a residual global symmetry mixing gauge field and ghosts: the BRST symmetry of Becchi, Rouet and Stora [Becchi:1974] [Becchi:1976] and, independently, Tyutin [Tyutin:1975]. Its generator is nilpotent (Theorem 106.57), the physical state space is its cohomology (Definition 106.59), and it is the modern statement of what gauge invariance becomes after the gauge has been fixed. Its practical consequence is a set of identities among renormalization constants, the Slavnov–Taylor identities [Taylor:1971] [Slavnov:1972], generalizing the Ward–Takahashi identity Theorem 100.48 of the abelian theory. Where the abelian identity gives the single relation \(Z_{1}=Z_{2}\) of Corollary 100.49 — charge renormalization is governed by the photon field renormalization alone — the non-abelian identities relate the quark–gluon, three-gluon, four-gluon and ghost–gluon vertex renormalizations to each other, so that the one coupling \(g_{s}\) of Equation (102.28) remains one coupling after renormalization. That is not automatic: three vertices renormalized independently would be three couplings, and the theory would lose its predictive content.

Theorem 102.26 (Non-abelian gauge theories are renormalizable).

A Yang–Mills theory with fermions in any representation, quantized with Theorem 102.25, is renormalizable: all ultraviolet divergences are removed to every order by a finite number of counterterms with the same structure as the terms already present [tHooft:1971a] [tHooft:1972]. Rests on Theorems 100.17 and 102.25.

Derivation pending.

Renormalizability of Yang–Mills theory. The proof combines the power-counting theorem, which for a gauge theory in four dimensions gives superficial degree of divergence four minus the number of external lines weighted by their dimension exactly as in the abelian case, with the Slavnov–Taylor identities, which show that the divergent structures allowed by power counting are constrained to be gauge invariant and therefore already present in the Lagrangian. The regularization must respect the gauge symmetry, and dimensional regularization does. This is a long argument of the same shape as the BPHZ theorem of the electromagnetic chapter and belongs beside it in Appendix A.

The Lagrangian and its parameters

Definition 102.27 (The QCD Lagrangian density).
\begin{equation}\tag{102.29} \Lag_{\mathrm{QCD}} =-\frac{1}{4\mu_{0}}F^{a}_{\mu\nu}F^{a\,\mu\nu} +\sum_{f}\bar{q}_{f}\left(\ii\hbar c\,\gamma^{\mu}D_{\mu} -m_{f}c^{2}\right)q_{f}\ec \end{equation}

with \(F^{a}_{\mu\nu}\) from Equation (102.24), \(D_{\mu}\) from Equation (102.21) acting on the colour triplet index, and the flavour sum running over \(f=u,d,s,c,b,t\). Both terms carry \(\mathrm{J}/\mathrm{m}^{3}\). This is the Lagrangian written down by Fritzsch, Gell-Mann and Leutwyler [Fritzsch:1973]. A further gauge-invariant term is allowed and is deferred to Section 102.10.

Proposition 102.28 (The parameters of QCD).

Equation (102.29) contains seven free parameters: the six quark masses and the coupling \(g_{s}\), equivalently \(\alpha_{s}\) at one reference scale. Nothing else — not the number of gluons, not their couplings to each other, not the relative strength with which any flavour couples to them — is adjustable. Rests on Equation (102.29), Proposition 14.50 and Theorem 102.26.

Proof.

Derives Proposition 102.28. The gauge sector is fixed by the group: the number of gauge fields is \(\dim\mathfrak{su}(3)=8\) by Proposition 14.50, and the cubic and quartic couplings of Equation (102.28) carry the structure constants Equation (14.69), which are numbers, multiplied by powers of the same \(g_{s}\) that appears in the quark vertex. The quark sector contributes one mass per flavour. That the coupling is the same for every flavour is the non-abelian version of Proposition 100.6: \(D_{\mu}\) depends on the matter field only through its colour representation, and every quark is a \(\vect{3}\). Renormalizability (Theorem 102.26) is what guarantees the count is not enlarged order by order, and the Slavnov–Taylor identities are what keep the three gauge vertices tied to one number.

ParameterValueSI equivalent
$m_{u}$$2.16(7)\,\mathrm{MeV}/c^{2}$$3.85\times 10^{-30}\,\mathrm{kg}$
$m_{d}$$4.70(7)\,\mathrm{MeV}/c^{2}$$8.38\times 10^{-30}\,\mathrm{kg}$
$m_{s}$$93.5(8)\,\mathrm{MeV}/c^{2}$$1.667\times 10^{-28}\,\mathrm{kg}$
$m_{c}$$1.273(5)\,\mathrm{GeV}/c^{2}$$2.269\times 10^{-27}\,\mathrm{kg}$
$m_{b}$$4.183(7)\,\mathrm{GeV}/c^{2}$$7.457\times 10^{-27}\,\mathrm{kg}$
$m_{t}$$172.57(29)\,\mathrm{GeV}/c^{2}$$3.0763\times 10^{-25}\,\mathrm{kg}$
$\alpha_{s}(m_{Z}c^{2})$$0.1180(9)$dimensionless
The seven parameters of Equation (102.29). Quark masses are $\overline{\mathrm{MS}}$ values, those of $u$, $d$ and $s$ quoted at the scale $2\,\mathrm{GeV}$ and those of $c$, $b$, $t$ at their own mass, in the scheme of Definition 100.54; the scheme dependence is not an uncertainty but part of the definition, because a quark mass is not the pole of a propagator of anything observable. Values from [Navas:2024].
Proposition 102.29 (Feynman rules for QCD in SI form).

The rules of Proposition 100.14 carry over with the following replacements and additions, each factor written with its \(\hbar\) and \(c\) in place.

  1. Quark propagator: Equation (100.23) times \(\delta_{ij}\) in colour.

  2. Gluon propagator: Equation (100.17) times \(\delta^{ab}\).

  3. Quark–gluon vertex:

    \begin{equation}\tag{102.30} -\frac{\ii g_{s}}{\hbar}\gamma^{\mu}\left(T_{a}\right)_{ij}\ec \end{equation}

    of dimension \(\mathrm{C}/\mathrm{J}/\mathrm{s}\), identical to Equation (100.24) with \(q\to g_{s}T_{a}\).

  4. Three-gluon vertex: from the second line of Equation (102.28), a factor proportional to \((g_{s}/\hbar\mu_{0})f^{abc}\) times a momentum-dependent tensor antisymmetric in the simultaneous exchange of any two legs.

  5. Four-gluon vertex: from the third line, proportional to \((g_{s}/\hbar)^{2}\mu_{0}^{-1}\) times sums of products \(f^{abe}f^{cde}\).

  6. Ghost propagator and ghost–gluon vertex from Theorem 102.25, with a factor \(-1\) for each closed ghost loop as for a fermion loop.

Every colour factor is a contraction of the \(T_{a}\) and \(f^{abc}\) computed once and for all in Proposition 14.52. The two that occur constantly are

\begin{equation}\tag{102.31} \left(T_{a}T_{a}\right)_{ij}=C_{F}\,\delta_{ij}\ec\quad C_{F}=\frac{4}{3}\ec\qquad f^{acd}f^{bcd}=C_{A}\,\delta^{ab}\ec\quad C_{A}=3\ec \end{equation}

being the quadratic Casimirs Equation (14.77) and Equation (14.78) of the fundamental and adjoint representations, and \(\tr\left(T_{a}T_{b}\right)=T_{F}\delta_{ab}\) with \(T_{F}=\tfrac{1}{2}\) from Equation (14.66). Rests on Proposition 100.14, Equation (102.28) and Theorem 102.25.

Proof.

Derives Proposition 102.29. Rules 1–3 are Proposition 100.14 with the substitution \(q\gamma^{\mu}\to g_{s}\gamma^{\mu}T_{a}\) forced by Equation (102.21), together with the observation that the free parts of Equation (102.29) are diagonal in colour; rules 4–6 are read off the interaction terms of Equation (102.28) and of Theorem 102.25 by the same expansion of \(\exp(\ii S/\hbar)\). For Equation (102.31), \(T_{a}T_{a}\) commutes with every generator and is therefore a multiple of the identity on the irreducible \(\vect{3}\); Equation (14.77) evaluates the multiple as \(4/3\), and Equation (14.78) does the same for the adjoint, where the generators are \(-\ii f^{abc}\).

Remark 102.30 (Flavour independence is a prediction, and it is tested).

Proposition 102.28 asserts that one number governs the coupling of the gluon to every quark flavour. The claim is testable because \(\alpha_{s}\) is extracted from processes dominated by different flavours: \(\tau\) decays involve \(u\), \(d\) and \(s\); bottomonium spectroscopy involves \(b\); the \(Z\) hadronic width involves all five accessible flavours democratically. That these determinations agree after being evolved to a common scale, as Section 102.4.2 records, is a test of flavour independence and not merely a measurement of a number.

Asymptotic freedom

The sign of the beta function

Remark 100.34 read the running of \(\alpha\) as the polarization of a medium: virtual pairs align against an inserted charge, so that a distant probe sees less charge than a close one, and the effective coupling grows with momentum transfer. The reading is literal, and in SI units it can be pushed further than it usually is, because SI keeps the vacuum's electric and magnetic responses as separate named quantities. That separation is what makes the sign of the non-abelian beta function computable by a short argument instead of a two-loop-sized diagram count.

Lemma 102.31 (The vacuum's two responses are reciprocal).

Let the polarization of the vacuum by virtual pairs be described by an effective permittivity \(\varepsilon_{\mathrm{eff}}\) and permeability \(\mu_{\mathrm{eff}}\). Lorentz invariance of the vacuum requires

\begin{equation}\tag{102.32} \varepsilon_{\mathrm{eff}}\,\mu_{\mathrm{eff}}=\frac{1}{c^{2}}\ec \qquad\text{hence}\qquad \frac{\varepsilon_{\mathrm{eff}}}{\varepsilon_{0}} =\frac{\mu_{0}}{\mu_{\mathrm{eff}}}\ep \end{equation}

Consequently a paramagnetic vacuum, \(\mu_{\mathrm{eff}}>\mu_{0}\), is one with \(\varepsilon_{\mathrm{eff}}<\varepsilon_{0}\): it antiscreens. Rests on Postulate 38.2 and Equation (102.19).

Proof.

Derives Lemma 102.31. In a medium of permittivity \(\varepsilon\) and permeability \(\mu\) the phase velocity of light is \(\left(\varepsilon\mu\right)^{-1/2}\). The vacuum, polarized or not, is Lorentz invariant, and by the postulates of Lorentz Transformations light propagates in it at \(c\) in every frame; a medium with \(\varepsilon\mu\neq c^{-2}\) would define a rest frame. Hence Equation (102.32). The second form follows from \(\varepsilon_{0}\mu_{0}c^{2}=1\), and the statement about signs is immediate. Finally, the coupling measured at a scale is \(\alpha_{s}=g_{s}^{2}/ \left(4\pi\varepsilon_{\mathrm{eff}}\hbar c\right)\) by Equation (102.19), so \(\varepsilon_{\mathrm{eff}}<\varepsilon_{0}\) means a larger effective coupling at long distance and a smaller one at short distance, which is asymptotic freedom.

The problem is therefore reduced to a magnetostatic one: is the QCD vacuum paramagnetic or diamagnetic? A magnetic response has two sources, and they compete. Orbital motion of a charged particle in a field is diamagnetic — Lenz's law, and Landau diamagnetism. Alignment of an intrinsic magnetic moment with the field is paramagnetic. For a charged particle of spin \(s\) and gyromagnetic ratio \(g\) the two enter the energy in a fixed ratio, and the following lemma is where the ratio is fixed.

Lemma 102.32 (Landau levels with $g=2$).

A massless particle of charge \(q\), spin projection \(s_{z}\) along a uniform magnetic field \(B\hat{z}\) and gyromagnetic ratio \(g=2\) has energies

\begin{equation}\tag{102.33} E_{n}^{2}(p_{z})=c^{2}p_{z}^{2} +\abs{q}\hbar c^{2}B\left(2n+1+\kappa\right)\ec\qquad \kappa:=-2s_{z}\ec\quad n=0,1,2,\dots\ec \end{equation}

each level carrying the degeneracy \(\abs{q}B/(2\pi\hbar)\) per unit area transverse to the field. The quantity \(\abs{q}\hbar c^{2}B\) has the dimension of a squared energy. Rests on Equations (93.2) and (102.28).

Proof.

Derives Lemma 102.32. For a spinless particle the transverse motion is a harmonic oscillator of frequency \(\omega_{c}=\abs{q}c^{2}B/E\), whose quantization gives \(E^{2}=c^{2}p_{z}^{2}+(2n+1)\abs{q}\hbar c^{2}B\) — the relativistic Landau problem, obtained by squaring the wave operator so that \(\vect{p}_{\perp}^{2}\to(2n+1)\abs{q}\hbar B\). A magnetic moment \(\vect{\mu}=g\left(q/2m\right)\vect{S}\) adds an interaction energy \(-\vect{\mu}\cdot\vect{B}\), which in the squared equation appears as \(-g\abs{q}\hbar c^{2}Bs_{z}\); putting \(g=2\) gives Equation (102.33). That \(g=2\) is the Dirac value for a spin-\(\tfrac{1}{2}\) particle (The Dirac Equation), and its small radiative correction (Theorem 100.41) is of higher order in the coupling and irrelevant to a one-loop coefficient. That \(g=2\) also for the charged gluons is a property of Equation (102.28): expanding \(-F^{a}_{\mu\nu}F^{a\,\mu\nu}/4\mu_{0}\) to second order about a background field produces, besides the covariant Laplacian, exactly the term \(-2\ii\mathcal{F}_{\mu\nu}\) acting on the vector index of the fluctuation, and a coefficient \(2\) in front of the field-strength coupling to the spin generator is the statement \(g=2\). The degeneracy \(\abs{q}B/(2\pi\hbar)\) per unit area is the standard one: one state per flux quantum \(2\pi\hbar/\abs{q}\).

Theorem 102.33 (The one-loop coefficient, from a magnetic response).

Let a gauge theory contain species \(i\) of spin \(s_{i}\) carrying charge \(x_{i}g_{s}\) under a chosen Cartan direction, the label \(i\) running over all one-particle states (particles and antiparticles counted separately). Then the vacuum behaves as a medium with

\begin{equation}\tag{102.34} \frac{\varepsilon_{\mathrm{eff}}}{\varepsilon_{0}} =1+\frac{\alpha_{s}}{24\pi}\,S\, \ln\frac{\Lambda_{E}^{2}}{E^{2}}\ec\qquad S:=\sum_{i}x_{i}^{2}\,(-1)^{2s_{i}} \sum_{s_{z}}\left(1-12s_{z}^{2}\right)\ec \end{equation}

with \(\Lambda_{E}\) an ultraviolet cutoff energy and \(E\) the energy scale at which the coupling is measured; and consequently

\begin{equation}\tag{102.35} \frac{\dd\alpha_{s}}{\dd\ln E}=\frac{S}{12\pi}\,\alpha_{s}^{2} +O\!\left(\alpha_{s}^{3}\right)\ep \end{equation}

Inside the spin sum the term \(1\) is the diamagnetic contribution of the orbital motion and the term \(-12s_{z}^{2}\) the paramagnetic contribution of the magnetic moment. Rests on Lemma 102.31, Lemma 102.32 and Equation (102.19).

Proof.

Derives Theorem 102.33. Put the vacuum in a uniform field \(B\) pointing along the chosen Cartan direction in colour space, so that every field component has a definite charge \(x_{i}g_{s}\) and the problem is the abelian one of Lemma 102.32. The vacuum energy density is the sum of the zero-point energies of all modes,

\begin{equation}\tag{102.36} u=\frac{1}{2}\sum_{i}(-1)^{2s_{i}} \frac{\abs{q_{i}}B}{2\pi\hbar} \sum_{s_{z}}\sum_{n=0}^{\infty} \int_{-\infty}^{\infty}\frac{\dd p_{z}}{2\pi\hbar}\, E_{n}(p_{z})\ec \end{equation}

the factor \((-1)^{2s}\) being the sign of a fermionic zero-point energy, \(-\tfrac{1}{2}\hbar\omega\) per mode instead of \(+\tfrac{1}{2}\hbar\omega\).

Cut the longitudinal momentum off at \(\abs{p_{z}}\leq P\) and write \(\varepsilon_{n}^{2}=A\left(2n+1+\kappa\right)\) with \(A=\abs{q}\hbar c^{2}B\). Then

\begin{equation}\tag{102.37} \int_{-P}^{P}\frac{\dd p_{z}}{2\pi\hbar} \sqrt{c^{2}p_{z}^{2}+\varepsilon_{n}^{2}} =\frac{cP^{2}}{2\pi\hbar}+\frac{\varepsilon_{n}^{2}}{4\pi\hbar c} -\frac{\varepsilon_{n}^{2}}{4\pi\hbar c} \ln\frac{\varepsilon_{n}^{2}}{4c^{2}P^{2}} +O\!\left(P^{-2}\right)\ec \end{equation}

the first two terms being independent of \(B\) and of \(n\) in a way that contributes no \(B^{2}\ln B\), and the third being the one that matters. Split \(\ln\varepsilon_{n}^{2}=\ln A+\ln(2n+1+\kappa)\): the second piece carries no \(B\) and produces a term proportional to \(B^{2}\) with no logarithm, i.e. a renormalization of the field energy and not a running. Keeping the first piece,

\begin{equation}\tag{102.38} u_{\ln}=-\frac{1}{2}\sum_{i}(-1)^{2s_{i}} \frac{\abs{q_{i}}B}{2\pi\hbar}\cdot \frac{A_{i}}{4\pi\hbar c}\, \ln\frac{A_{i}}{4c^{2}P^{2}} \sum_{s_{z}}\sum_{n=0}^{\infty}\left(2n+1+\kappa\right)\ep \end{equation}

The sum over Landau levels diverges and is defined by analytic continuation: writing \(2n+1+\kappa=2(n+x)\) with \(x=\left(1+\kappa\right)/2\), and using the Hurwitz zeta function \(\zeta(\sigma,x)=\sum_{n\geq0}(n+x)^{-\sigma}\) continued to \(\sigma=-1\), where \(\zeta(-1,x)=-\tfrac{1}{2}\left(x^{2}-x+\tfrac{1}{6}\right)\),

\begin{equation}\tag{102.39} \sum_{n=0}^{\infty}\left(2n+1+\kappa\right) =2\zeta(-1,x)=-\left(x^{2}-x+\tfrac{1}{6}\right) =\frac{1-3\kappa^{2}}{12} =\frac{1-12s_{z}^{2}}{12}\ec \end{equation}

using \(\kappa=-2s_{z}\). The two terms of Equation (102.39) are the two magnetic responses: the \(1\) comes from the \(2n+1\) of the orbital motion and the \(-12s_{z}^{2}\) from the \(\kappa\) of the moment.

Substituting \(A_{i}=\abs{q_{i}}\hbar c^{2}B\) into Equation (102.38) and writing \(L_{i}=\ln\left[4c^{2}P^{2}/(\abs{q_{i}}\hbar c^{2}B)\right]\),

\begin{equation}\tag{102.40} u_{\ln}=\frac{cB^{2}}{16\pi^{2}\hbar}\sum_{i}q_{i}^{2} (-1)^{2s_{i}}L_{i} \sum_{s_{z}}\frac{1-12s_{z}^{2}}{12}\ep \end{equation}

Comparing with the field energy density \(u=B^{2}/2\mu_{\mathrm{eff}}\) and using Equation (102.32) together with \(q_{i}^{2}=x_{i}^{2}g_{s}^{2}\) and \(g_{s}^{2}/(\varepsilon_{0}\hbar c)=4\pi\alpha_{s}\) from Equation (102.19),

\begin{equation}\tag{102.41} \frac{\varepsilon_{\mathrm{eff}}}{\varepsilon_{0}} =\frac{\mu_{0}}{\mu_{\mathrm{eff}}} =1+\frac{\alpha_{s}}{24\pi}\sum_{i}x_{i}^{2}(-1)^{2s_{i}} \left[\sum_{s_{z}}\left(1-12s_{z}^{2}\right)\right]L_{i}\ep \end{equation}

The species-dependence of \(L_{i}\) is a \(B\)-independent additive constant under the logarithm and is therefore a choice of subtraction scheme, not physics: only the coefficient of \(\ln B\) is common to all schemes, which is the standard statement that a one-loop beta-function coefficient is scheme-independent (Proposition 100.55). Identifying the energy scale probed by the field, \(E^{2}\sim g_{s}\hbar c^{2}B\) — the field sets the transverse momentum of the modes — and writing \(\Lambda_{E}=2cP\), all \(L_{i}\) become \(\ln(\Lambda_{E}^{2}/E^{2})\) up to constants, which is Equation (102.34).

Finally \(\alpha_{s}(E) =g_{s}^{2}/\left(4\pi\varepsilon_{\mathrm{eff}}\hbar c\right)\), so \(\alpha_{s}(E)^{-1}=\alpha_{s}(\Lambda_{E})^{-1} +\left(S/24\pi\right)\ln(\Lambda_{E}^{2}/E^{2})\). Differentiating with respect to \(\ln E\) gives \(-\alpha_{s}^{-2}\,\dd\alpha_{s}/\dd\ln E=-S/12\pi\), which is Equation (102.35).

Corollary 102.34 (The master formula reproduces QED).

For a single Dirac fermion of unit charge, Equation (102.35) gives \(\dd\alpha/\dd\ln E=2\alpha^{2}/3\pi\), which is Equation (100.84). Rests on Equations (100.84), (102.35) and (102.39).

Proof.

Derives Corollary 102.34. The one-particle states are the fermion and the antifermion, each of charge magnitude \(e\), so \(\sum_{i}x_{i}^{2}=2\). The statistics factor is \((-1)^{2s}=-1\) and the spin sum over \(s_{z}=\pm\tfrac{1}{2}\) is \(2\left(1-12\cdot\tfrac{1}{4}\right)=-4\). Hence \(S=2\cdot(-1)\cdot(-4) =8\) and \(\dd\alpha/\dd\ln E=8\alpha^{2}/12\pi=2\alpha^{2}/3\pi\). The sign is positive — QED screens — and by Equation (102.39) the reason is arithmetical: for \(s_{z}=\pm\tfrac{1}{2}\) the paramagnetic term \(12s_{z}^{2}=3\) beats the diamagnetic \(1\), so the electron vacuum is in fact paramagnetic, and it is the Fermi statistics — the minus sign of a fermion loop — that turns the result back into screening.

Theorem 102.35 (Asymptotic freedom).

For \(\SU(3)_{c}\) with \(n_{f}\) quark flavours light enough to be excited at the scale considered,

\begin{equation}\tag{102.42} \frac{\dd\alpha_{s}}{\dd\ln E} =-\frac{b_{0}}{2\pi}\,\alpha_{s}^{2} +O\!\left(\alpha_{s}^{3}\right)\ec\qquad b_{0}=\frac{11}{3}C_{A}-\frac{4}{3}T_{F}n_{f} =11-\frac{2}{3}n_{f}\ec \end{equation}

which is negative — the coupling decreases with energy — for every \(n_{f}\leq16\) [Gross:1973] [Politzer:1973]. Nature supplies \(n_{f}=6\). Rests on Theorem 102.33, Equation (14.80) and Equation (102.31).

Proof.

Derives Theorem 102.35. Apply Theorem 102.33 twice.

Quarks. A quark flavour is a colour triplet; taking the Cartan direction to be \(T_{3}\), its three colour states carry charges \(x=+\tfrac{1}{2},-\tfrac{1}{2},0\), and each is a Dirac field with its own antiparticle. Hence \(\sum_{i}x_{i}^{2}=2\tr\left(T_{3}^{2}\right)=2T_{F}=1\) per flavour, the factor \(2\) counting antiparticles and \(T_{F}=\tfrac{1}{2}\) being Equation (14.66). With the spin factor \((-1)\times(-4)=+4\) from Corollary 102.34, quarks contribute \(S_{q}=4n_{f}\).

Gluons. The gluons form the adjoint \(\vect{8}\), whose charges under \(\ad_{T_{3}}\) are the first components of the roots Equation (14.80), namely \(\pm1\), \(\pm\tfrac{1}{2}\), \(\pm\tfrac{1}{2}\) and two zeros, so that \(\sum_{a}x_{a}^{2}=f^{3cd}f^{3cd}=C_{A}=3\) by Equation (102.31). Here the charged components already occur in \(\pm\) pairs — a gluon of charge \(+1\) is the antiparticle of the one with charge \(-1\) — so no further doubling is applied. A gluon is massless and has two physical polarizations, \(s_{z}=\pm1\); the unphysical ones are removed by the ghosts of Theorem 102.25, which is where that machinery earns its keep. The statistics factor is \((-1)^{2}=+1\) and the spin sum is

\begin{equation}\tag{102.43} \sum_{s_{z}=\pm1}\left(1-12s_{z}^{2}\right)=2\left(1-12\right)=-22\ec \end{equation}

so \(S_{g}=3\times(-22)=-66\).

Adding, \(S=4n_{f}-66=-6\left(11-\tfrac{2}{3}n_{f}\right)=-6b_{0}\), and Equation (102.35) becomes Equation (102.42). Restoring the group factors in place of their numerical values, the quark term is \(-\tfrac{4}{3}T_{F}n_{f}\) and the gluon term \(+\tfrac{11}{3}C_{A}\), which is the form used for any gauge group. \(b_{0}>0\) requires \(n_{f}<33/2\), i.e. \(n_{f}\leq16\).

Remark 102.36 (Eleven is twelve minus one).

The whole of asymptotic freedom is in Equation (102.43). A gluon carries spin \(1\), so its paramagnetic term is \(12s_{z}^{2}=12\) against a diamagnetic term of \(1\): the moment wins by eleven to one, the vacuum is strongly paramagnetic, and by Lemma 102.31 a paramagnetic vacuum antiscreens. The famous \(11\) of \(b_{0}\) is literally \(12-1\). The electron's moment also wins, three to one, but a fermion loop enters with the opposite sign and the result is screening. So the two theories differ not because the gluon is charged in some vague sense, but because spin one triples the moment term relative to spin one half while Bose statistics declines to flip its sign. This is the contrast Remark 100.34 promised, made quantitative: QED has \(S=8\) per unit-charge fermion and QCD has \(S=-66\) from the gluons alone.

Remark 102.37 (What the short calculation costs).

Three things in the derivation above are shortcuts, and none of them affects the coefficient. First, the sum Equation (102.39) is divergent and was assigned a value by analytic continuation. A proper-time or dimensional regulator (Definition 100.30) gives the same coefficient of \(\ln B\), which is the only quantity used; the finite parts differ and are scheme, as Equation (102.41) already noted. Second, the constant chromomagnetic field is not the true vacuum: the \(n=0\), \(s_{z}=+1\) gluon mode of Equation (102.33) has \(E^{2}=c^{2}p_{z}^{2}-\abs{q}\hbar c^{2}B\), which is negative for small \(p_{z}\), so the configuration is unstable and the effective action acquires an imaginary part [Nielsen:1978]. The instability is real physics but it does not touch the real part, which is where the logarithm and hence \(b_{0}\) live. Third, only the physical gluon polarizations were counted, the unphysical ones being assumed cancelled by ghosts; the gauge-invariant background-field computation does this explicitly and confirms Equation (102.42) [Gross:1973] [Politzer:1973] [tHooft:1972]. The argument is given in this form because it exhibits the mechanism — a competition between two magnetic responses that SI units keep visibly distinct — rather than only the answer.

Definition 102.38 (The scale of QCD).

Integrating Equation (102.42) between two energies gives

\begin{equation}\tag{102.44} \frac{1}{\alpha_{s}(E)}=\frac{1}{\alpha_{s}(E_{0})} +\frac{b_{0}}{2\pi}\ln\frac{E}{E_{0}}\ec \end{equation}

so that the pair \(\left(\alpha_{s},E_{0}\right)\) may be traded for the single energy \(\Lambda_{\mathrm{QCD}}\) at which the right-hand side formally vanishes:

\begin{equation}\tag{102.45} \alpha_{s}(E)=\frac{2\pi}{b_{0}\, \ln\left(E/\Lambda_{\mathrm{QCD}}\right)}\ec\qquad \Lambda_{\mathrm{QCD}} =E_{0}\exp\!\left[-\frac{2\pi}{b_{0}\alpha_{s}(E_{0})}\right]\ep \end{equation}
Phenomenon 102.39 (A mass scale from a dimensionless coupling).

The Lagrangian Equation (102.29) with the light quark masses set to zero contains no dimensionful parameter whatever: \(\hbar\) and \(c\) are conversion factors, and \(\alpha_{s}\) is a pure number by Equation (102.19). Nevertheless the theory it defines has a scale, and the scale is of the observed size: \(\Lambda_{\mathrm{QCD}}\) is a few hundred \(\mathrm{MeV}\), corresponding to a length \(\hbar c/\Lambda_{\mathrm{QCD}}\) of about \(10^{-15}\,\mathrm{m}\) — the size of a hadron. Almost all of the mass of ordinary matter is set by this number, which appears in no line of the Lagrangian. Rests on Equations (102.19), (102.29) and (102.45).

Derivation. Derives Phenomenon 102.39. This is dimensional transmutation, and Equation (102.45) is the whole of it. A statement of the form “\(\alpha_{s}=0.118\)” is meaningless without naming the scale at which it holds, because Equation (102.44) says the number changes with the scale; and a statement of the form “\(\alpha_{s}=0.118\) at \(E_{0}\)” contains one dimensionless number and one energy. The combination in Equation (102.45) is invariant under changing \(E_{0}\): differentiating \(\ln\Lambda_{\mathrm{QCD}}=\ln E_{0}-2\pi/ \left(b_{0}\alpha_{s}(E_{0})\right)\) with respect to \(\ln E_{0}\) gives \(1+\left(2\pi/b_{0}\alpha_{s}^{2}\right) \dd\alpha_{s}/\dd\ln E_{0}=1-1=0\) by Equation (102.42). So the theory has exactly one free parameter in its massless limit, and that parameter is an energy. The classical scale invariance of the massless Lagrangian is broken by the quantum theory — it is the trace anomaly, the same phenomenon in another guise — and \(\Lambda_{\mathrm{QCD}}\) is what the breaking produces. Numerically, taking \(E_{0}=m_{Z}c^{2}=91.188\,\mathrm{GeV}\) and \(\alpha_{s}=0.1180\) [Navas:2024] with \(n_{f}=5\), so that \(b_{0}=11-\tfrac{10}{3}=\tfrac{23}{3}\),

\begin{equation}\tag{102.46} \Lambda_{\mathrm{QCD}}^{(5)}\Big|_{\text{one loop}} =91.188\,\mathrm{GeV}\times\ee^{-6.945} =88\,\mathrm{MeV}\ec \end{equation}

and at four loops in the \(\overline{\mathrm{MS}}\) scheme the same data give about \(0.21\,\mathrm{GeV}\) for five active flavours and about \(0.33\,\mathrm{GeV}\) for three [Navas:2024]. The factor of two between Equation (102.46) and the four-loop value is not a discrepancy: \(\Lambda\) is defined by a truncation of Equation (102.42) and by a scheme, and changes when either changes, whereas \(\alpha_{s}\) at a stated scale does not. That is why the Particle Data Group quotes \(\alpha_{s}(m_{Z}c^{2})\) and not \(\Lambda\). What is robust, and what Phenomenon 102.39 asserts, is the order of magnitude and the fact that a scale exists at all.

Remark 102.40 (The scale that no parameter contains).

It is worth pausing on what has happened. The proton mass is \(938\,\mathrm{MeV}\)\(/c^{2}\); the quark masses in Table 102.1 that build it sum to \(9\,\mathrm{MeV}\)\(/c^{2}\). The remaining \(99\,\mathrm{\%}\) is \(\Lambda_{\mathrm{QCD}}\), and \(\Lambda_{\mathrm{QCD}}\) is not a parameter of the theory: it is the scale at which a dimensionless coupling, run down from wherever it was measured, becomes of order one. Had \(\alpha_{s}(m_{Z}c^{2})\) been \(0.09\) instead of \(0.118\), Equation (102.45) would have put \(\Lambda_{\mathrm{QCD}}\) at a few \(\mathrm{MeV}\) and nucleons would have been a hundred times lighter and a hundred times larger. The exponential in Equation (102.45) is what converts a modest change in a dimensionless number into an enormous change in a mass, and it is the only mechanism in the Standard Model that generates a scale without being given one.

The measured running of $\alpha_{s}$

Phenomenon 102.41 (The strong coupling weakens with energy).

Determinations of the strong coupling from \(\tau\) decays, from lattice computation, from deep inelastic scattering, from hadronic event shapes and from \(Z\)-pole observables span more than two decades in momentum transfer and fall on one decreasing curve, with

\begin{equation}\tag{102.47} \alpha_{s}(m_{Z}c^{2})=0.1180(9) \end{equation}

at the \(Z\) mass [Navas:2024]. The strong interaction is therefore weak at short distance and strong at hadronic distances — the opposite of the behaviour established in Quantum Electrodynamics and Renormalization, and the reason any perturbative calculation of a strong process is possible at all. Rests on Equations (102.42) and (102.44).

Derivation. Derives Phenomenon 102.41. The curve is Equation (102.44) with Equation (102.42), together with the rule that a quark decouples below its own mass — the decoupling theorem [Appelquist:1975] already used in Proposition 100.63 — so that \(n_{f}\) and hence \(b_{0}\) change at each flavour threshold and \(\alpha_{s}\) is matched continuously across it. Starting from Equation (102.47) and integrating downwards through the \(b\) threshold and upwards through the \(t\) threshold with the masses of Table 102.1 gives Table 102.2. The prediction is not one number but the shape of a curve over three decades, and the fact that determinations using entirely different physics land on it is the content of the phenomenon.

$E$$n_{f}$$\alpha_{s}(E)$ one loop$\hbar c/E$
$1.777\,\mathrm{GeV}$ ($m_{\tau}c^{2}$)4$0.279$$1.1\times 10^{-16}\,\mathrm{m}$
$4.183\,\mathrm{GeV}$ ($m_{b}c^{2}$)5$0.212$$4.7\times 10^{-17}\,\mathrm{m}$
$10\,\mathrm{GeV}$5$0.173$$2.0\times 10^{-17}\,\mathrm{m}$
$91.188\,\mathrm{GeV}$ ($m_{Z}c^{2}$)5$0.1180$$2.2\times 10^{-18}\,\mathrm{m}$
$1\,\mathrm{TeV}$6$0.089$$2.0\times 10^{-19}\,\mathrm{m}$
The one-loop running Equation (102.44) evolved from Equation (102.47), with $n_{f}$ changing at the quark masses of Table 102.1. The last column is the distance $\hbar c/E$ resolved at that energy. The one-loop truncation is accurate to about one percent above \(10\,\mathrm{GeV}\) and drifts low by roughly a tenth of its value at the $\tau$ mass, where the measured $\alpha_{s}$ is about $0.31$ [Navas:2024]; this is the expected size of the $O(\alpha_{s}^{3})$ term of Equation (102.42).

The determinations that must agree come from unrelated physics, and listing them is the point.

Remark 102.42 (Why the running is the strongest test of QCD).

Any theory with one coupling can be made to fit one measurement of that coupling. What cannot be arranged is that five classes of measurement, at energies differing by a factor of five hundred and using different final states, different detectors and different theoretical machinery, agree after being evolved by Equation (102.44) with a coefficient \(b_{0}\) that is not adjustable — it is \(11-\tfrac{2}{3}n_{f}\), fixed by Theorem 102.35 from the group and the particle content alone. The agreement of the world data with one curve of predicted shape, rather than the value \(0.1180\) at one point, is the quantitative evidence for QCD. The general renormalization-group apparatus behind the evolution is the subject of The Renormalization Group; a systematic survey of the determinations is [Bethke:2009] [Navas:2024].

Deep inelastic scattering and the parton model

Bjorken scaling and partons

Everything so far has been inference from a spectrum. Deep inelastic scattering is different in kind: it is Rutherford's experiment performed on the proton, and what it sees is not a symmetry but an object.

Notation 102.43 (Kinematics of inelastic lepton scattering).

An electron of four-momentum \(k\) scatters off a proton of four-momentum \(P\) and mass \(M\), emerging with \(k'\) and leaving an unobserved hadronic system \(X\). Write \(q=k-k'\) for the momentum transferred by the virtual photon and

\begin{equation}\tag{102.48} Q^{2}:=-q^{2}\ec\qquad \nu:=\frac{P\cdot q}{M}\ec\qquad x:=\frac{Q^{2}}{2M\nu}=\frac{Q^{2}}{2P\cdot q}\ec\qquad y:=\frac{P\cdot q}{P\cdot k}\ep \end{equation}

By Notation 100.1 a four-momentum carries \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), so \(Q^{2}\) carries \(\mathrm{kg}^{2}\,\mathrm{m}^{2}/\mathrm{s}^{2}\) and \(\nu\), being \(P\cdot q/M\), is an energy: in the proton rest frame \(\nu=E-E'\), the energy lost by the electron. Both \(x\) and \(y\) are dimensionless, and \(0\leq x\leq 1\) because \(W^{2}c^{2}:=(P+q)^{2}c^{2}\geq M^{2}c^{4}\) for any hadronic final state. Elastic scattering is \(x=1\); deep inelastic scattering means \(Q^{2}c^{2}\gg M^{2}c^{4}\) at fixed \(x\).

Definition 102.44 (Structure functions).

Lorentz invariance, current conservation and parity restrict the inclusive cross-section, at one-photon exchange, to two independent functions of the two invariants:

\begin{equation}\tag{102.49} \frac{\dd^{2}\sigma}{\dd\Omega\,\dd E'} =\sigma_{\mathrm{Mott}} \left[W_{2}\left(x,Q^{2}\right) +2W_{1}\left(x,Q^{2}\right)\tan^{2}\frac{\vartheta}{2}\right]\ec \end{equation}

with the Mott cross-section for a structureless target

\begin{equation}\tag{102.50} \sigma_{\mathrm{Mott}} =\frac{\alpha^{2}\left(\hbar c\right)^{2} \cos^{2}\left(\vartheta/2\right)} {4E^{2}\sin^{4}\left(\vartheta/2\right)}\ec \end{equation}

of dimension \(\mathrm{m}^{2}\), \(\vartheta\) the electron scattering angle and \(E\) its incident energy. The dimensionless structure functions are

\begin{equation}\tag{102.51} F_{1}\left(x,Q^{2}\right):=Mc^{2}W_{1}\ec\qquad F_{2}\left(x,Q^{2}\right):=\nu W_{2}\ec \end{equation}

\(W_{1}\) and \(W_{2}\) each carrying \(/\mathrm{J}\).

Phenomenon 102.45 (The proton contains pointlike constituents).

In inelastic electron–proton scattering at large momentum transfer the cross-section falls far more slowly with \(Q^{2}\) than the elastic one does, and the structure functions depend, to a good first approximation, only on the dimensionless combination \(x=Q^{2}/2M\nu\) — with \(Q\) the momentum transferred, \(\nu\) the energy transferred and \(M\) the proton mass — and not on \(Q^{2}\) separately [Bloom:1969] [Breidenbach:1969], exactly as Bjorken had predicted from current algebra before the data existed [Bjorken:1969a]. The two structure functions are moreover not independent:

\begin{equation}\tag{102.52} F_{2}(x)=2xF_{1}(x)\ec \end{equation}

which the data obey [Navas:2024]. Scaling says the target contains constituents that carry no scale of their own; Equation (102.52) says those constituents carry spin \(\tfrac{1}{2}\). The experiment is Experiment: Deep Inelastic Scattering. Rests on Definition 102.44 and Equation (102.50).

Derivation. Derives Phenomenon 102.45. Scaling. Suppose the photon, at large \(Q^{2}\), resolves a time short compared with the time over which the proton's constituents interact with one another, so that it scatters elastically off one of them and the rest are spectators. This is the impulse approximation, and Feynman's name for the constituents is partons [Feynman:1969]. Let the struck parton carry a fraction \(\xi\) of the proton's four-momentum, \(p=\xi P\), and let it remain on its mass shell after the collision, its mass being \(\xi M\) in the same approximation. Then

\begin{equation}\tag{102.53} \left(\xi P+q\right)^{2}=\xi^{2}M^{2}c^{2} \quad\Longrightarrow\quad \xi^{2}M^{2}c^{2}+2\xi\,P\cdot q-Q^{2}=\xi^{2}M^{2}c^{2} \quad\Longrightarrow\quad \xi=\frac{Q^{2}}{2P\cdot q}=x\ep \end{equation}

So the measured variable \(x\) is the momentum fraction of the struck constituent: this is what makes \(x\) physical rather than a convenient abbreviation. The cross-section is then the incoherent sum over constituents of pointlike elastic cross-sections, each of which is the Mott formula Equation (102.50) times the squared charge, and the only place a scale could enter — a form factor, i.e. a size — is absent by hypothesis. Hence

\begin{equation}\tag{102.54} F_{2}(x)=\sum_{q}Q_{q}^{2}\, x\left[q(x)+\bar{q}(x)\right]\ec \end{equation}

with \(q(x)\dd x\) the number of quarks of flavour \(q\) carrying momentum fraction in \(\dd x\): a function of \(x\) alone. Any structure of the constituents themselves would reintroduce a length and destroy the scaling, so scaling is the statement that the constituents are pointlike down to the resolution \(\hbar c/Qc\) reached — about \(10^{-16}\,\mathrm{m}\) at the original SLAC energies, a tenth of the proton radius \(8.4075\times 10^{-16}\,\mathrm{m}\) [Mohr:2025]. The contrast is with elastic scattering, whose cross-section carries the squared dipole form factor and therefore falls as \(Q^{-12}\): the same beam sees a soft object elastically and hard objects inelastically, which is precisely Rutherford's situation.

The Callan–Gross relation. Decompose the virtual photon into its three polarization states in the frame where it moves along \(+z\) and the struck parton along \(-z\). A massless spin-\(\tfrac{1}{2}\) parton coupling through the vector current \(\bar{q}\gamma^{\mu}q\) conserves helicity, because \(\gamma^{\mu}\) commutes with \(\gamma^{5}\) up to a sign that leaves the chiral projectors intact and, for a massless fermion, chirality is helicity. After absorbing the photon the parton moves along \(+z\); its helicity is unchanged, so its spin projection on the \(z\) axis has flipped, \(\Delta J_{z}=\pm1\). Only a transverse photon, \(\lambda=\pm1\), can supply that. The longitudinal photon, \(\lambda=0\), therefore cannot be absorbed at all:

\begin{equation}\tag{102.55} \sigma_{L}=0\qquad\text{for spin-}\tfrac{1}{2}\text{ partons.} \end{equation}

Expressing the two cross-sections in terms of the structure functions, the transverse one is proportional to \(F_{1}\) and the longitudinal to the combination \(F_{L}:=F_{2}-2xF_{1}\) (up to terms of order \(M^{2}c^{4}/Q^{2}c^{2}\), negligible in the deep inelastic region), so Equation (102.55) is exactly Equation (102.52) [Callan:1969]. A spin-\(0\) parton gives the opposite conclusion by the same argument — it has no spin to flip, so only \(\lambda=0\) is absorbed, \(\sigma_{T}=0\) and \(F_{1}=0\) — and the data exclude it decisively: the measured ratio \(R_{L}=\sigma_{L}/\sigma_{T}\) is small, as QCD predicts; the experimental write-up belongs to Section 116.4.3.

Proposition 102.46 (The partons have fractional charge).

Comparing charged-lepton and neutrino deep inelastic scattering on an isoscalar target measures the mean square charge of the constituents directly:

\begin{equation}\tag{102.56} \frac{F_{2}^{eN}(x)}{F_{2}^{\nu N}(x)} =\frac{1}{2}\left[\left(\frac{2}{3}\right)^{2} +\left(\frac{1}{3}\right)^{2}\right]=\frac{5}{18}\ec \end{equation}

which the data confirm [Navas:2024]. Rests on Definition 102.5 and Equation (102.54).

Proof.

Derives Proposition 102.46. The photon couples to a quark with strength \(Q_{q}e\), so Equation (102.54) weights each flavour by \(Q_{q}^{2}\). The charged weak current of Weak Interactions couples with a strength independent of electric charge, so the corresponding neutrino structure function is \(F_{2}^{\nu N}(x)=x\sum_{q}[q(x)+\bar{q}(x)]\) with unit weights. On an isoscalar target the \(u\) and \(d\) distributions enter symmetrically, and the ratio of the two sums is the mean of \((2/3)^{2}\) and \((1/3)^{2}\), which is Equation (102.56). Had the constituents carried integer charges the ratio would have been of order unity instead of \(0.28\). This is a measurement of the charges themselves, independent of the spectroscopy of Section 102.1.3, and it is why neutrino scattering was decisive.

Phenomenon 102.47 (Half the proton's momentum is electrically neutral).

Integrating the measured \(F_{2}\) over \(x\) and dividing out the charge weights gives the total momentum fraction carried by charged constituents,

\begin{equation}\tag{102.57} \int_{0}^{1}\dd x\;x\sum_{q}\left[q(x)+\bar{q}(x)\right] \approx0.5\ec \end{equation}

so that about half the proton's momentum is carried by something that does not couple to the photon or to the weak current [Navas:2024]. The natural candidate is the gauge field itself. Rests on Equations (102.53), (102.54) and (102.56).

Derivation. Derives Phenomenon 102.47. The momentum of the proton is shared among its constituents by definition, \(\int_{0}^{1}\dd x\,x\sum_{\text{all}}f_{i}(x)=1\), since \(x\) is the momentum fraction (Equation (102.53)). The measured left-hand side of Equation (102.57) is obtained from Equation (102.54) together with Equation (102.56), the two providing enough information to separate the charge weights from the distributions; the details of the extraction belong in Section 116.5.2. The deficit is a fact about the data and not an inference: whatever carries the missing half is electrically neutral and weakly inert. In QCD it is the gluon, which appears here for the first time as an experimental necessity rather than a theoretical convenience — and it appeared this way, in the momentum sum rule, before it was seen in the three-jet events of Section 102.6.2.

Scaling violation and the DGLAP equations

Scaling is not exact, and QCD says exactly how it fails. A quark is not free even at short distance: it can radiate a gluon, and the probability of doing so is not small because the emission is logarithmically enhanced when the gluon is soft or collinear — the same enhancement as the soft photon of Section 100.6.1. Raising \(Q^{2}\) means resolving a smaller transverse distance, hence seeing emissions that were previously unresolved, hence finding more partons, each carrying less momentum. Structure functions therefore drift: down at large \(x\), up at small \(x\), logarithmically in \(Q^{2}\).

Proposition 102.48 (Collinear emission and its logarithm).

The probability that a quark of momentum fraction \(y\) radiates, leaving a quark of fraction \(x=zy\), is

\begin{equation}\tag{102.58} \dd\mathcal{P}=\frac{\alpha_{s}}{2\pi}\,P_{qq}(z)\, \dd z\,\frac{\dd k_{T}^{2}}{k_{T}^{2}}\ec\qquad P_{qq}(z)=C_{F}\left[\frac{1+z^{2}}{1-z}\right]_{+}\ec \end{equation}

with \(k_{T}\) the transverse momentum of the emitted gluon, \(C_{F}=4/3\) from Equation (102.31), and the plus prescription defined by \(\int_{0}^{1}\dd z\,[g(z)]_{+}f(z)=\int_{0}^{1}\dd z\,g(z) [f(z)-f(1)]\). Integrating \(\dd k_{T}^{2}/k_{T}^{2}\) up to the resolution \(Q^{2}\) produces \(\ln Q^{2}\), which is the entire origin of the scaling violation. Rests on Equation (102.31), Theorem 100.60 and Equation (100.76).

Proof.

Derives Proposition 102.48. Two features are derived here and the third is quoted. The factor \(\dd k_{T}^{2}/k_{T}^{2}\) is the collinear singularity of a massless emitter: the propagator of the quark after emission has virtuality \(\left(p-k\right)^{2}\propto -k_{T}^{2}/(1-z)\), and the squared amplitude carries its inverse square against a two-body phase space \(\propto\dd k_{T}^{2}\), leaving \(\dd k_{T}^{2}/k_{T}^{2}\). It is logarithmic precisely because the emitter is massless, and it is the mass singularity of Theorem 100.60 in its QCD form. The factor \(1/(1-z)\) is the soft singularity: as \(z\to1\) the gluon energy goes to zero and the amplitude reduces to the eikonal current Equation (100.76), whose square gives exactly this pole, with \(C_{F}\) replacing the squared electric charge because the emission vertex is \(T_{a}\) and \(T_{a}T_{a}=C_{F}\identity\). The numerator \(1+z^{2}\) is the spin structure of the vertex, the two terms being the helicity-conserving and helicity-flip configurations of the emitting quark; it is computed from the same trace technology as Section 100.3.2 [Altarelli:1977] [Gribov:1972].

The plus prescription is not an extra assumption but a normalization forced by conservation of quark number. The number of quarks of a given flavour minus antiquarks is a conserved charge and cannot change with \(Q^{2}\), so \(\int_{0}^{1}\dd z\,P_{qq}(z)=0\); the unregulated \(1/(1-z)\) is not integrable, and the divergence is cancelled by the virtual correction, which lives entirely at \(z=1\). The plus prescription is exactly that cancellation written as a distribution: it subtracts, at \(z=1\), whatever the real emission supplies. Equivalently \([(1+z^{2})/(1-z)]_{+}=(1+z^{2})\left[1/(1-z)\right]_{+} +\tfrac{3}{2}\delta(1-z)\), and the \(\tfrac{3}{2}\) is the virtual term.

Theorem 102.49 (The DGLAP evolution equations).

The parton distributions depend on the scale at which they are defined, and the dependence is

\begin{equation}\tag{102.59} \frac{\pp q_{i}(x,Q^{2})}{\pp\ln Q^{2}} =\frac{\alpha_{s}(Q^{2})}{2\pi}\int_{x}^{1}\frac{\dd y}{y} \left[P_{qq}\!\left(\frac{x}{y}\right)q_{i}(y,Q^{2}) +P_{qg}\!\left(\frac{x}{y}\right)g(y,Q^{2})\right]\ec \end{equation}

together with the companion equation for the gluon distribution involving \(P_{gq}\) and \(P_{gg}\) [Gribov:1972] [Altarelli:1977] [Dokshitzer:1977]. The right-hand side is a multiplicative convolution on \((0,1]\) in the sense of Definition 17.97. Rests on Proposition 102.48, Equation (102.58) and Definition 17.97.

Proof.

Derives Theorem 102.49. A parton of fraction \(x\) at resolution \(Q^{2}+\dd Q^{2}\) is either one that was already there at \(Q^{2}\), or one produced by an emission from a parton of larger fraction \(y>x\) that the coarser resolution could not separate. Multiplying Equation (102.58) by the density of emitters \(q_{i}(y)\), changing variables from \(z\) to \(x=zy\) at fixed \(y\) so that \(\dd z=\dd x/y\), and integrating \(\dd k_{T}^{2}/k_{T}^{2}\) over the shell between the two resolutions gives the first term of Equation (102.59). The second term is the same statement for a gluon converting to a quark–antiquark pair, with \(P_{qg}(z)=T_{F}\left[z^{2}+(1-z)^{2}\right]\). That the kernel depends on \(x\) and \(y\) only through \(x/y\) is a consequence of the collinear factorization: the emission probability knows only the momentum fraction taken, not the absolute momenta. Comparing with Equation (102.63) below identifies the integral as \(\left(P\star q\right)(x)\) with the Haar measure \(\dd y/y\) of Definition 17.97, which is the structure that Corollary 102.50 exploits.

Corollary 102.50 (Moments diagonalise the evolution).

Define the moments of a distribution as in Equation (17.114), \(M_{N}[f]=\int_{0}^{1}f(x)x^{N}\dd x\), so that \(N=0\) counts partons and \(N=1\) is the momentum fraction. For a non-singlet combination — one in which the gluon term of Equation (102.59) cancels, for instance \(u-d\) or \(q-\bar{q}\) — the evolution becomes an ordinary differential equation for each moment separately,

\begin{equation}\tag{102.60} \frac{\dd M_{N}}{\dd\ln Q^{2}} =\frac{\alpha_{s}(Q^{2})}{2\pi}\,\gamma_{N}\,M_{N}\ec\qquad \gamma_{N}:=\int_{0}^{1}\dd z\;z^{N}P_{qq}(z)\ec \end{equation}

whose solution, using Equation (102.42), is a pure power of the coupling:

\begin{equation}\tag{102.61} \frac{M_{N}\left(Q^{2}\right)}{M_{N}\left(Q_{0}^{2}\right)} =\left[\frac{\alpha_{s}(Q^{2})}{\alpha_{s}(Q_{0}^{2})} \right]^{-2\gamma_{N}/b_{0}}\ep \end{equation}

The anomalous dimensions are

\begin{equation}\tag{102.62} \gamma_{N}=C_{F}\left[\frac{3}{2}-H_{N}-H_{N+2}\right]\ec\qquad H_{m}:=\sum_{k=1}^{m}\frac{1}{k}\ec \end{equation}

with \(\gamma_{0}=0\) and \(\gamma_{1}=-\tfrac{4}{3}C_{F}=-\tfrac{16}{9}\). Rests on Theorem 102.49, Corollary 17.99 and Equation (102.42).

Proof.

Derives Corollary 102.50. Write the convolution of Theorem 102.49 in the form

\begin{equation}\tag{102.63} \left(P\star q\right)(x)=\int_{x}^{1}P(y)\, q\!\left(\frac{x}{y}\right)\frac{\dd y}{y}\ec \end{equation}

which is Equation (17.113): both factors vanish for argument greater than one, so the range is \(x\leq y\leq1\) and the convolution stays supported on the unit interval. By Corollary 17.99, and ultimately by the Mellin convolution theorem Theorem 17.98, the moments of a multiplicative convolution are the products of the moments, Equation (17.115). Taking \(M_{N}\) of Equation (102.59) therefore turns the integral operator into multiplication by \(\gamma_{N}=M_{N}[P_{qq}]\), giving Equation (102.60). This is the whole reason the Mellin transform is the right instrument here: it diagonalises dilations, and Equation (102.59) is an equation about how momentum fractions dilate.

To solve, divide Equation (102.60) by Equation (102.42) rewritten as \(\dd\alpha_{s}/\dd\ln Q^{2}=-b_{0}\alpha_{s}^{2}/4\pi\) — the factor \(\tfrac{1}{2}\) relative to Equation (102.42) because \(\ln Q^{2}=2\ln Q\). Then \(\dd\ln M_{N}/\dd\alpha_{s}=-2\gamma_{N}/(b_{0}\alpha_{s})\), which integrates to \(M_{N}\propto\alpha_{s}^{-2\gamma_{N}/b_{0}}\) and hence to Equation (102.61).

For Equation (102.62), use the plus prescription:

\[ \gamma_{N}=C_{F}\int_{0}^{1}\dd z\, \frac{\left(z^{N}-1\right)\left(1+z^{2}\right)}{1-z}\ec \]

the subtraction at \(z=1\) making the integrand finite. Split the numerator: \(\left(z^{N}-1\right)+\left(z^{N+2}-z^{2}\right)\), and write the second bracket as \(\left(z^{N+2}-1\right)-\left(z^{2}-1\right)\). With \(\int_{0}^{1}\left(z^{m}-1\right)/(1-z)\,\dd z=-H_{m}\) — immediate from \(\left(1-z^{m}\right)/(1-z)=1+z+\dots+z^{m-1}\) integrated term by term — this gives \(\gamma_{N}=C_{F}\left[-H_{N}-H_{N+2}+H_{2}\right]\) and \(H_{2}=\tfrac{3}{2}\), which is Equation (102.62). Checking: \(\gamma_{0} =C_{F}[\tfrac{3}{2}-0-\tfrac{3}{2}]=0\), as quark-number conservation demands, and \(\gamma_{1}=C_{F}[\tfrac{3}{2}-1-\tfrac{11}{6}] =-\tfrac{4}{3}C_{F}\).

Example 102.51 (How fast a moment moves).

Take the momentum moment \(N=1\) of a non-singlet distribution and evolve from \(Qc=10\,\mathrm{GeV}\) to \(Qc=100\,\mathrm{GeV}\) with \(n_{f}=5\), so \(b_{0}=23/3\) and the exponent in Equation (102.61) is \(-2\gamma_{1}/b_{0}=\left(32/9\right)\left(3/23\right)=0.464\). From Table 102.2, \(\alpha_{s}\) falls from \(0.173\) to \(0.116\) over that range, a ratio \(0.673\), so

\begin{equation}\tag{102.64} \frac{M_{1}\left(Q^{2}\right)}{M_{1}\left(Q_{0}^{2}\right)} =0.673^{0.464}=0.83\ep \end{equation}

The momentum carried by the valence quarks falls by about a sixth over one decade in \(Q\): a large, slow, logarithmic effect, exactly what distinguishes a scaling violation from a scaling failure.

Remark 102.52 (The prediction is a shape, not a number).

QCD does not predict \(q(x,Q_{0}^{2})\): the distributions at one scale are non-perturbative and must be measured. What it predicts is their entire \(Q^{2}\) dependence, with no free parameter beyond \(\alpha_{s}\). Equation (102.61) makes the point sharply: the ratio of two moments' evolution,

\begin{equation}\tag{102.65} \frac{\ln\left[M_{N}(Q^{2})/M_{N}(Q_{0}^{2})\right]} {\ln\left[M_{M}(Q^{2})/M_{M}(Q_{0}^{2})\right]} =\frac{\gamma_{N}}{\gamma_{M}}\ec \end{equation}

contains neither \(\alpha_{s}\) nor \(\Lambda_{\mathrm{QCD}}\) nor the initial distributions — only the two rational numbers of Equation (102.62). Plotting one moment against another on logarithmic axes must give a straight line of predicted slope, and it does. Global fits to data spanning four orders of magnitude in \(Q^{2}\) and five in \(x\) are described by Equation (102.59) with one coupling; this is the most demanding quantitative test QCD has passed, and its experimental side belongs to Section 116.6.4.

Factorization and parton distributions

Equation (102.54) and Equation (102.59) together contain a claim that goes far beyond deep inelastic scattering: that a hadronic cross-section splits into a part depending on the hard process, which is calculable, and a part depending on the hadron, which is not but is universal.

Theorem 102.53 (Collinear factorization).

For an inclusive process with a large scale \(Q\), the cross-section can be written to leading power in \(\Lambda_{\mathrm{QCD}}/Q\) as

\begin{equation}\tag{102.66} \sigma=\sum_{i,j}\int\dd x_{1}\dd x_{2}\; f_{i}\!\left(x_{1},\mu_{F}^{2}\right) f_{j}\!\left(x_{2},\mu_{F}^{2}\right)\, \hat{\sigma}_{ij}\!\left(x_{1},x_{2},Q^{2},\mu_{F}^{2}, \alpha_{s}\right)\ec \end{equation}

where \(\hat{\sigma}_{ij}\) is a partonic cross-section computable as a series in \(\alpha_{s}\) and the \(f_{i}\) are parton distributions that depend on the hadron but not on the process. The arbitrary scale \(\mu_{F}\) separating the two is a factorization scale; its \(\mu_{F}\)-dependence cancels between the two factors order by order, and \(\pp f_{i}/\pp\ln\mu_{F}^{2}\) is Equation (102.59). Rests on Equations (102.54) and (102.59).

Derivation pending.

Collinear factorization to all orders. What is available inline is the leading-order statement, which is the impulse approximation together with the observation that the collinear logarithms are universal because they come from the region where the emitted parton is nearly on shell and therefore long-lived compared with the hard collision. The all-orders proof requires showing that the leading regions of the loop momenta are exhausted by hard, collinear and soft subgraphs, that the soft gluons cancel or are absorbed into eikonal lines, and that what remains is the same operator matrix element in every process. That analysis, and the counterexamples at subleading power, belong in Appendix A.

The physical content is that the distributions may be measured in one process and used in another. They are extracted from deep inelastic scattering on fixed targets and at HERA, and then used — with the evolution Equation (102.59) carrying them from \(Q\sim10\,\mathrm{GeV}/c\) to \(Q\sim1\,\mathrm{TeV}/c\) — to predict cross-sections at hadron colliders. Every Standard Model rate measured at the LHC, the Higgs production of Experiment: The Higgs Boson Discovery included, is computed with Equation (102.66) and would be unpredictable without it. That the same \(f_{i}(x,\mu_{F}^{2})\) describe electron–proton scattering at \(10\,\mathrm{GeV}\) and proton–proton collisions at \(13\,\mathrm{TeV}\) is a test of universality across three decades in scale and two different initial states, and it is the reason Theorem 102.53 is stated as a theorem rather than as a modelling assumption [Collins:1989].

Remark 102.54 (What the factorization scale is not).

\(\mu_{F}\) is not a physical boundary between “inside” and “outside” the proton. It is a subtraction point, exactly like the renormalization scale of Definition 100.54: emissions harder than \(\mu_{F}\) are put into \(\hat{\sigma}\) and those softer into \(f_{i}\), and the split is a bookkeeping choice. A complete calculation is \(\mu_{F}\)-independent; a calculation truncated at order \(\alpha_{s}^{n}\) has a residual dependence of order \(\alpha_{s}^{n+1}\), and varying \(\mu_{F}\) by a factor of two is the standard — and admittedly crude — estimate of the missing higher orders. A parton distribution is therefore not a probability density in any scheme-independent sense; it is a scheme-dependent object, and quoting one without its scheme and scale is meaningless.

Jets and the gluon

Two-jet events and the quark spin

Deep inelastic scattering sees the constituents by their recoil. Electron–positron annihilation makes a pair of them directly, and what the detector records is the closest thing to a photograph of a quark that confinement permits.

Proposition 102.55 (Angular distribution of a produced fermion pair).

For \(e^{+}e^{-}\to f\bar{f}\) through one virtual photon, with all masses negligible,

\begin{equation}\tag{102.67} \frac{\dd\sigma}{\dd\Omega} =\frac{\alpha^{2}\left(\hbar c\right)^{2}} {4E_{\mathrm{cm}}^{2}}\,Q_{f}^{2}\, N_{c}^{(f)}\left(1+\cos^{2}\vartheta\right)\ec \end{equation}

\(\vartheta\) being the angle between the outgoing fermion and the beam. A produced spin-\(0\) pair would give \(\sin^{2}\vartheta\) instead. Rests on Equation (100.40), Equation (100.32) and Definition 102.10.

Proof.

Derives Proposition 102.55. Section 100.3.1 gives the spin-averaged squared amplitude for this process as \(\overline{\abs{\mathcal{M}}^{2}} =32\pi^{2}\alpha^{2}\hbar^{4}\left(t^{2}+u^{2}\right)/s^{2}\). For massless kinematics in the centre-of-mass frame, \(t=-\tfrac{1}{2}s\left(1-\cos\vartheta\right)\) and \(u=-\tfrac{1}{2}s\left(1+\cos\vartheta\right)\), so

\[ \frac{t^{2}+u^{2}}{s^{2}} =\frac{1}{4}\left[\left(1-\cos\vartheta\right)^{2} +\left(1+\cos\vartheta\right)^{2}\right] =\frac{1+\cos^{2}\vartheta}{2}\ep \]

Inserting this into the two-body phase space of Equation (100.40) and multiplying by \(Q_{f}^{2}\) for the charge and by \(N_{c}\) for the colour multiplicity gives Equation (102.67); integrating over angles returns Equation (100.32) times the same factors, which is Equation (102.15). The \(\sin^{2}\vartheta\) for scalars follows because a scalar pair produced by a vector current must be in a \(p\)-wave with \(J_{z}=0\) along the pair axis, and the corresponding Wigner function is \(\sin\vartheta\).

Phenomenon 102.56 (Hadrons emerge as jets, and some events have three).

High-energy \(e^{+}e^{-}\) annihilation does not produce hadrons isotropically. Most events consist of two back-to-back collimated sprays whose common axis follows the \(1+\cos^{2}\vartheta\) distribution characteristic of a produced spin-\(\tfrac{1}{2}\) pair [Hanson:1975], and a calculable minority are planar events with three distinct sprays [Brandelik:1979]. The energy sharing and the angular correlations of the third spray are those of radiation emitted by a coloured source into a massless spin-one quantum. This is the only direct evidence for the gluon, and it is direct only in the sense that a jet, not a parton, is what a detector records. Rests on Equations (102.58) and (102.67).

Derivation of the two-jet structure. Derives Phenomenon 102.56. Two facts make a jet. First, the parton pair is produced back to back with energy \(E_{\mathrm{cm}}/2\) each, and the angular distribution of the pair axis is Equation (102.67) — an observable that survives hadronization because the subsequent radiation is soft. Second, the radiation is soft and collinear: by Equation (102.58) the emission probability is concentrated at small \(k_{T}\), so the hadrons that materialize share the parton's direction to within an angle of order \(\Lambda_{\mathrm{QCD}}/E\). The transverse momentum inside a jet is therefore bounded by a hadronic scale while its longitudinal momentum grows with the beam energy, and the spray narrows as \(1/E\). Hanson and collaborators established this at SPEAR by measuring the sphericity tensor and showing that its principal axis is distributed as \(1+\cos^{2}\vartheta\) [Hanson:1975]: the axis of a spray of hadrons remembers the spin of a particle that was never observed.

Remark 102.57 (What a jet has to be to be calculable).

A jet is defined by an algorithm, and not every algorithm defines a calculable quantity. Theorem 100.58 and Theorem 100.60 say that infrared and collinear divergences cancel between real and virtual corrections only for observables that are inclusive over soft emission and over collinear splitting. An observable is called infrared and collinear safe if it is unchanged by adding a zero-energy particle and by replacing any particle with two collinear ones carrying its total momentum [Sterman:1977]. The jet algorithms in use are constructed to have that property, and quantities computed from algorithms that lack it are not merely imprecise but divergent order by order. This is the practical content of Theorem 100.60 in QCD, and it is why the comparison of Phenomenon 102.56 with theory compares jet rates and never parton rates.

Three-jet events: the gluon

Proposition 102.58 (The three-jet energy distribution).

Let \(x_{i}=2E_{i}/E_{\mathrm{cm}}\) be the scaled energies of the quark, antiquark and gluon, so that \(x_{1}+x_{2}+x_{3}=2\). Then to first order in \(\alpha_{s}\)

\begin{equation}\tag{102.68} \frac{1}{\sigma_{0}}\, \frac{\dd^{2}\sigma}{\dd x_{1}\,\dd x_{2}} =\frac{\alpha_{s}}{2\pi}\,C_{F}\, \frac{x_{1}^{2}+x_{2}^{2}} {\left(1-x_{1}\right)\left(1-x_{2}\right)} =\frac{2\alpha_{s}}{3\pi}\, \frac{x_{1}^{2}+x_{2}^{2}} {\left(1-x_{1}\right)\left(1-x_{2}\right)}\ec \end{equation}

with \(\sigma_{0}\) the two-jet cross-section Equation (102.67) integrated over angles [Ellis:1981]. Had the emitted quantum been a scalar rather than a vector, the numerator would have been \(x_{3}^{2}\). Rests on Proposition 102.48, Equation (100.76) and Equation (102.31).

Derivation of the singular structure. Derives Proposition 102.58. The two limits that dominate the rate are derivable from results already established, and they already distinguish a vector from a scalar.

Soft. As \(x_{3}\to0\) both \(x_{1}\) and \(x_{2}\) tend to \(1\), and the amplitude factorizes into the two-jet amplitude times the eikonal current Equation (100.76) of Section 100.6.1, with the electric charge replaced by \(g_{s}T_{a}\) and the colour factor \(T_{a}T_{a}=C_{F}\) of Equation (102.31). Squaring gives the double pole \(\left[(1-x_{1})(1-x_{2})\right]^{-1}\) of Equation (102.68) with numerator \(2\), which is \(x_{1}^{2}+x_{2}^{2}\) at \(x_{1}=x_{2}=1\). A scalar emitter has no such factorization: a soft scalar does not couple to the eikonal current at all, because the coupling to a fast fermion line is suppressed by the emitted energy, and \(x_{3}^{2}\to0\) in the same limit. The observed events pile up at the soft edge, and that alone excludes a scalar.

Collinear. As \(x_{1}\to1\) at fixed \(x_{2}\) the gluon becomes collinear with the antiquark, and Equation (102.68) reduces to \(\left(\alpha_{s}/2\pi\right)C_{F} \left(1+x_{2}^{2}\right)/\left[(1-x_{1})(1-x_{2})\right]\), whose \(x_{2}\) dependence is the splitting function Equation (102.58) with \(z=x_{2}\) — as it must be, since Proposition 102.48 and Equation (102.68) are two views of the same emission. That the same function \(C_{F}(1+z^{2}) /(1-z)\) governs scaling violation in Section 102.5.2 and the three-jet rate here is a consistency requirement of the theory and is tested by the agreement of the two determinations of \(\alpha_{s}\) in Section 102.4.2.

Derivation pending.

The exact order-\(\alpha_{s}\) three-jet distribution. What is derived inline is the soft and collinear structure, which is where the rate is concentrated and which already excludes a scalar gluon. The full matrix element requires the spin-summed trace over the two interfering diagrams in which the gluon is emitted from the quark and from the antiquark, together with the three-body phase space in the Dalitz variables; and the corresponding statement for a scalar emitter, needed to make the exclusion quantitative rather than qualitative. Both are trace computations of the kind carried out for the electromagnetic processes and belong beside them in Appendix A.

The experimental history is short and decisive. Planar three-jet events were reported by TASSO at PETRA in 1979 [Brandelik:1979] and confirmed within months by the other three experiments at the same machine. The interpretation rests on more than the existence of a third spray: the distribution of energy between the three jets follows Equation (102.68), and the orientation of the three-jet plane relative to the beam — and of the jets within it — is that of a vector quantum and not a scalar. The measurement that established this is the distribution in the Ellis–Karliner angle, the angle between the two most energetic jets in the rest frame of the two least energetic, which is sensitive to the spin of the radiated quantum and insensitive to almost everything else.

Remark 102.59 (Testing the non-abelian structure itself).

The existence of a spin-one mediator is not yet the existence of a non-abelian one. What distinguishes QCD from an abelian theory of coloured quarks and a neutral gluon is the triple-gluon vertex of Equation (102.28), and that vertex first contributes to a four-jet final state, where a gluon splits into two gluons. The angular correlations among four jets — notably the angle between the planes defined by the two hardest and the two softest jets — depend on the ratio \(C_{A}/C_{F}\), which the group theory of Equation (102.31) fixes at \(9/4\). The LEP experiments measured \(C_{A}\) and \(C_{F}\) from four-jet rates and correlations and found values consistent with \(3\) and \(4/3\) at the ten-percent level [Navas:2024], with the abelian alternative (\(C_{A}=0\)) excluded by many standard deviations. This is the closest thing to a direct measurement of the gluon self-coupling, and therefore of the mechanism that Theorem 102.35 says produces asymptotic freedom.

Confinement

Lattice gauge theory

Everything in Sections 102.4 and 102.5 was perturbative and therefore restricted to short distance. Confinement is a long-distance statement, and the only systematic method for it is Wilson's [Wilson:1974]: define the theory on a Euclidean lattice, where the functional integral is a finite-dimensional ordinary integral, and evaluate it. The construction is carried out in Section 106.5.1 and is not repeated here; what is repeated is what it buys and what it does not.

Three of its results are used throughout this chapter. The link variables Equation (106.79) put the gauge field in the group rather than the algebra, so gauge invariance is exact at finite spacing and there is no gauge fixing and no ghost (Proposition 106.67). The Wilson action Equation (106.81) reduces to Equation (102.27) as the spacing goes to zero (Theorem 106.68), with lattice coupling \(\beta=6/\hat{g}^{2}=6/(4\pi\alpha_{s})\) for \(\SU(3)\). And the Wilson loop Definition 106.70 measures the static potential, an area law being equivalent to a linearly rising \(V(R)\) (Proposition 106.71).

Theorem 102.60 (Confinement at strong coupling, and what it proves).

For \(\beta\ll1\) the Wilson loop obeys an area law with string tension \(\sigma=\left(\hbar c/a^{2}\right)\ln\left(2N^{2}/\beta\right)\), which is Equation (106.84). This is an exact statement about the lattice theory at strong bare coupling. It is not a proof of confinement in QCD. Rests on Theorem 106.72, Equation (106.84) and Theorem 100.25.

Proof.

Derives Theorem 102.60. The area law itself is Theorem 106.72. The second sentence is the point at issue and is established by a counterexample. Compact \(\U(1)\) lattice gauge theory — lattice electrodynamics — obeys the same strong-coupling area law, by the same plaquette-tiling argument, since that argument uses only the group integration formula \(\int\dd U\,U=0\) and never the non-abelian structure. But continuum electrodynamics manifestly does not confine: the Coulomb potential of Theorem 100.25 is measured every day. The resolution is that compact \(\U(1)\) in four dimensions has a genuine phase transition at intermediate \(\beta\) separating a confining strong-coupling phase from a Coulomb weak-coupling phase, and the continuum limit is taken in the latter. Hence a strong-coupling area law establishes nothing about the continuum theory unless one also knows that no such transition intervenes between \(\beta\ll1\) and the continuum limit \(\beta\to\infty\) demanded by Remark 106.69. For \(\SU(3)\) the absence of a transition is a numerical finding, not a theorem.

The first such demonstration is Creutz's, for the gauge group \(\SU(2)\) [Creutz:1980b]. Computing the string tension by Monte Carlo across the region where the strong-coupling expansion fails, he found no discontinuity, and found in addition that \(\sigma a^{2}\) falls with \(\beta\) at the rate that asymptotic freedom demands for that group — asymptotic scaling: the lattice spacing shrinks with the bare coupling exactly as the running of the coupling requires, so that \(\sigma\) expressed in physical units is \(\beta\)-independent. That is the statement that the lattice theory and the asymptotically free continuum theory are the same theory. It is a numerical demonstration to finite accuracy. For \(\SU(3)\) the same continuity and scaling have since been established on the lattice, most sharply by the spectroscopy computations of [Duerr:2008], and that is what the honest claim in Section 102.7.3 rests on.

Remark 102.61 (Lattice QCD as a source of numbers).

Since Creutz the method has become quantitative. With improved actions, physical quark masses, several lattice spacings extrapolated to zero and several volumes extrapolated to infinity, lattice QCD computes: the light hadron spectrum to about one percent [Duerr:2008]; the quark masses of Table 102.1, which are defined by the lattice or by sum rules since no quark is ever isolated; \(\alpha_{s}\), currently the most precise determination entering Equation (102.47); decay constants and form factors needed to extract the quark mixing parameters of Flavour Physics and Neutrinos; the equation of state of hot matter (Section 102.9); and the hadronic vacuum polarization entering the muon anomaly [Borsanyi:2021]. What it does not do is prove anything: it is a controlled numerical evaluation of a functional integral with quantified statistical and systematic uncertainties, which is a different epistemic object from a theorem and is treated as such throughout this chapter. Its relationship to the statistical mechanics of Statistical Mechanics is exact — the Euclidean weight \(\ee^{-S_{E}/\hbar}\) is a Boltzmann factor — and the continuum limit is a critical point in the sense of Phase Transitions and Critical Phenomena.

The static potential and hadron masses

Phenomenon 102.62 (The static potential rises linearly).

The energy of a static colour source and its anticolour separated by \(r\) is, over the range \(0.2\times 10^{-15}\,\mathrm{m}\lesssim r\lesssim 10^{-15}\,\mathrm{m}\) accessible both to lattice computation and to quarkonium spectroscopy,

\begin{equation}\tag{102.69} V(r)=-\frac{C_{F}\,\alpha_{s}\hbar c}{r}+\sigma r\ec\qquad \sigma\approx0.9\,\mathrm{GeV}/\mathrm{fm} =1.4\times 10^{5}\,\mathrm{N}\ec \end{equation}

the first term being the short-distance Coulomb-like exchange with the colour factor \(C_{F}=4/3\) of Equation (102.31) and the second the confining term. The string tension \(\sigma\) is a force, and it is about fourteen tonnes weight. Rests on Equation (102.31), Equation (14.77) and Proposition 106.71.

Derivation of the two terms and a numerical test. Derives Phenomenon 102.62. The Coulomb term is one-gluon exchange between a colour source in \(\vect{3}\) and one in \(\bar{\vect{3}}\) coupled to a singlet. The colour factor is \(\avg{T_{a}\otimes T_{a}}_{\vect{1}} =\tfrac{1}{2}\left[C_{2}(\vect{1})-2C_{2}(\vect{3})\right] =-C_{F}\), using \(C_{2}(\vect{1})=0\) and Equation (14.77); the sign is attraction, and the magnitude is \(4/3\) times the electromagnetic case with \(\alpha\to\alpha_{s}\). Note that the same computation for a \(\vect{3}\otimes\vect{3}\) pair coupled to \(\bar{\vect{3}}\) gives \(-C_{F}/2\), attractive but weaker, and for the symmetric \(\vect{6}\) gives \(+C_{F}/4\), repulsive: colour singlets are the most tightly bound configurations available, which is Postulate 102.11 made plausible.

The linear term is not derived here — it is measured on the lattice and inferred from the spectrum — but it has a check that costs one page and is worth doing, because it connects Equation (102.69) to a completely independent body of data. Model the confining flux as a relativistic string of tension \(\sigma\) with massless ends rotating rigidly at the speed of light. If the string has total length \(2R\), the transverse speed at distance \(r\) from the centre is \(v=cr/R\), and

\begin{align} E&=2\int_{0}^{R}\frac{\sigma\,\dd r}{\sqrt{1-r^{2}/R^{2}}} =\pi\sigma R\ec\nn\\ J&=\frac{2}{c^{2}}\int_{0}^{R} \frac{\sigma\,v\,r\,\dd r}{\sqrt{1-r^{2}/R^{2}}} =\frac{2\sigma}{cR}\int_{0}^{R} \frac{r^{2}\dd r}{\sqrt{1-r^{2}/R^{2}}} =\frac{\pi\sigma R^{2}}{2c}\ec \tag{102.70} \end{align}

using \(\int_{0}^{1}u^{2}(1-u^{2})^{-1/2}\dd u=\pi/4\). Eliminating \(R\) and writing \(E=Mc^{2}\) and \(J=\hbar j\),

\begin{equation}\tag{102.71} j=\frac{M^{2}c^{3}}{2\pi\sigma\hbar}\ec \end{equation}

a linear Regge trajectory: spin proportional to squared mass, with a slope fixed by the string tension alone. The name records Regge's poles in complex angular momentum [Regge:1959]; that the observed hadronic trajectories are in fact linear is an empirical statement about the spectrum, tested on the \(\rho\) data below [Navas:2024]. Numerically, \(\sigma=1.442\times 10^{5}\,\mathrm{N}\) gives \(j=0.896\) for \(Mc^{2}=1\,\mathrm{GeV}\): by Equation (102.71) the slope of \(j\) in the squared rest energy \((Mc^{2})^{2}\) is \(1/(2\pi\sigma\hbar c)=3.49\times 10^{19}\,/\mathrm{J}^{2}\), that is \(0.896\,/\mathrm{GeV}^{2}\) in the customary units. The measured \(\rho\) trajectory — \(\rho(775)\) with \(j=1\), \(a_{2}(1320)\) with \(j=2\), \(\rho_{3}(1690)\) with \(j=3\) [Navas:2024] — has slope \(0.882\,/\mathrm{GeV}^{2}\) between the first two states and \(0.895\,/\mathrm{GeV}^{2}\) between the second and third. Agreement to about one percent, from a string tension measured in an entirely different way, and with no adjustable parameter: the linearity of Equation (102.69) and the linearity of the observed meson trajectories are the same fact.

Remark 102.63 (Quarkonium as the short-distance probe).

Charmonium and bottomonium are non-relativistic bound states of heavy quarks and therefore sample Equation (102.69) at small \(r\), where the Coulomb term dominates. Their spectra test the potential directly. The signature fact is that the spacing between the ground state and the first radial excitation is nearly the same in the two systems — \(0.589\,\mathrm{GeV}/c^{2}\) for \(\psi(2S)-J/\psi(1S)\) and \(0.563\,\mathrm{GeV}/c^{2}\) for \(\Upsilon(2S)-\Upsilon(1S)\) [Navas:2024], with \(m_{J/\psi}c^{2}=3.0969\,\mathrm{GeV}\) and \(m_{\Upsilon}c^{2}=9.4604\,\mathrm{GeV}\) — although the reduced masses differ by a factor of three. A pure Coulomb potential would give spacings scaling as the reduced mass; a pure linear potential would give them scaling as \(m^{-1/3}\); the observed mass-independence is what an intermediate, roughly logarithmic, potential produces, and Equation (102.69) is exactly such a potential over the relevant range.

Phenomenon 102.64 (Almost none of the proton's mass comes from quark masses).

The proton mass is \(938.272\,\mathrm{MeV}/c^{2}\), while the current masses of its three valence quarks sum to less than \(10\,\mathrm{MeV}/c^{2}\) [Navas:2024] — about one percent of the total. The remainder is the energy of the gluon field and of the confined motion of the quarks. That this is not an evasion is shown by computation: lattice evaluations of the light hadron spectrum, taking as input only the quark masses and the coupling, reproduce the measured masses at the percent level [Duerr:2008]. Rests on Phenomenon 102.39, Equation (102.29) and Equation (106.81).

Derivation. Derives Phenomenon 102.64. The arithmetic is immediate from Table 102.1: \(2m_{u}+m_{d}=9.02\,\mathrm{MeV}/c^{2}\) against \(938.272\,\mathrm{MeV}/c^{2}\), so \(99\,\mathrm{\%}\) of the mass is not quark mass. What has to be explained is not the deficit but the presence of a scale at all, and that is Phenomenon 102.39: in the limit \(m_{u},m_{d}\to0\) the Lagrangian Equation (102.29) contains no dimensionful parameter, yet the quantum theory generates \(\Lambda_{\mathrm{QCD}}\) through Equation (102.45), and every mass in the spectrum is a pure number times \(\Lambda_{\mathrm{QCD}}/c^{2}\). The proton mass is therefore predicted up to that one overall scale, which is fixed by measuring any single hadron mass; everything else is then a prediction with no freedom.

Turning that into numbers is a computation and not a derivation, and the distinction is worth keeping. The lattice programme [Duerr:2008] proceeds by: choosing the bare quark masses and \(\beta\); computing hadron correlators by Monte Carlo evaluation of Equation (106.81) with dynamical quarks; setting the scale by matching one measured quantity, so that \(a\) is known in metres; tuning the light quark masses by matching two more, typically \(m_{\pi}\) and \(m_{K}\); extrapolating to zero spacing, to infinite volume and, where necessary, to the physical light quark masses; and then predicting the remaining spectrum. Three inputs, and the masses of the nucleon, the \(\Delta\), the \(\Lambda\), the \(\Sigma\), the \(\Xi\) and the \(\Omega\) come out right to a few percent. That is the strongest single piece of evidence that Equation (102.29) is the correct theory of the strong interaction, and it is the reason this chapter can call confinement an observed regularity of a known Lagrangian rather than a mystery.

No free quark has ever been observed

Phenomenon 102.65 (No free quark has ever been observed).

No particle carrying a fractional electric charge has been detected in any experiment. Millikan-type searches for residual fractional charge in bulk matter set limits below \(10^{-20}\) fractionally charged particles per nucleon in the best samples [Halyo:2000], and accelerator and cosmic-ray searches are likewise null [Navas:2024]. The positive counterpart of this null result is that the potential energy between a static colour source and its anticolour grows linearly with their separation instead of falling off, so that pulling them apart creates a new pair rather than liberating a colour charge. Rests on Equations (102.12) and (102.69).

Derivation of the string-breaking mechanism. Derives Phenomenon 102.65. Take Equation (102.69) at face value and try to separate a quark from an antiquark. At separation \(r\) the stored energy is \(\sigma r\), and it grows without bound. But the vacuum is not empty: as soon as \(\sigma r\) exceeds the energy needed to create a light quark–antiquark pair, roughly twice the constituent mass Equation (102.12), it is energetically favourable for the flux tube to break, the new pair capping the two broken ends. The critical separation is

\begin{equation}\tag{102.72} r_{\mathrm{break}}\approx\frac{2m_{q}c^{2}}{\sigma} =\frac{2\times0.336\,\mathrm{GeV}} {0.9\,\mathrm{GeV}/\mathrm{fm}} \approx0.7\times 10^{-15}\,\mathrm{m}\ec \end{equation}

which is smaller than a hadron. So the attempt to isolate a quark produces, at a separation of less than one hadron radius, two hadrons instead: what would have been a free colour source is always screened by a newly created one before it can be separated. This is why the null results above are expected, and it is also why they cannot be strengthened into an observation of the mechanism: the mechanism predicts precisely that nothing anomalous is ever seen.

Derivation pending.

Confinement: no analytic derivation from the continuum Lagrangian exists, and none should be implied. What is available is Wilson's area law for the loop expectation value, exact in the strong-coupling lattice expansion, together with the numerical demonstration that the lattice theory joins continuously onto the asymptotically free continuum. Closing the gap between those two statements is an open problem, not an exercise.

Remark 102.66 (The honest status of confinement).

Three statements must be kept apart, and conflating them is the commonest overclaim in expositions of QCD. (i) No free quark has been observed. This is an experimental fact, Phenomenon 102.65. (ii) The lattice theory confines at strong coupling. This is a theorem, Theorem 102.60, and by the compact-\(\U(1)\) counterexample given in its proof it does not by itself say anything about the continuum theory. (iii) Continuum \(\SU(3)\) Yang–Mills theory has a mass gap and confines. This is neither observed directly nor proved. It is supported by numerical evidence of high quality — the asymptotic scaling first demonstrated for \(\SU(2)\) [Creutz:1980b] and the \(\SU(3)\) spectrum of [Duerr:2008] — and it is one of the recognized open problems of mathematical physics, recorded as such in What We Observe but Do Not Understand. This chapter asserts (i) and (ii) and reports (iii) as well supported and unproved.

Chiral symmetry and its breaking

The chiral limit

Table 102.1 contains a striking accident. The masses of the two lightest quarks, \(2.16\,\mathrm{MeV}\)\(/c^{2}\) and \(4.70\,\mathrm{MeV}\)\(/c^{2}\), are of the order of one percent of \(\Lambda_{\mathrm{QCD}}\), which is the only other scale in the theory. To that accuracy they may be set to zero, and the resulting theory has a symmetry that the real one nearly has.

Proposition 102.67 (The chiral symmetry of massless quarks).

Write \(q=q_{L}+q_{R}\) with the projectors \(P_{L,R}=\tfrac{1}{2}\left(\identity\mp\gamma^{5}\right)\). In Equation (102.29) the gauge coupling is diagonal in chirality and the mass term is not:

\begin{equation}\tag{102.73} \bar{q}\ii\hbar c\gamma^{\mu}D_{\mu}q =\bar{q}_{L}\ii\hbar c\gamma^{\mu}D_{\mu}q_{L} +\bar{q}_{R}\ii\hbar c\gamma^{\mu}D_{\mu}q_{R}\ec\qquad \bar{q}mc^{2}q=c^{2}\left(\bar{q}_{L}mq_{R} +\bar{q}_{R}mq_{L}\right)\ep \end{equation}

Hence for \(n_{f}\) massless flavours the Lagrangian is invariant under independent unitary rotations of the left- and right-handed fields,

\begin{equation}\tag{102.74} \SU(n_{f})_{L}\times\SU(n_{f})_{R}\times\U(1)_{V}\times\U(1)_{A}\ec \end{equation}

of which \(\U(1)_{V}\) is baryon number and \(\U(1)_{A}\) is destroyed by the axial anomaly of Theorem 106.87. Rests on Equation (102.29), Equation (102.21) and Theorem 106.87.

Proof.

Derives Proposition 102.67. \(\gamma^{5}\) anticommutes with every \(\gamma^{\mu}\), so \(P_{L}\gamma^{\mu}=\gamma^{\mu}P_{R}\) and \(\bar{q}_{L}\gamma^{\mu}q_{R} =q^{\dagger}P_{L}\gamma^{0}\gamma^{\mu}P_{R}q =q^{\dagger}\gamma^{0}P_{R}\gamma^{\mu}P_{R}q=0\) because \(P_{R}\gamma^{\mu}P_{R}=\gamma^{\mu}P_{L}P_{R}=0\); the cross terms in the kinetic piece therefore vanish, which is the first half of Equation (102.73). For the mass term the same algebra keeps only the cross terms, since \(\bar{q}_{L}q_{L}=0\). The covariant derivative Equation (102.21) contains no \(\gamma^{5}\) and acts only on colour, so it commutes with the flavour rotations \(q_{L}\to U_{L}q_{L}\), \(q_{R}\to U_{R}q_{R}\); with the mass matrix set to zero nothing else connects the two chiralities, giving Equation (102.74). Splitting \(\U(n_{f})=\SU(n_{f})\times\U(1)\) on each side and recombining the two \(\U(1)\)s into vector (\(U_{L}=U_{R}\)) and axial (\(U_{L}=U_{R}^{\dagger}\)) combinations names the last two factors. The vector one is the conservation of baryon number. The axial one is not a symmetry of the quantum theory at all: the measure of the functional integral is not invariant under it, which is Theorem 106.87, and the classical current has an anomalous divergence proportional to the gluon topological density.

Phenomenon 102.68 (The spectrum has no parity doublets).

If \(\SU(2)_{L}\times\SU(2)_{R}\) were realized in the ordinary way — as a symmetry of the states as well as of the Lagrangian — every hadron would have a partner of equal mass and opposite parity. It does not. The lightest vector meson is the \(\rho(775)\) with \(J^{PC}=1^{--}\) and its chiral partner the \(a_{1}(1260)\) with \(1^{++}\), some \(455\,\mathrm{MeV}\)\(/c^{2}\) heavier; the nucleon at \(939\,\mathrm{MeV}\)\(/c^{2}\) with \(J^{P}=\tfrac{1}{2}^{+}\) has as its partner the \(N(1535)\) with \(\tfrac{1}{2}^{-}\), some \(596\,\mathrm{MeV}\)\(/c^{2}\) heavier [Navas:2024]. Meanwhile the vector subgroup \(\SU(2)_{V}\) — ordinary isospin — is realized in the usual way, the proton and neutron differing by \(1.3\,\mathrm{MeV}\)\(/c^{2}\), one part in seven hundred [Navas:2024]. Rests on Proposition 102.67 and Equation (102.74).

Derivation. Derives Phenomenon 102.68. If a symmetry generator \(Q\) annihilates the vacuum then it commutes with the Hamiltonian and maps energy eigenstates to degenerate energy eigenstates. The axial generators \(Q_{5}^{a}\) change parity, because the axial current is a pseudovector; so \(Q_{5}^{a}\ket{h}\) would be a state of the same mass and opposite parity to \(\ket{h}\), for every hadron. The splittings quoted are not small corrections to such a degeneracy: they are of order \(\Lambda_{\mathrm{QCD}}\), which is the largest scale available, whereas the explicit breaking by the quark masses is of order \(7\,\mathrm{MeV}\). Symmetry breaking by the Lagrangian cannot produce an effect a hundred times its own size. Hence the axial generators do not annihilate the vacuum, while the vector ones evidently do — the isospin multiplets are degenerate to the accuracy at which electromagnetic and \(m_{d}-m_{u}\) effects enter. The symmetry is spontaneously broken, in the pattern

\begin{equation}\tag{102.75} \SU(2)_{L}\times\SU(2)_{R}\longrightarrow\SU(2)_{V}\ep \end{equation}

Dynamical symmetry breaking

Definition 102.69 (The quark condensate).

The order parameter of Equation (102.75) is the vacuum expectation value

\begin{equation}\tag{102.76} \avg{\bar{q}q}=\avg{\bar{q}_{L}q_{R}+\bar{q}_{R}q_{L}}\neq0\ec \end{equation}

of SI dimension \(/\mathrm{m}^{3}\) by Definition 100.2: it is a number density, and it is negative.

The mechanism was identified by Nambu and Jona-Lasinio [Nambu:1961] by explicit analogy with superconductivity. In the BCS theory of superconductivity, treated in the condensed-matter part of this book, an arbitrarily weak attraction between electrons destabilizes the Fermi sea and produces a condensate of pairs, and with it a gap: an excitation that was massless acquires a mass generated by the interaction rather than written into the Hamiltonian. Replace electron pairs by quark–antiquark pairs and the gap by a quark mass, and Equation (102.76) is the same statement. A massless quark propagating through a vacuum full of \(\bar{q}q\) pairs acquires an effective mass, the constituent mass Equation (102.12) of about \(336\,\mathrm{MeV}\)\(/c^{2}\) — a hundred times its Lagrangian mass, and manufactured entirely by the interaction.

Theorem 102.70 (Goldstone).

Let \(j^{\mu}_{a}\) be a conserved current of a symmetry that does not annihilate the vacuum, in the sense that some local operator \(\Phi\) has \(\avg{\comm{Q_{a}}{\Phi}}\neq0\) with \(Q_{a}=c^{-1}\int\dd^{3}x\,j^{0}_{a}\). Then the theory contains a massless one-particle state \(\ket{\pi_{a}}\) with \(\avg{0|j^{\mu}_{a}|\pi_{a}}\neq0\) [Goldstone:1961]. Rests on Phenomenon 102.68 and Equation (102.75).

Proof.

Derives Theorem 102.70. Write \(v=\avg{\comm{Q_{a}}{\Phi(0)}}\neq0\) and insert a complete set of states between the operators:

\[ v=\frac{1}{c}\int\dd^{3}x\sum_{n}\left[ \avg{0|j^{0}_{a}(x)|n}\avg{n|\Phi(0)|0} -\avg{0|\Phi(0)|n}\avg{n|j^{0}_{a}(x)|0}\right]\ep \]

Translation invariance gives \(\avg{0|j^{0}_{a}(x)|n}=\avg{0|j^{0}_{a}(0)|n} \ee^{-\ii p_{n}\cdot x/\hbar}\), so the integral over \(\dd^{3}x\) produces \((2\pi\hbar)^{3}\delta^{3}(\vect{p}_{n})\): only states of zero three-momentum contribute. For those the remaining \(x\) dependence is \(\ee^{-\ii E_{n}t/\hbar}\), and therefore

\[ v=\frac{\left(2\pi\hbar\right)^{3}}{c}\sum_{n,\;\vect{p}_{n}=0} \left[\avg{0|j^{0}_{a}|n}\avg{n|\Phi|0}\ee^{-\ii E_{n}t/\hbar} -\text{c.c.\ term}\right]\ep \]

But \(Q_{a}\) is time independent, because \(\pp_{\mu}j^{\mu}_{a}=0\) and the spatial part integrates to a surface term; so \(v\) cannot depend on \(t\). Since \(v\neq0\), at least one term must survive, and a surviving term must have \(E_{n}=0\) at \(\vect{p}_{n}=0\). A state with zero energy at zero momentum is a massless particle, and by construction it has a non-vanishing matrix element with the current. Its quantum numbers are those of the current: for the axial current \(A^{\mu}_{a}\), which is a pseudovector, the state must be a pseudoscalar, since \(\avg{0|A^{\mu}_{a}|\pi_{b}}\) can only be proportional to \(p^{\mu}\) and \(p^{\mu}\) is the only vector available for a spinless state. Applying this to Equation (102.75): three generators are broken — \(\SU(2)_{L}\times\SU(2)_{R}\) has six and \(\SU(2)_{V}\) has three — so there are three massless pseudoscalars, an isotriplet.

Remark 102.71 (Where the mass of ordinary matter comes from).

Electroweak Unification and the Higgs Boson is often described as explaining the origin of mass. It does not explain the origin of our mass. The Higgs mechanism supplies the Lagrangian quark masses of Table 102.1, which contribute \(9\,\mathrm{MeV}\)\(/c^{2}\) of the proton's \(938.272\,\mathrm{MeV}\)\(/c^{2}\). The remaining \(99\,\mathrm{\%}\) is supplied by Equation (102.76) and Phenomenon 102.39: a dimensionless coupling that runs, a scale generated by that running, and a vacuum that condenses. Two different mechanisms, in two different sectors, and the one that matters for the mass of a table is the strong one.

The pion as a pseudo-Goldstone boson

Phenomenon 102.72 (The pion is anomalously light).

The lightest hadron is far lighter than the rest of the spectrum: \(m_{\pi^{\pm}}c^{2}=139.570\,\mathrm{MeV}\) and \(m_{\pi^{0}}c^{2}=134.977\,\mathrm{MeV}\), against \(775\,\mathrm{MeV}\) for the \(\rho\) and \(938\,\mathrm{MeV}\) for the nucleon [Navas:2024]. The gap is not a small effect to be absorbed into a fit; it is a factor of five or six, and it is specific to the pseudoscalar triplet. Furthermore it is \(m_{\pi}^{2}\), and not \(m_{\pi}\), that is proportional to the light quark masses, so that the pion alone becomes massless when those masses are set to zero. Rests on Theorem 102.70, Phenomenon 102.68 and Definition 102.69.

Derivation. Derives Phenomenon 102.72. Theorem 102.70 supplies three massless pseudoscalars in the limit \(m_{u}=m_{d}=0\), and Phenomenon 102.68 identifies the broken symmetry as the one whose Goldstone bosons they are. The pions are those states, made massive only by the explicit breaking. It remains to find how their mass depends on the breaking, and the answer is the Gell-Mann–Oakes–Renner relation [GellMann:1968]:

\begin{equation}\tag{102.77} \left(m_{\pi}c^{2}\right)^{2}F_{\pi}^{2} =-\left(m_{u}+m_{d}\right)c^{2}\, \avg{\bar{q}q}\left(\hbar c\right)^{3}\ec \end{equation}

where \(F_{\pi}\) is the pion decay constant of Equation (106.103), an energy, and every factor is written so that both sides carry (energy)\(^{4}\): \(\avg{\bar{q}q}\) carries \(/\mathrm{m}^{3}\) by Definition 102.69 and \(\left(\hbar c\right)^{3}\) converts it.

The derivation is a Ward identity. The divergence of the isotriplet axial current, exact in QCD, is \(\pp_{\mu}A^{\mu}_{a}=\left(m_{u}+m_{d}\right)c\,P_{a}\) with \(P_{a}=\ii\bar{q}\gamma^{5}\tau_{a}q\) the pseudoscalar density; it vanishes in the chiral limit, which is Proposition 102.67. Take its matrix element between the vacuum and a one-pion state. The left-hand side is fixed by the definition of the decay constant, \(\avg{0|A^{\mu}_{a}|\pi_{b}(p)}=\ii F_{\pi}p^{\mu}\delta_{ab}/c\), so that \(\avg{0|\pp_{\mu}A^{\mu}_{a}|\pi_{b}} =F_{\pi}\left(m_{\pi}c\right)^{2}\delta_{ab}/\hbar c\) on using \(p^{2}=m_{\pi}^{2}c^{2}\). The right-hand side is \(\left(m_{u}+m_{d}\right)c\avg{0|P_{a}|\pi_{b}}\), and the remaining matrix element is related to the condensate by the same chiral rotation that produced the Goldstone boson: acting with the broken charge on the pseudoscalar density returns the scalar density, so \(\avg{0|P_{a}|\pi_{b}}\) is proportional to \(\avg{\bar{q}q}/F_{\pi}\). Assembling the factors gives Equation (102.77). The structural point survives any bookkeeping: \(m_{\pi}^{2}\), not \(m_{\pi}\), is linear in the quark masses, because the quark mass enters the Lagrangian linearly and the pion mass enters the Lagrangian squared.

Example 102.73 (Two numbers out of the relation).

The condensate. Insert \(m_{\pi^{0}}c^{2}=134.977\,\mathrm{MeV}\), \(F_{\pi}=92.07\,\mathrm{MeV}\) from \(f_{\pi}=130.2(12)\,\mathrm{MeV}\), and \(\left(m_{u}+m_{d}\right)c^{2}=6.86\,\mathrm{MeV}\) [Navas:2024] into Equation (102.77). Writing \(-\avg{\bar{q}q}\left(\hbar c\right)^{3}=\Sigma^{3}\) with \(\Sigma\) an energy,

\begin{equation}\tag{102.78} \Sigma=\left[\frac{\left(m_{\pi}c^{2}\right)^{2}F_{\pi}^{2}} {\left(m_{u}+m_{d}\right)c^{2}}\right]^{1/3} =282\,\mathrm{MeV}\ec\qquad \avg{\bar{q}q}=-2.9\times 10^{45}\,/\mathrm{m}^{3}\ec \end{equation}

about three quark–antiquark pairs per cubic femtometre, and of the order of \(\Lambda_{\mathrm{QCD}}\) as Phenomenon 102.39 requires — there being nothing else for it to be made of. Lattice and sum-rule determinations of the condensate agree with Equation (102.78) at the ten-percent level [Navas:2024], which is the accuracy the leading order of the chiral expansion deserves.

A ratio with no condensate in it. Equation (102.77) extended to three flavours gives \(\left(m_{K}c^{2}\right)^{2}F_{K}^{2} =-\left(m_{d}+m_{s}\right)c^{2}\avg{\bar{q}q}(\hbar c)^{3}\) for the \(K^{0}\), so that in the limit \(F_{K}\to F_{\pi}\) the condensate cancels:

\begin{equation}\tag{102.79} \frac{m_{K^{0}}^{2}}{m_{\pi^{0}}^{2}} =\frac{m_{d}+m_{s}}{m_{u}+m_{d}} =\frac{4.70+93.5}{2.16+4.70}=14.3\ec \end{equation}

against the measured \(\left(497.611/134.977\right)^{2}=13.6\) [Navas:2024]: agreement to \(5\,\mathrm{\%}\). The neutral mesons are used because the charged ones carry an electromagnetic self-energy that Equation (102.77) does not describe. This is a genuine test, because the quark masses entering it are determined elsewhere — on the lattice and from sum rules — and not fitted to the meson masses.

Remark 102.74 (Chiral perturbation theory, and the ninth meson).

Equation (102.77) is the first term of a systematic expansion. Because the pions are the light degrees of freedom at low energy, and because their interactions are constrained by the broken symmetry to vanish as the momenta go to zero, one may write an effective Lagrangian for the pion field alone, organized in powers of momenta and quark masses: chiral perturbation theory [Weinberg:1979a] [Gasser:1984]. It is an effective field theory in the sense developed in the renormalization-group chapter, with a finite number of measurable constants at each order, and it predicts pion–pion scattering lengths, form factors and mass ratios to a few percent.

The pseudoscalar octet has eight members and the Goldstone count for \(\SU(3)_{L}\times\SU(3)_{R}\to\SU(3)_{V}\) is eight, so the \(\eta'\) is one too many. It is not a Goldstone boson of \(\U(1)_{A}\), because \(\U(1)_{A}\) is not a symmetry: the anomaly of Theorem 106.87 destroys it, with the gluon topological density of Section 106.6.2 on the right-hand side of the divergence equation. Weinberg's observation [Weinberg:1975] is that if \(\U(1)_{A}\) had been a spontaneously broken symmetry the ninth pseudoscalar would have had to weigh at most \(\sqrt{3}m_{\pi}\), about \(234\,\mathrm{MeV}\)\(/c^{2}\); the \(\eta'\) weighs \(957.78\,\mathrm{MeV}\)\(/c^{2}\) [Navas:2024]. Its heaviness is therefore evidence for the anomaly, and through it for the topological structure that Section 102.10 is about.

QCD under extreme conditions

Asymptotic freedom implies that matter compressed or heated far enough must eventually become a gas of quarks and gluons: at short distances the coupling is weak, and raising the temperature or the density reduces the mean separation. The transition is not a prediction of perturbation theory — it happens at \(\alpha_{s}\sim1\) — but of the lattice, and it has been looked for in heavy-ion collisions.

Phenomenon 102.75 (Hadronic matter deconfines when heated).

Lattice computation of QCD at finite temperature shows a rapid but smooth rise of the energy density, by a factor of order ten over a narrow range of temperature centred on \(k_{B}T_{c}\approx155\,\mathrm{MeV}\), that is

\begin{equation}\tag{102.80} T_{c}\approx\frac{155\,\mathrm{MeV}}{k_{B}} =1.8\times 10^{12}\,\mathrm{K}\ec \end{equation}

some \(10^{5}\) times the temperature at the centre of the Sun. It is a crossover, not a phase transition with a discontinuity. Above it the degrees of freedom are those of quarks and gluons rather than of hadrons. Rests on Theorem 102.35 and Definition 102.10.

Derivation of the size of the jump. Derives Phenomenon 102.75. The energy density of an ideal ultrarelativistic gas is the standard blackbody result, generalized to \(g\) internal states and to fermions by Fermi–Dirac counting:

\begin{equation}\tag{102.81} \varepsilon=g_{\mathrm{eff}}\, \frac{\pi^{2}\left(k_{B}T\right)^{4}}{30\left(\hbar c\right)^{3}}\ec \qquad g_{\mathrm{eff}}=g_{B}+\frac{7}{8}g_{F}\ec \end{equation}

the factor \(7/8\) being the ratio of the Fermi to the Bose integral \(\int_{0}^{\infty}x^{3}\dd x/(\ee^{x}\pm1)\). Counting: a gluon has \(8\) colours and \(2\) polarizations, so \(g_{B}=16\); a quark flavour has \(3\) colours, \(2\) spins, and particle and antiparticle, so \(g_{F}=12\) per flavour and \(g_{F}=36\) for the three light flavours. Hence \(g_{\mathrm{eff}}=16+\tfrac{7}{8}\cdot36=47.5\). Below the transition the degrees of freedom are three pions, \(g_{\mathrm{eff}}=3\). The ratio is nearly sixteen, and that factor — not the precise value of \(T_{c}\) — is what the lattice sees as a steep rise in \(\varepsilon/T^{4}\).

Numerically at \(k_{B}T=155\,\mathrm{MeV}\), Equation (102.81) gives \(\varepsilon\approx1.2\,\mathrm{GeV}\) per cubic femtometre, that is

\begin{equation}\tag{102.82} \varepsilon\approx1.9\times 10^{35}\,\mathrm{J}/\mathrm{m}^{3}\ec \end{equation}

against about \(0.4\,\mathrm{GeV}\) per cubic femtometre inside a nucleon at rest. The lattice value near \(T_{c}\) is some tens of percent below the ideal-gas limit Equation (102.81), and the deficit does not go away quickly as the temperature rises: the medium is not a weakly coupled gas, which is exactly what the heavy-ion data below also say.

The experimental programme is the collision of heavy nuclei at relativistic energies, and its conclusion after the first five years at RHIC [Adams:2005] rests on two observations that point the same way.

Jet quenching. A parton produced in a hard scattering inside the collision volume must traverse the medium before fragmenting into the jet of Section 102.6. If the medium is coloured it radiates gluons and loses energy. The observable is the suppression of high transverse-momentum hadrons relative to the same quantity in proton–proton collisions scaled by the number of binary collisions, and it is large — a factor of about five in central gold–gold collisions — while being absent in deuteron–gold collisions, which excludes an initial-state explanation.

Elliptic flow. A non-central collision has an almond-shaped overlap region. If the matter produced behaves as a fluid, the pressure gradients are larger along the short axis and the final momentum distribution is anisotropic in a way that remembers the initial geometry. The measured anisotropy is large and is described by relativistic hydrodynamics with a shear viscosity to entropy density ratio close to the smallest values known for any fluid. The conclusion of [Adams:2005] was accordingly not “a weakly coupled quark–gluon gas has been made” but “a strongly coupled, nearly ideal liquid has been made”, which was not what the perturbative picture had led anyone to expect and is the reason the result mattered.

Two further regimes are treated elsewhere. Cold dense quark matter is the possible interior of a neutron star and belongs to Compact Stars and Relativistic Astrophysics. And the whole universe passed through Equation (102.80) downwards at an age of a few microseconds, converting a quark–gluon plasma into hadrons; that is the cosmological quark–hadron transition, and it is the last epoch before the nucleosynthesis of Stellar Structure and Nucleosynthesis.

The strong CP problem

Definition 102.27 deferred one term, and it is the only known defect of the theory.

Proposition 102.76 (QCD permits one more term).

Gauge invariance, Lorentz invariance, locality and renormalizability permit, besides Equation (102.29), the term

\begin{equation}\tag{102.83} \Lag_{\theta}=\theta\hbar\,q(x) \propto\theta\,\frac{g_{s}^{2}}{\hbar}\, \sum_{a}\vect{E}^{a}\cdot\vect{B}^{a}\ec \end{equation}

with \(q(x)\) the gluon topological charge density of Definition 105.43 and \(\theta\) a dimensionless number, and nothing else. The term violates \(P\) and \(CP\) (Proposition 105.44), and only the combination \(\bar{\theta}=\theta+\arg\det M_{q}\) of Equation (105.111) is physical, because a chiral rotation of the quark fields shifts \(\theta\) through the anomaly of Theorem 106.87. Rests on Equation (102.29), Definition 105.43 and Theorem 100.17.

Proof.

Derives Proposition 102.76. Power counting (Theorem 100.17) allows operators of mass dimension at most four. The gauge-invariant, Lorentz-invariant operators of dimension four that can be built from \(F^{a}_{\mu\nu}\) are \(F^{a}_{\mu\nu}F^{a\,\mu\nu}\), already in Equation (102.29), and \(\epsilon^{\mu\nu\rho\sigma}F^{a}_{\mu\nu}F^{a}_{\rho\sigma}\), which is Equation (102.83); from the quark fields they are the kinetic and mass terms. There is no other. The identification of Equation (102.83) with the topological density, its periodicity in \(\theta\), and its origin in the vacuum structure are Corollary 106.84 and Section 106.6.2. The statement that only \(\bar{\theta}\) is observable is Equation (105.111): it ties the gluon sector to the phases of the quark mass matrix, which come from a different sector of the Standard Model altogether.

Phenomenon 102.77 (The neutron has no measurable electric dipole moment).

Every symmetry of QCD permits a term in the Lagrangian proportional to \(F\tilde{F}\); it violates \(P\) and \(CP\), and it would give the neutron a permanent electric dipole moment along its spin. None is seen. The best measurement is consistent with zero and bounds \(\abs{d_{n}}<1.8\times 10^{-28}\,e\,\mathrm{m}\) at 90 percent confidence [Abel:2020], which forces the coefficient of the permitted term below about \(10^{-10}\). Nothing in the theory explains why a term it allows should be absent to that accuracy. Rests on Proposition 102.76, Proposition 105.45 and Equation (105.112).

Derivation. Derives Phenomenon 102.77. The step that converts a measured bound on a dipole moment into a bound on a Lagrangian parameter is the chiral estimate Equation (105.112), derived in Proposition 105.45: \(d_{n}\approx\bar{\theta}\times 1.6\times 10^{-37}\text{–}6.4\times 10^{-37}\,\mathrm{C}\,\mathrm{m}\), the parametric form following from the requirement that the answer vanish with the quark masses — because a massless quark makes \(\theta\) unobservable by the chiral rotation of Proposition 102.76 — and the coefficient from a chiral loop. Combining with the ultracold-neutron measurement [Abel:2020] gives \(\abs{\bar{\theta}}\lesssim10^{-10}\), which is Equation (105.113). The uncertainty of the estimate is a factor of a few, which is why the bound is quoted as an order of magnitude.

Remark 102.78 (A fine-tuning puzzle, honestly labelled).

Nothing is inconsistent. \(\bar{\theta}=0\) is a perfectly good value of a parameter, and QCD with \(\bar{\theta}=0\) describes every measurement in this chapter. What is peculiar is that \(\bar{\theta}\) is an angle, periodic with period \(2\pi\), assembled from the gluon sector and the phases of the Yukawa couplings of Electroweak Unification and the Higgs Boson, which are otherwise of order one and which supply the large \(CP\) violation of Discrete Symmetries and CPT; and the two contributions cancel to ten decimal places. That is a fine-tuning puzzle, not a contradiction, and this book records it as such.

The best-known proposal promotes \(\bar{\theta}\) to a dynamical field whose potential is minimized at zero [Peccei:1977], which predicts a light pseudoscalar particle [Wilczek:1978]. No such particle has been detected in any search to date. In the terms of this book that makes it a proposal and not physics, and the strong \(CP\) problem is carried forward unresolved to What We Observe but Do Not Understand. It is worth noting what would settle it: detecting the particle, or measuring a nonzero neutron electric dipole moment, which would show \(\bar{\theta}\neq0\) and remove the puzzle by removing its premise.