Electroweak Unification and the Higgs Boson
The weak and electromagnetic interactions are two faces of one gauge theory. This chapter builds it: the gauge group \(\SU(2)_{L}\times\U(1)_{Y}\) of Glashow [Glashow:1961], its spontaneous breaking to the \(\U(1)_{\text{em}}\) of The Maxwell Equations by a scalar doublet, and the resulting massive \(W^{\pm}\) and \(Z^{0}\) carriers that replace the contact interaction of Fermi [Fermi:1934] used in Weak Interactions. The chapter therefore sits after Generalized Classical Field Theory, where the gauge-field formalism belongs, and after the renormalization of Quantum Electrodynamics and Renormalization, and before the strong sector of Quantum Chromodynamics with which it completes the Standard Model. Two threads run through it. The first is theoretical: how a gauge boson acquires mass without destroying the gauge invariance that makes the theory renormalizable — the mechanism of Anderson [Anderson:1963], Englert and Brout [Englert:1964], Higgs [Higgs:1964a] [Higgs:1964] and Guralnik, Hagen and Kibble [Guralnik:1964], whose renormalizability 't Hooft and Veltman established [tHooft:1971a] [tHooft:1972].
The second thread is evidential, and it is the reason the theory is in this book at all. Weak neutral currents were seen in Gargamelle [Hasert:1973a] [Hasert:1973b]; the \(W\) and \(Z\) were produced and weighed at the CERN \(p\bar{p}\) collider [Arnison:1983a] [Banner:1983] [Arnison:1983b] [Bagnaia:1983]; four LEP experiments and SLD measured the \(Z\) resonance to parts in \(10^{4}\) [Schael:2006]; and the scalar itself was found in 2012 — the discovery account belongs to Experiment: The Higgs Boson Discovery — and has since been mapped coupling by coupling [Aad:2022] [Tumasyan:2022]. Where the picture is incomplete — the metastability of the vacuum, the \(W\)-mass tension, the origin of the Yukawa hierarchy — the chapter says so. The standard reference treatment is [Weinberg:1996].
Conventions carried into this chapter
Everything below is written in SI units, in the conventions fixed once in Section 100.1.1: the metric is \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\), a four-momentum is \(p^{\mu}=(E/c,\vect{p})\) of dimension \(\mathrm{kg}\,\mathrm{m}/\mathrm{s}\), the mass shell is \(p^{2}=m^{2}c^{2}\) (Equation (100.1)), Mandelstam variables are squared momenta so that \(\sqrt{s}\,c\) is a centre-of-mass energy, and every Lagrangian density is an energy density, \(\mathrm{J}/\mathrm{m}^{3}\), with the action \(S=c^{-1}\int\dd^{4}x\,\Lag\) of Equation (100.2). The gauge fields and their couplings are those of Definition 101.66: \(W^{a}_{\mu}\) and \(B_{\mu}\) carry \(\mathrm{V}\,\mathrm{s}/\mathrm{m}\) like the electromagnetic potential of Definition 100.2, and \(g\) and \(g'\) carry the dimension of electric charge, \(\mathrm{C}\).
Three abbreviations recur so often that they are given names here.
For any coupling \(q\) of the dimension of charge, write
a pure number, since \(q^{2}/\varepsilon_{0}\) and \(\hbar c\) both carry \(\mathrm{J}\,\mathrm{m}\). In particular \(\hat{e}^{2}=e^{2}/(\varepsilon_{0}\hbar c)=4\pi\alpha\) by Equation (100.3), so \(\hat{e}=\sqrt{4\pi\alpha}=0.302822\) at zero momentum transfer, and \(\hat{g}\), \(\hat{g}'\) are the numbers that the references cited below write simply as \(g\) and \(g'\) (see Remark 104.2).
The scalar doublet \(\Phi\) carries the dimension of a scalar field, \(\mathrm{J}^{1/2}\,\mathrm{m}^{-1/2}\) (Equation (106.33)), so its vacuum expectation value \(v\) does too. The combination
is an energy, and it is what the literature means when it writes “\(v=246\,\mathrm{GeV}\)”. Its value is computed in Proposition 104.26.
Every source cited below sets \(\hbar=c=1\). The dictionary is the one given once in Remark 100.4, plus three entries special to this chapter: a coupling of charge dimension maps to its hat by Equation (104.1), a vacuum expectation value maps to \(E_{v}\) by Equation (104.2), and a scalar self-coupling \(\lambda\) — which carries \(/\mathrm{J}/\mathrm{m}\) here, by Equation (106.34) — maps to \(\hat{\lambda}=\lambda\hbar c\). With those three substitutions every formula below becomes the one in the references, and conversely. No derivation in this chapter is performed with \(\hbar\) or \(c\) set to one; this remark exists so that the reader can check the chapter against its sources.
From Fermi's theory to a gauge theory
The four-fermion interaction and its breakdown
Fermi's theory of beta decay [Fermi:1934] is a product of two currents evaluated at one point, Equation (101.6), with a single constant \(G_{F}\) in front. Weak Interactions develops it in full and measures it: the definitive determination comes from the positive-muon lifetime, which the MuLan experiment fixed to about one part per million [Tishchenko:2013], giving
by Equation (101.8). The Lorentz structure of the currents is \(V-A\), forced on the theory by the observation of maximal parity violation and written down by Feynman and Gell-Mann [Feynman:1958] and by Sudarshan and Marshak [Sudarshan:1958]; Section 101.4 derives it from the data and Experiment: Parity Violation is the apparatus.
Three facts about that theory are the reason this chapter exists, and all three are established in Weak Interactions. They are collected here because everything the electroweak theory does is a response to them.
Let \(G_{F}\) be as in Equation (104.3) and let \(E\) denote a centre-of-mass energy. Then
-
the dimensionless number governing every weak amplitude is \(G_{F}E^{2}/(\hbar c)^{3}\), which reaches unity at \(E_{F}=292.8\,\mathrm{GeV}\) (Equation (101.9)), so the perturbation expansion has an expiry date;
-
the \(J=0\) cross-section \(\sigma=G_{F}^{2}E^{2}/\pi(\hbar c)^{4}\) of Equation (101.60) breaches the partial-wave unitarity bound at \(E=1.04\,\mathrm{TeV}\) (Theorem 101.78);
-
the superficial degree of divergence \(D=4-\tfrac{3}{2}E_{\text{ext}}+2V\) of Equation (101.91) grows with the order \(V\) of perturbation theory, so the theory needs infinitely many counterterms and predicts nothing beyond the order at which it is fitted.
All three are consequences of one fact: \(G_{F}\) carries \(\mathrm{J}\,\mathrm{m}^{3}\) (Proposition 101.7), a dimension no rescaling can remove. Rests on Equation (104.3), Proposition 101.7 and Theorem 101.78.
Derivation. Derives Proposition 104.3. Each item is proved where it is stated in Weak Interactions; what is added here is that the three are the same statement counted three ways. A coupling of dimension \(\mathrm{J}\,\mathrm{m}^{3}\) can only enter an amplitude in the combination \(G_{F}E^{2}/(\hbar c)^{3}\), because that is the unique dimensionless product of \(G_{F}\), \(\hbar\), \(c\) and an energy; hence item (i). A cross-section is an area and the only area available is \(G_{F}^{2}E^{2}/(\hbar c)^{4}\) times a pure number, which grows while the unitarity bound \(16\pi(\hbar c)^{2}/E^{2}\) falls, so they must cross; hence item (ii). And in the momentum counting of Theorem 100.17 each vertex contributes its coupling's negative mass dimension to \(D\) with a positive sign, which is Corollary 101.82; hence item (iii).
∎The obvious repair is Yukawa's: give the interaction a carrier [Yukawa:1935]. Two currents joined by a heavy vector boson of mass \(M_{W}\) reproduce the contact term at momentum transfers small compared with \(M_{W}c\) (Proposition 101.65), with the matching condition Equation (101.78),
the last form using \(\hat{g}^{2}=4\pi\alpha_{W}\) from Equation (101.71) and Equation (104.1). This trades one dimensionful constant for a dimensionless coupling and a mass, which is progress only if the mass is explained. It is not, and the cure is worse than it looks.
Let \(W_{\mu}\) be a vector field with the Proca mass term
added by hand to a gauge-invariant kinetic term. Then
-
\(\Lag_{M}\) is not invariant under \(W_{\mu}\mapsto W_{\mu}+\pp_{\mu}\Lambda\);
-
its propagator Equation (101.77) tends to a constant rather than to zero at large momentum, so the power counting of Theorem 100.17 fails and the theory is non-renormalizable;
-
the longitudinal polarization vector grows linearly with momentum,
\begin{equation}\tag{104.6} \varepsilon^{\mu}_{(L)}(k) =\frac{k^{\mu}}{M_{W}c}+\mathcal{O}\! \left(\frac{M_{W}c}{\abs{\vect{k}}}\right)\ec \end{equation}so amplitudes with external longitudinal \(W\)s grow with energy and the unitarity problem of Proposition 104.3(ii) reappears at a higher scale instead of being cured.
Rests on Equation (101.77), Theorem 100.17 and Proposition 104.3.
Derives Theorem 104.4. (i) Under the shift, \(W^{+}_{\mu}W^{-\mu}\) acquires \(\pp_{\mu}\Lambda^{+}W^{-\mu}+W^{+}_{\mu}\pp^{\mu}\Lambda^{-} +\pp_{\mu}\Lambda^{+}\pp^{\mu}\Lambda^{-}\), which no choice of \(\Lambda\) makes vanish identically. The mass term breaks the very invariance that Theorem 100.5 used to construct the interaction.
(ii) The massive vector kernel is \(-\mu_{0}^{-1}\hbar^{-2}\left[(k^{2}-M_{W}^{2}c^{2})\eta_{\mu\nu} -k_{\mu}k_{\nu}\right]\), whose inverse is Equation (101.77). At large \(k\) the transverse part falls as \(k^{-2}\), exactly as the photon propagator Equation (100.17) does, but the term \(k_{\mu}k_{\nu}/M_{W}^{2}c^{2}\) divided by \(k^{2}-M_{W}^{2}c^{2}\) tends to \(\eta\)-independent constant behaviour: it does not fall at all. Repeating the count of Theorem 100.17 with internal boson lines contributing \(k^{0}\) instead of \(k^{-2}\) raises \(D\) by two for every such line, and \(D\) again grows with the order.
(iii) A massive vector of four-momentum \(k^{\mu}=(E/c,0,0,k)\) with \(E^{2}/c^{2}-k^{2}=M_{W}^{2}c^{2}\) has three polarization vectors satisfying \(k_{\mu}\varepsilon^{\mu}=0\) and \(\varepsilon\cdot\varepsilon^{*}=-1\); the longitudinal one is
as one checks by expanding \(E/c=\sqrt{k^{2}+M_{W}^{2}c^{2}}\). The second term is bounded; the first grows without limit, and it is the one that appears in Equation (104.6). Since an amplitude is linear in each external polarization, an amplitude with \(n\) longitudinal \(W\)s carries \(n\) factors of \(k/M_{W}c\) relative to the transverse one. The physical content is that the longitudinal mode is the piece a genuinely massless gauge field does not have; inserting a mass by hand supplies it without supplying a field for it to come from.
∎Theorem 104.4 is the pivot of the chapter and it is worth naming its resolution in advance. The longitudinal mode is not manufactured out of nothing: it is a scalar degree of freedom that existed before the symmetry broke, and the growth of Equation (104.6) in one description is the ordinary, non-growing scalar amplitude in another. That equivalence, made exact in Proposition 104.47, is what saves unitarity, and the requirement that it hold fixes the couplings of the scalar (Theorem 104.51) and hence everything the LHC measures in Section 104.7.2. Remark 101.69 states the same conclusion from the other side.
Yang–Mills fields and the mass problem
The gauge principle of Theorem 100.5 takes a global phase symmetry and makes it local, at the cost of one vector field. Yang and Mills [Yang:1954] asked what happens when the global group is non-abelian, and the answer is the framework in which the rest of this chapter is written. The general construction belongs to Generalized Classical Field Theory and its geometry to Section 14.6; what is needed here is the explicit \(\SU(2)\) case in SI units.
Let \(T^{a}=\tfrac{1}{2}\sigma^{a}\) be the generators of \(\SU(2)\), with \(\comm{T^{a}}{T^{b}}=\ii\varepsilon^{abc}T^{c}\) and \(\tr(T^{a}T^{b})=\tfrac{1}{2}\delta^{ab}\) (Section 14.3). Let \(\psi\) be a field in some representation with generators \(T^{a}\), and \(W^{a}_{\mu}\) three vector fields of dimension \(\mathrm{V}\,\mathrm{s}/\mathrm{m}\). The covariant derivative is
with \(g\) of dimension \(\mathrm{C}\), and the field strength is
so that the Yang–Mills Lagrangian density is
of dimension \(\mathrm{J}/\mathrm{m}^{3}\), exactly as in Equation (100.12). This is Equation (106.65) with the coupling displayed rather than absorbed.
Under the local transformation \(\psi\mapsto U(x)\psi\) with \(U(x)=\exp(-\ii\theta^{a}(x)T^{a})\), Equation (104.7) is covariant, \(D_{\mu}\psi\mapsto U D_{\mu}\psi\), if and only if the potential transforms as
and then \(W_{\mu\nu}:=W^{a}_{\mu\nu}T^{a}\) transforms homogeneously, \(W_{\mu\nu}\mapsto UW_{\mu\nu}U^{-1}\), so that Equation (104.9) is invariant. Expanding Equation (104.8) inside Equation (104.9) produces, besides the free quadratic term, a cubic vertex of strength \(g/\hbar\mu_{0}\) and a quartic vertex of strength \(g^{2}/\hbar^{2}\mu_{0}\): the gauge bosons carry the charge they mediate. Rests on Equations (104.7), (104.8) and (104.9).
Derives Proposition 104.7. The first statement is Equation (100.10) with the phase replaced by a matrix and the extra term coming from the failure of \(U\) and \(W_{\mu}\) to commute; writing \(\mathcal{W}_{\mu}=gW_{\mu}/\hbar\) of dimension \(/\mathrm{m}\), the requirement \((\pp_{\mu}+\ii\mathcal{W}'_{\mu})U\psi =U(\pp_{\mu}+\ii\mathcal{W}_{\mu})\psi\) for all \(\psi\) gives \(\ii\mathcal{W}'_{\mu}U=U\ii\mathcal{W}_{\mu}-(\pp_{\mu}U)\), which is Equation (104.10). For the field strength, note that \(\ii\mathcal{W}_{\mu\nu}=\comm{D_{\mu}}{D_{\nu}}\) is a commutator of covariant objects, hence a covariant object: it is Equation (106.64), and in components it is Equation (104.8) because \(\comm{T^{b}}{T^{c}}=\ii\varepsilon^{bcd}T^{d}\) contributes \(-\varepsilon^{abc}\mathcal{W}^{b}_{\mu}\mathcal{W}^{c}_{\nu}\). Since \(\tr(W_{\mu\nu}W^{\mu\nu})\) is invariant under conjugation, Equation (104.9) is invariant. Finally, squaring Equation (104.8) gives cross terms \((\pp W)(g\varepsilon WW/\hbar)\) and \((g\varepsilon WW/\hbar)^{2}\), which are the cubic and quartic self-interactions; they have no counterpart in Equation (100.12), where \(\varepsilon^{abc}\) has nothing to contract. That the gauge bosons interact with one another is a prediction, and it is measured directly in Phenomenon 104.57.
∎The obstruction is now stated in its final form.
No term of the form \(\tfrac{1}{2\mu_{0}}(Mc/\hbar)^{2}W^{a}_{\mu}W^{a\mu}\) is invariant under Equation (104.10), and no gauge-invariant local function of \(W^{a}_{\mu}\) alone gives a mass to the vector field. Rests on Equation (104.10).
Derives Theorem 104.8. For an infinitesimal transformation, \(\delta(gW^{a}_{\mu}/\hbar)=\pp_{\mu}\theta^{a} +\varepsilon^{abc}\theta^{b}(gW^{c}_{\mu}/\hbar)\), obtained by expanding Equation (104.10) to first order. Then
The second term vanishes because \(\varepsilon^{abc}\) is antisymmetric in \(ac\) while \(W^{c}_{\mu}W^{a\mu}\) is symmetric; the first does not, and it cannot be cancelled by anything, since \(\theta^{a}\) is an arbitrary function while \(W^{a\mu}\) is a field. For the second claim, observe that a mass term is by definition quadratic in the field with no derivatives, and Equation (104.10) contains an inhomogeneous piece \(\pp_{\mu}\theta\) that only a derivative can absorb; every gauge-invariant local functional of \(W\) alone is therefore built from \(W_{\mu\nu}\) and its covariant derivatives, each of which carries at least one \(\pp_{\mu}\) and hence at least two powers of momentum. Such a term contributes to the kinetic operator, never to a mass.
∎Theorem 104.8 forbids writing a mass down. It does not prove that the quantum is massless, and Schwinger made the distinction sharp [Schwinger:1962]: what fixes the mass is the position of the pole in the full propagator, not the form of a term in the Lagrangian. Write the exact photon two-point function of Corollary 100.32 as
which is transverse, and hence gauge invariant, whatever \(\Pi\) is. If the vacuum polarization \(\Pi(k^{2})\) is regular at \(k^{2}=0\), the pole sits at \(k^{2}=0\) and the quantum is massless. If instead \(\Pi(k^{2})\to -M^{2}c^{2}/k^{2}\) as \(k^{2}\to0\) — a pole in \(\Pi\) — the propagator's pole moves to \(k^{2}=M^{2}c^{2}\) and the quantum is massive, with the transversality, and therefore the gauge invariance, untouched. Schwinger exhibited this in two-dimensional electrodynamics, where the massless-fermion loop produces exactly such a pole. The physical realization in three space dimensions is the superconductor: the photon inside one is massive, with \(M=\hbar/\lambda_{L}c\) for the London penetration depth \(\lambda_{L}\), and nothing about electromagnetic gauge invariance is lost (Superconductivity and Superfluidity). What supplies the pole in \(\Pi\) is a massless scalar excitation, and the rest of this section is the search for one.
Glashow's $\SU(2)_{L}\times\U(1)_{Y}$
The charged weak currents of Equation (101.35) do not close under commutation: \(\comm{J^{+}}{J^{-}}\) produces a third, electrically neutral current. A gauge theory containing them must therefore have at least three generators, and the smallest group with three is \(\SU(2)\). Glashow [Glashow:1961] took the further step of adjoining a \(\U(1)\) so that electromagnetism could be accommodated without identifying the photon with the neutral member of the triplet, which is impossible because the photon does not violate parity while the \(\SU(2)\) currents maximally do. Salam and Ward reached the same structure independently [Salam:1964].
The gauge group is \(\SU(2)_{L}\times\U(1)_{Y}\) with gauge fields \(W^{a}_{\mu}\) and \(B_{\mu}\) and couplings \(g\), \(g'\), acting through Equation (101.70),
The subscript \(L\) is not decoration: \(T^{a}\) acts on the left-chiral projection \(\psi_{L}=\tfrac{1}{2}(\identity-\gamma^{5})\psi\) only, and annihilates \(\psi_{R}\). One generation of matter is
with the quarks additionally carrying colour, on which \(\SU(2)_{L}\times\U(1)_{Y}\) acts trivially. The hypercharge \(Y\) is assigned so that
reproduces the observed electric charges, \(Q\) measured in units of \(e\).
| Field | $T$ | $T^{3}$ | $Y$ | $Q$ |
|---|---|---|---|---|
| $\nu_{eL}$ | $\tfrac{1}{2}$ | $+\tfrac{1}{2}$ | $-1$ | $0$ |
| $e_{L}$ | $\tfrac{1}{2}$ | $-\tfrac{1}{2}$ | $-1$ | $-1$ |
| $e_{R}$ | $0$ | $0$ | $-2$ | $-1$ |
| $u_{L}$ | $\tfrac{1}{2}$ | $+\tfrac{1}{2}$ | $+\tfrac{1}{3}$ | $+\tfrac{2}{3}$ |
| $d_{L}$ | $\tfrac{1}{2}$ | $-\tfrac{1}{2}$ | $+\tfrac{1}{3}$ | $-\tfrac{1}{3}$ |
| $u_{R}$ | $0$ | $0$ | $+\tfrac{4}{3}$ | $+\tfrac{2}{3}$ |
| $d_{R}$ | $0$ | $0$ | $-\tfrac{2}{3}$ | $-\tfrac{1}{3}$ |
| $\Phi^{+}$ | $\tfrac{1}{2}$ | $+\tfrac{1}{2}$ | $+1$ | $+1$ |
| $\Phi^{0}$ | $\tfrac{1}{2}$ | $-\tfrac{1}{2}$ | $+1$ | $0$ |
This chapter and Weak Interactions write \(Q=T^{3}+Y/2\), which is the convention of the original papers and of Table 104.1. Proposition 106.96 writes \(Q=T^{3}+Y\), with every hypercharge half as large — so the left-handed quark doublet carries \(Y=1/6\) there and \(Y=1/3\) here. Both are in wide use and neither is wrong; the ratios that carry physics, and every anomaly condition, are identical. The reader comparing the two must halve or double, and nothing else.
That the assignments of Table 104.1 are consistent at all is a nontrivial statement, since \(Y\) must be constant on each \(\SU(2)_{L}\) multiplet while \(Q\) is not.
Given Equation (104.13), the requirement that \(Y\) take one value on each irreducible \(\SU(2)_{L}\) multiplet fixes \(Y\) on that multiplet to twice the average electric charge of its members, and is consistent only if the charges within a doublet differ by exactly one unit. Rests on Equation (104.13).
Derives Proposition 104.12. For a doublet with members of charge \(Q_{\uparrow}\), \(Q_{\downarrow}\) and \(T^{3}=\pm\tfrac{1}{2}\), Equation (104.13) gives \(Y=2Q_{\uparrow}-1=2Q_{\downarrow}+1\); the two are equal if and only if \(Q_{\uparrow}-Q_{\downarrow}=1\), and then \(Y=Q_{\uparrow}+Q_{\downarrow}\), twice the average. The leptonic doublet \((\nu_{e},e)\) has \(0-(-1)=1\) and the quark doublet \((u,d)\) has \(\tfrac{2}{3}-(-\tfrac{1}{3})=1\): both pass, and it is the charged current that requires it, since \(W^{+}\) carries exactly one unit of charge. The right-chiral fields are singlets, so \(T^{3}=0\) and \(Y=2Q\) with no constraint at all. Nothing so far explains why the charges come in these ratios; that is the content of Section 104.5.3, where the same numbers are forced by consistency at one loop.
∎Two consequences follow immediately and were Glashow's main results. First, since \(B_{\mu}\) and \(W^{3}_{\mu}\) are both neutral, the photon must be a combination of them and the orthogonal combination is a second neutral boson: the model predicts a weak neutral current, of a kind Fermi's theory does not contain and no experiment had then seen. Proposition 101.67 carries out the rotation and reads off the fermion couplings; the experimental discovery is Section 104.6.1. Second, the neutral current so obtained is flavour diagonal at tree level provided the quark doublets are complete, which is the mechanism of Glashow, Iliopoulos and Maiani [Glashow:1970] proved as Theorem 101.59 and used again in Section 104.4.3.
The 1961 paper inserts the \(W\) and \(Z\) masses by hand [Glashow:1961], and says so: its closing paragraph records that the resulting theory is not renormalizable. By Theorem 104.4 that is not a blemish to be tidied up later but the same disease Fermi's theory had, moved to a different place. The group structure was right and the mass terms were fatal, and the six years between 1961 and 1967 were spent finding out how to have the first without the second.
Spontaneous symmetry breaking
The superconducting analogy
A symmetry of the dynamics need not be a symmetry of the state the system actually occupies. A ferromagnet below its Curie temperature has a magnetization pointing somewhere, although the exchange Hamiltonian singles out no direction; a crystal has axes, although the interatomic potential is rotationally invariant. Nothing about this is quantum mechanical, and Phase Transitions and Critical Phenomena treats it as the general phenomenon of an ordered phase. What Nambu saw [Nambu:1960] is that the same thing can happen to the vacuum of a relativistic field theory, and that when it does, the consequences are as sharp as in the laboratory.
Let a theory have an action invariant under a group \(G\) and let \(\ket{0}\) be its ground state. The symmetry is spontaneously broken to the subgroup \(H=\set{h\in G:U(h)\ket{0}=\ket{0}}\) when \(H\neq G\). The set of states \(\set{U(g)\ket{0}:g\in G}\) is then in bijection with the coset space \(G/H\), all of them degenerate in energy, and the choice among them is not made by the dynamics.
The ancestor of the Higgs field is not an elementary particle at all but the Ginzburg–Landau order parameter of a superconductor [Ginzburg:1950]: a complex field \(\psi\) whose free energy density, near the transition, is \(a\abs{\psi}^{2}+\tfrac{1}{2}b\abs{\psi}^{4}\) with \(a\) changing sign at \(T_{c}\). The following proposition is the entire mathematical content of that model, and it will be used unchanged for the electroweak potential.
Let \(\phi\) be a complex field of dimension \(\mathrm{J}^{1/2}\,\mathrm{m}^{-1/2}\) and let
with \(\mu^{2}>0\) and \(\lambda>0\). Then \(\phi=0\) is a local maximum, the minima form the circle
and the depth of the well is \(V_{\min}=-\lambda v^{4}/4\). Rests on Definition 104.14.
Derives Proposition 104.15. Write \(\abs{\phi}^{2}=\rho\). Then \(V=-\mu^{2}\rho+\lambda\rho^{2}\), \(\dv{V}{\rho}=-\mu^{2}+2\lambda\rho\), which vanishes at \(\rho=\mu^{2}/2\lambda\), and \(\dd^{2}V/\dd\rho^{2}=2\lambda>0\) there, so it is the minimum; \(\rho=0\) gives \(\dd V/\dd\rho=-\mu^{2}<0\), so the origin is a maximum. Setting \(\rho=v^{2}/2\) gives \(v^{2}=\mu^{2}/\lambda\), which is Equation (104.15), and \(V_{\min}=-\mu^{2}v^{2}/2+\lambda v^{4}/4=-\lambda v^{4}/4\) on using \(\mu^{2}=\lambda v^{2}\). The dimensions check: \([\mu^{2}\abs{\phi}^{2}] =/\mathrm{m}^{2}\times\mathrm{J}/\mathrm{m} =\mathrm{J}/\mathrm{m}^{3}\), an energy density, and likewise for the quartic. Every point of the circle is a minimum and the phase of \(\phi\) at the minimum is undetermined: that undetermined phase is the degeneracy of Definition 104.14, with \(G=\U(1)\) and \(H\) trivial.
∎In the Bardeen–Cooper–Schrieffer ground state (Superconductivity and Superfluidity) the electron number symmetry is spontaneously broken by the condensate, and the single-particle spectrum acquires a gap \(\Delta\). Nambu [Nambu:1960] recognized the gap equation as a self-consistent mass generation and asked what the relativistic analogue would be. With Jona-Lasinio he built it [Nambu:1961]: a theory of massless fermions with a chirally invariant four-fermion interaction develops, above a critical coupling, a non-zero \(\avg{\bar{\psi}\psi}\) and hence a fermion mass that no term in the Lagrangian contains. The model also produced a massless pseudoscalar bound state, which Nambu identified with the pion — the first appearance of what the next subsection proves must always be there. Two lessons carry over. Mass can be a property of the vacuum rather than of the Lagrangian; and a broken continuous symmetry leaves a massless excitation behind. The strong interaction realizes both, and Quantum Chromodynamics takes up the pion as an approximate Goldstone boson of chiral symmetry.
Goldstone's theorem
Goldstone's model [Goldstone:1961] is Proposition 104.15 promoted to a field theory. Write \(\phi=\frac{1}{\sqrt{2}}(v+h+\ii\chi)\) with \(h\), \(\chi\) real fields measuring the departure from one chosen minimum, radially and tangentially. Substituting into \(\Lag=\pp_{\mu}\phi^{*}\pp^{\mu}\phi-V(\phi)\) and expanding gives
in which the field \(\chi\) has no mass term at all. By the dictionary of Notation 106.23, a mass term \(\tfrac{1}{2}(mc/\hbar)^{2}\) means
The massless mode is the one that moves along the circle of minima, where the potential is flat by construction. That this is no accident of the quartic potential is Goldstone's theorem.
Let \(\phi_{i}\), \(i=1,\dots,n\), be real scalar fields with a potential \(V(\phi)\) invariant under a compact connected group \(G\) acting as \(\delta\phi_{i}=\theta^{a}\left(t^{a}\right)_{ij}\phi_{j}\) with real antisymmetric generators \(t^{a}\), \(a=1,\dots,\dim G\). Let \(\phi_{0}\) minimize \(V\) and let \(H\subseteq G\) be its stabilizer. Then the mass matrix
annihilates every vector \(t^{a}\phi_{0}\), so the spectrum contains exactly \(\dim G-\dim H\) massless scalars, one for each broken generator. Rests on Definition 104.14.
Derives Theorem 104.17. Invariance of \(V\) under an infinitesimal transformation reads, for every \(\phi\) and every \(a\),
This is an identity in \(\phi\), so it may be differentiated. Applying \(\pp/\pp\phi_{k}\) gives
Now evaluate at \(\phi=\phi_{0}\). The second term vanishes because \(\phi_{0}\) is a stationary point of \(V\), and the first becomes
Hence each vector \(t^{a}\phi_{0}\) lies in the kernel of \(M^{2}\). If \(t^{a}\) belongs to the Lie algebra of \(H\) then \(t^{a}\phi_{0}=0\) and Equation (104.20) is empty; otherwise \(t^{a}\phi_{0}\neq0\) and it is a genuine null eigenvector. The vectors \(\set{t^{a}\phi_{0}}\) span a space of dimension \(\dim G-\dim H\), because the map \(t\mapsto t\phi_{0}\) from the Lie algebra of \(G\) to \(\R^{n}\) has kernel exactly the Lie algebra of \(H\) by definition of the stabilizer; and \(M^{2}\), being a real symmetric matrix, has as many zero eigenvalues as the dimension of its kernel. Since \(\phi_{0}\) is a minimum, \(M^{2}\) is positive semi-definite and no eigenvalue is negative. Therefore the spectrum consists of \(\dim G-\dim H\) massless modes and \(n-\dim G+\dim H\) massive ones.
Geometrically, \(t^{a}\phi_{0}\) is tangent to the orbit \(G\phi_{0}\) at \(\phi_{0}\), and \(V\) is constant on that orbit by Equation (104.19): a direction along which the potential does not change is a direction in which the field costs no energy to excite at zero momentum, which is what a massless particle is.
∎Theorem 104.17 is a statement about a classical potential, and Proposition 106.45 shows how loops modify a potential. The result nevertheless survives: Goldstone, Salam and Weinberg [Goldstone:1962] proved it to all orders, and the proof uses only the existence of a conserved current \(j^{\mu}\) with \(\avg{0|\comm{Q}{\mathcal{O}}|0}\neq0\) for some local operator \(\mathcal{O}\), together with Lorentz invariance, locality and a positive-definite state space. The spectral representation of Proposition 106.47, applied to \(\avg{0|j^{\mu}(x)\mathcal{O}(0)|0}\), then forces a zero-mass contribution. Each assumption is load-bearing, and their precise statement belongs to Axiomatic Quantum Field Theory. The one that fails in this chapter is the last but one: gauge fixing a gauge theory either breaks manifest Lorentz invariance (radiation gauge) or introduces negative-norm states (covariant gauge), and that is the loophole the next subsection walks through. It is not a technicality — it is why the photon in a superconductor can be massive without a massless mode appearing in the spectrum.
The Anderson–Higgs mechanism
Anderson [Anderson:1963] pointed out, in a non-relativistic setting and before any of the relativistic papers, that a gauge field coupled to a system with a broken symmetry does not leave a massless scalar behind: the would-be Goldstone mode and the gauge field combine, and what is left is a massive vector. In a plasma the resulting mass is the plasma frequency; in a superconductor it is the inverse London penetration depth. The relativistic version is worked here in full, for one abelian gauge field, because every feature of the electroweak case is already present.
Let \(\phi\) be a complex scalar of charge \(q\) coupled to a \(\U(1)\) gauge field through \(D_{\mu}=\pp_{\mu}+\ii qA_{\mu}/\hbar\), with
Then the spectrum consists of one real scalar of mass \(m_{h}=\hbar\sqrt{2\lambda v^{2}}/c\) and one massive vector of mass
with no massless scalar. The count of degrees of freedom is unchanged: two (complex scalar) plus two (massless vector) before, one plus three after. Rests on Equations (100.8), (104.6) and (104.16).
Derives Theorem 104.19. Write \(\phi(x)=\frac{1}{\sqrt{2}}\left[v+h(x)\right] \exp\!\left[\ii\chi(x)/v\right]\), which is legitimate wherever \(\phi\neq0\) and which parametrizes the same two real degrees of freedom as before, \(h\) radial and \(\chi\) tangential; \(\chi/v\) is dimensionless because \([\chi]=[v]\). The potential depends only on \(\abs{\phi}\), so it is exactly Equation (104.16) without the \(\chi\) terms: \(h\) gets the mass Equation (104.17) and \(\chi\) appears nowhere in \(V\).
Now perform the gauge transformation Equation (100.8) with \(\Lambda=-\hbar\chi/qv\), under which \(\phi\mapsto\ee^{-\ii q\Lambda/\hbar}\phi =\frac{1}{\sqrt{2}}(v+h)\): the phase is gone. This choice is the unitary gauge, and it is available precisely because \(\chi\) shifts inhomogeneously under a gauge transformation, exactly as \(A_{\mu}\) does. What was a field is now a gauge parameter. The Lagrangian becomes
using \(\abs{D_{\mu}\phi}^{2}=\tfrac{1}{2}(\pp h)^{2} +\tfrac{1}{2}(q/\hbar)^{2}(v+h)^{2}A^{2}\) for a real \(\phi\). The second term of Equation (104.23) is a mass term. Comparing it with the Proca form \(\tfrac{1}{2\mu_{0}}(m_{A}c/\hbar)^{2}A^{2}\) gives
which is Equation (104.22). Dimensionally, \(q^{2}v^{2}\mu_{0}/c^{2}\) carries \(\mathrm{C}^{2}\times\mathrm{J}/\mathrm{m} \times\mathrm{kg}\,\mathrm{m}/\mathrm{C}^{2} \times\mathrm{s}^{2}/\mathrm{m}^{2} =\mathrm{kg}^{2}\), so \(m_{A}\) is a mass.
For the counting: before breaking, \(\phi\) has two real components and a massless vector has two transverse polarizations, four in all; after, \(h\) has one and a massive vector has three, again four. The Goldstone mode has not disappeared, it has become the longitudinal polarization whose absence from the massless theory Equation (104.6) identified as the problem. The last two terms of Equation (104.23) are the interactions of the surviving scalar with the vector, and their strengths are not free: both are fixed by \(m_{A}\) and \(v\). That is the statement measured in Phenomenon 104.62.
∎In the static limit, Equation (104.23) gives the field equation \(\left(\nabla^{2}-m_{A}^{2}c^{2}/\hbar^{2}\right)\vect{A}=0\) for the transverse potential, so a magnetic field entering the ordered medium is screened exponentially over the length
the Compton wavelength of the massive photon. This is the Meissner effect, observed in 1933 [Meissner:1933] and measured as a penetration depth of tens of nanometres in elemental superconductors (Superconductivity and Superfluidity). Rests on Equations (104.22) and (104.23).
Derives Corollary 104.20. Varying Equation (104.23) with respect to \(A_{\mu}\) gives \(\pp_{\mu}F^{\mu\nu} =-\mu_{0}q^{2}v^{2}\hbar^{-2}A^{\nu}\), and in a static configuration with \(\pp_{0}=0\) and \(\nabla\cdot\vect{A}=0\) the spatial part reads \(\nabla^{2}\vect{A}=\mu_{0}q^{2}v^{2}\hbar^{-2}\vect{A} =(m_{A}c/\hbar)^{2}\vect{A}\) by Equation (104.22). A solution decaying into the half-space \(z>0\) is \(\vect{A}\propto\ee^{-z/\lambda_{L}}\) with \(\lambda_{L}\) as stated, and \(\vect{B}=\curl\vect{A}\) decays with it.
∎Theorem 104.17 is not contradicted; its hypotheses are not met. In unitary gauge the massless scalar is simply absent from the field content, and the theory manifestly has no massless spin-zero particle — but unitary gauge is not manifestly Lorentz covariant in its propagator structure, and the vector propagator does not fall off (Theorem 104.4(ii)), so the assumptions used in Remark 104.18 fail. In a covariant gauge the Goldstone field \(\chi\) is present, but so are negative-norm states, and the positivity assumption fails instead. Either way one hypothesis goes. Which one is a matter of bookkeeping, and Section 104.5.1 shows that the freedom to choose is exactly what makes the theory both unitary and renormalizable.
The 1964 papers
Three groups published the relativistic mechanism in 1964, within four months of one another, and they proved different things.
Englert and Brout [Englert:1964] computed the vacuum polarization of a gauge field coupled to a scalar with a non-zero expectation value and exhibited the pole in \(\Pi(k^{2})\) that Remark 104.9 identified as the criterion: their result is that the gauge boson mass emerges from a resummation, and they gave it for a general non-abelian group in the same paper.
Higgs's first paper, in Physics Letters [Higgs:1964a], is the negative half of the argument: it shows by explicit construction that the Goldstone theorem's proof breaks in a gauge theory quantized in a Lorentz gauge, because the current is not the generator of a symmetry of the physical state space. His second paper, in Physical Review Letters [Higgs:1964], works the abelian model of Theorem 104.19 and states the consequence the others did not emphasize: besides the massive vector there survives one massive scalar quantum, with definite couplings. That particle is the subject of Section 104.7, and it is the reason the mechanism is falsifiable rather than merely consistent. Higgs returned to its properties in 1966 [Higgs:1966], computing its decays into pairs of gauge bosons.
Guralnik, Hagen and Kibble [Guralnik:1964] worked in radiation gauge, where Lorentz invariance is not manifest, and showed explicitly that the massless mode still occurs in the field algebra but decouples from the physical states: their treatment is the one that makes the evasion of Theorem 104.17 transparent rather than mysterious.
Kibble then supplied what the Standard Model actually uses [Kibble:1967]: the extension to a non-abelian group, where the vector bosons need not all acquire the same mass and some may acquire none.
Let a set of scalars \(\Phi\) transform in a representation of a gauge group with generators \(T^{a}\) and couplings \(g_{a}\), coupled through \(D_{\mu}=\pp_{\mu}+\ii g_{a}A^{a}_{\mu}T^{a}/\hbar\), and let \(\Phi\) acquire the expectation value \(\Phi_{0}\). Then the gauge fields acquire the mass matrix
whose kernel is spanned by the generators annihilating \(\Phi_{0}\). Hence the gauge bosons of the unbroken subgroup \(H\) stay exactly massless, one massive vector appears for each broken generator, and that is also the number of Goldstone modes removed from the scalar spectrum by Theorem 104.17. Rests on Theorems 104.17 and 104.19.
Derives Theorem 104.22. Setting \(\Phi=\Phi_{0}\) in the scalar kinetic term, \(\left(D_{\mu}\Phi_{0}\right)^{\dagger}\left(D^{\mu}\Phi_{0}\right) =\hbar^{-2}\left(g_{a}T^{a}\Phi_{0}\right)^{\dagger} \left(g_{b}T^{b}\Phi_{0}\right)A^{a}_{\mu}A^{b\mu}\), since \(\pp_{\mu}\Phi_{0}=0\). Only the part symmetric in \(ab\) survives the contraction with \(A^{a}_{\mu}A^{b\mu}\), which is the real part. Comparing with the Proca form \(\tfrac{1}{2\mu_{0}}(c/\hbar)^{2}(M^{2})^{ab}A^{a}_{\mu}A^{b\mu}\) gives Equation (104.25).
The matrix is a Gram matrix of the vectors \(g_{a}T^{a}\Phi_{0}\), so it is positive semi-definite and a vector \(n^{a}\) lies in its kernel if and only if \(n^{a}g_{a}T^{a}\Phi_{0}=0\), that is, if and only if the corresponding generator annihilates the vacuum. The rank of \(M^{2}\) therefore equals the number of independent broken generators, which by Theorem 104.17 is the number of massless scalars the ungauged theory would have had. Each of them is removed from the scalar spectrum by the unitary-gauge argument of Theorem 104.19 applied direction by direction, and reappears as the longitudinal polarization of the corresponding massive vector: the bookkeeping balances generator by generator, not merely in total.
∎Theorem 104.22 is why the photon is exactly massless in the electroweak theory rather than approximately so. If a single generator combination annihilates \(\Phi_{0}\), the corresponding vector has strictly zero mass to all orders — because the unbroken \(\U(1)\) remains a genuine gauge symmetry and Theorem 104.8 still applies to it. The experimental bound on the photon mass, \(m_{\gamma}<1.6\times 10^{-50}\,\mathrm{kg}\) in the laboratory and some four orders of magnitude tighter from solar-system magnetic fields (Phenomenon 100.7), therefore tests a structural feature of the symmetry-breaking pattern and not a fitted parameter.
The Glashow–Weinberg–Salam model
The electroweak Lagrangian
Weinberg's 1967 paper [Weinberg:1967] applies Theorem 104.22 to Glashow's group with the smallest scalar multiplet that will do the job, and Salam presented the same construction the following year [Salam:1968]. The choice of multiplet is not free: it must break \(\SU(2)_{L}\times\U(1)_{Y}\) down to a \(\U(1)\) under which the electron is charged and the neutrino is not, and it must do so leaving exactly one unbroken generator.
Let \(\Phi\) be a complex \(\SU(2)_{L}\) doublet of hypercharge \(Y=1\),
with the Lagrangian density
\(D_{\mu}\) being Equation (104.11) and \(\mu^{2}>0\), \(\lambda>0\). By Equation (104.13) the upper component has \(Q=+1\) and the lower \(Q=0\): the hypercharge assignment \(Y=1\) is exactly what makes the lower component electrically neutral, and hence what makes it able to acquire an expectation value without breaking electromagnetism. The SI dimensions are \([\mu^{2}]=/\mathrm{m}^{2}\) and \([\lambda]=/\mathrm{J}/\mathrm{m}\), so that \(\hat{\lambda}=\lambda\hbar c\) is dimensionless (Equation (106.34)); \(\hbar\mu c\) is an energy.
The whole electroweak Lagrangian is now
each term of dimension \(\mathrm{J}/\mathrm{m}^{3}\): the two gauge kinetic terms as in Equation (104.9), the fermion kinetic terms as in Equation (100.5) but with no mass terms, the scalar sector Equation (104.27), and the Yukawa couplings of Section 104.4.3. The sum over \(\psi\) runs over the chiral fields of Table 104.1, with \(D_{\mu}\) carrying that field's own \(T^{a}\) and \(Y\). There is not one free mass in it.
The potential in Equation (104.27) is minimized on the three-sphere \(\Phi^{\dagger}\Phi=v^{2}/2\) with \(v=\mu/\sqrt{\lambda}\). An \(\SU(2)_{L}\times\U(1)_{Y}\) rotation brings any point of it to
whose stabilizer is generated by \(Q=T^{3}+Y/2\) alone. Hence
three of the four generators are broken, three of the four real scalar fields become the longitudinal polarizations of \(W^{\pm}\) and \(Z^{0}\), one real scalar survives as a particle, and the photon is exactly massless. Rests on Equation (104.27), Theorem 104.17 and Theorem 104.22.
Derives Theorem 104.25. The potential depends on \(\Phi\) only through \(\rho=\Phi^{\dagger}\Phi\), so Proposition 104.15 applies verbatim with \(\abs{\phi}^{2}\) replaced by \(\rho\): the minimum is at \(\rho=\mu^{2}/2\lambda=v^{2}/2\). The set of \(\Phi\) with \(\Phi^{\dagger}\Phi=v^{2}/2\) is a sphere \(S^{3}\) in the four real dimensions of the doublet, and \(\SU(2)\) acts transitively on \(S^{3}\) (it is \(S^{3}\) as a manifold), so every minimum is gauge-equivalent to Equation (104.29).
For the stabilizer, compute the action of each generator on \(\Phi_{0}\). With \(T^{a}=\tfrac{1}{2}\sigma^{a}\) and \(Y=1\),
none of which vanishes: each of the four generators moves the vacuum. But the combination \(Q=T^{3}+Y/2\) does annihilate it, by the last two lines of Equation (104.31), and it is the only combination that does, since the first two lines are linearly independent of one another and of the third. The unbroken subgroup is therefore one-dimensional and generated by \(Q\), which is Equation (104.30); that \(Q\) is the electric charge is Equation (104.13), so what survives is electromagnetism and not some other \(\U(1)\).
By Theorem 104.17 the ungauged model would have \(\dim G-\dim H=4-1=3\) massless scalars, leaving \(4-3=1\) massive one. By Theorem 104.22 the gauge-boson mass matrix has rank \(3\) and a one-dimensional kernel spanned by the generator \(Q\), so exactly one vector — the photon — stays massless, exactly to all orders by Remark 104.23, and three become massive by absorbing the three Goldstone modes.
∎Matching the low-energy limit of \(W\) exchange to the Fermi constant, Equation (104.4), fixes
that is \(E_{v}=3.9449\times 10^{-8}\,\mathrm{J}\), \(v=2.2186\times 10^{5}\,\mathrm{J}^{1/2}\,\mathrm{m}^{-1/2}\), and an associated length
Derivation. Derives Proposition 104.26. The \(W\) mass is computed in Theorem 104.28 below as \(M_{W}c^{2}=\tfrac{1}{2}\hat{g}E_{v}\); substituting it into the matching condition Equation (104.4),
so \(E_{v}^{2}=(\hbar c)^{3}/\sqrt{2}G_{F}\), which is Equation (104.32); the coupling \(\hat{g}\) has cancelled, which is why \(v\) is known far more precisely than either boson mass. Numerically, with \(\sqrt{2}\times1.1663787\times 10^{-5} =1.6495086\times 10^{-5}\) from Equation (104.3), the reciprocal is \(60624.1\,\mathrm{GeV}^{2}\) and its square root is \(246.220\,\mathrm{GeV}\). Converting, \(246.220\,\mathrm{GeV}\times1.602176634\times 10^{-10}\,\mathrm{J}/\mathrm{GeV} =3.94487\times 10^{-8}\,\mathrm{J}\); dividing by \(\sqrt{\hbar c}=1.7780683\times 10^{-13}\,\mathrm{J}^{1/2}\,\mathrm{m}^{1/2}\) gives \(v\); and dividing \(\hbar c=3.161527\times 10^{-26}\,\mathrm{J}\,\mathrm{m}\) by \(E_{v}\) gives Equation (104.33).
∎\(E_{v}\) is not a mass and not the energy of anything. It is the value of a field, converted to an energy by the two constants available; the physical content of Equation (104.33) is that the electroweak condensate has a structure length of \(0.8\,\mathrm{am}\), about a thousandth of a proton radius (\(8.4075\times 10^{-16}\,\mathrm{m}\), [Mohr:2025]). Every mass in the electroweak sector is \(E_{v}\) times a dimensionless coupling, so what the Standard Model explains is mass ratios; the overall scale is one input number, and why it is \(10^{-17}\) of any gravitational scale is Section 104.8.2.
With the vacuum Equation (104.29), the quadratic terms in the gauge fields obtained from \(\left(D_{\mu}\Phi_{0}\right)^{\dagger}\left(D^{\mu}\Phi_{0}\right)\) are
and therefore, defining
the masses
with \(\hat{g}\), \(\hat{g}'\) and \(E_{v}\) as in Notation 104.1. Rests on Equation (104.29), Notation 104.1 and Theorem 104.22.
Derives Theorem 104.28. Since \(\pp_{\mu}\Phi_{0}=0\), \(D_{\mu}\Phi_{0}=\frac{\ii}{\hbar} \left(gW^{a}_{\mu}T^{a}+\tfrac{1}{2}g'B_{\mu}\right)\Phi_{0}\), using \(Y=1\). Then
Insert \(T^{a}=\tfrac{1}{2}\sigma^{a}\) and \(\Phi_{0}=\tfrac{v}{\sqrt{2}}(0,1)\transpose\). The matrix in the middle is
so acting on \(\Phi_{0}\) it returns \(\tfrac{v}{2\sqrt{2}}\left(\sqrt{2}gW^{+}_{\mu}, -gW^{3}_{\mu}+g'B_{\mu}\right)\transpose\), whose squared norm is
which divided by \(\hbar^{2}\) is Equation (104.34). Note that \(B_{\mu}\) and \(W^{3}_{\mu}\) enter only through the single combination \(gW^{3}-g'B\): the orthogonal combination \(A_{\mu}\) of Equation (104.35) is absent, so it has no mass, which is the algebraic form of Theorem 104.22 here.
Reading off the masses requires the Proca normalizations. A complex vector has mass term \(\mu_{0}^{-1}\left(M_{W}c/\hbar\right)^{2}W^{+}_{\mu}W^{-\mu}\) and a real one \(\tfrac{1}{2\mu_{0}}\left(M_{Z}c/\hbar\right)^{2}Z_{\mu}Z^{\mu}\), so
Now convert to the hatted couplings. With \(g=\hat{g}\sqrt{\varepsilon_{0}\hbar c}\) from Equation (104.1) and \(v=E_{v}/\sqrt{\hbar c}\) from Equation (104.2),
because \(c\sqrt{\mu_{0}\varepsilon_{0}}=1\) exactly. The same substitution in \(M_{Z}\) gives the second of Equation (104.36). Dimensionally \(\mu_{0}g^{2}v^{2}/c^{2}\) carries \(\mathrm{kg}\,\mathrm{m}/\mathrm{C}^{2}\times\mathrm{C}^{2} \times\mathrm{J}/\mathrm{m}\times\mathrm{s}^{2}/\mathrm{m}^{2}=\mathrm{kg}^{2}\), as a squared mass must.
∎Writing \(\Phi=\frac{1}{\sqrt{2}}\left(0,v+h\right)\transpose\) in unitary gauge, the surviving real scalar \(h\) has
and the measured \(m_{h}c^{2}=125.20(11)\,\mathrm{GeV}\) [Navas:2024] gives
Rests on Equation (104.14), Equation (104.16) and Notation 106.23.
Derives Corollary 104.29. The potential restricted to the radial direction is Equation (104.14) with \(\abs{\phi}^{2}=(v+h)^{2}/2\), so Equation (104.16) applies unchanged: the \(h\) mass term is \(\tfrac{1}{2}(2\lambda v^{2})h^{2}\), which by Notation 106.23 means \((m_{h}c/\hbar)^{2}=2\lambda v^{2}\). Multiplying by \((\hbar c)^{2}\), \((m_{h}c^{2})^{2}=2(\hbar c)^{2}\lambda v^{2} =2(\lambda\hbar c)(v^{2}\hbar c)=2\hat{\lambda}E_{v}^{2}\), which is Equation (104.37); and \(\mu^{2}=\lambda v^{2}\) gives \(\hbar\mu c=m_{h}c^{2}/\sqrt{2}\). Numerically \(m_{h}c^{2}/E_{v}=125.20/246.22=0.50848\), whose square halved is \(0.12928\), and \(125.20/\sqrt{2}=88.53\).
∎Proposition 104.15 gives the depth of the well as \(V_{\min}=-\lambda v^{4}/4=-\hat{\lambda}E_{v}^{4}/4(\hbar c)^{3}\), which with Equation (104.38) is \(-2.5\times 10^{45}\,\mathrm{J}/\mathrm{m}^{3}\). Gravity responds to absolute energy density, and the observed dark-energy density is of order \(10^{-9}\,\mathrm{J}/\mathrm{m}^{3}\) [Aghanim:2020]: the electroweak condensate alone exceeds it by some \(54\) orders of magnitude. Nothing in this chapter, or anywhere in the Standard Model, cancels that. It is recorded here because it is a consequence of the mechanism just derived and not of anything speculative, and it is taken up as an open problem in What We Observe but Do Not Understand.
Mixing and the weak mixing angle
The rotation Equation (104.35) is a rotation by an angle, and Proposition 101.67 has already carried out the consequences for the fermion couplings. Collecting what Theorem 104.28 adds:
Define \(\theta_{W}\) by
Then
and, eliminating \(\hat{g}\) in favour of \(\hat{e}\) and \(E_{v}\) and using Equation (104.32),
Four measurable quantities — \(\alpha\), \(G_{F}\), \(M_{W}\), \(M_{Z}\) — are thus expressed through three parameters \(\hat{g}\), \(\hat{g}'\), \(v\), so one relation among them is a prediction. Rests on Equation (104.36), Equation (104.32) and Proposition 101.67.
Derives Proposition 104.31. The first two of Equation (104.40) are Equation (104.36) with Equation (104.39). The third is Equation (101.73), whose derivation is Proposition 101.67: substituting Equation (104.35) into Equation (104.11), the coefficient of \(A_{\mu}\) is \(\hat{g}\sin\theta_{W}T^{3} +\hat{g}'\cos\theta_{W}Y/2\), which is proportional to \(Q=T^{3}+Y/2\) if and only if the two coefficients coincide, and then their common value is \(\hat{e}\). For Equation (104.41), square the first relation and use \(\hat{g}\sin\theta_{W}=\hat{e}\): \((M_{W}c^{2})^{2}\sin^{2}\theta_{W} =\tfrac{1}{4}\hat{e}^{2}E_{v}^{2}=\pi\alpha E_{v}^{2}\), since \(\hat{e}^{2}=4\pi\alpha\); and \(E_{v}^{2}=(\hbar c)^{3}/\sqrt{2}G_{F}\) by Equation (104.32). Numerically, \(\pi\times7.2973525643\times 10^{-3}\times60624.1\,\mathrm{GeV}^{2} =1389.8\,\mathrm{GeV}^{2}\).
∎Two numbers can now be extracted from the measured masses.
With \(M_{W}c^{2}=80.369(13)\,\mathrm{GeV}\) and \(M_{Z}c^{2}=91.1880(20)\,\mathrm{GeV}\) [Navas:2024] and \(E_{v}\) from Equation (104.32),
The value of \(\hat{e}\) implied by Equation (104.40) is then \(0.3084\), corresponding to \(\alpha=\hat{e}^{2}/4\pi=1/132.1\) — which is neither \(\alpha(0)=1/137.036\) [Mohr:2025] nor \(\alpha(M_{Z}c^{2}) =1/128.94\) (Phenomenon 100.62), but lies between them. Rests on Equations (104.32) and (104.40).
Derivation. Derives Proposition 104.32. \(\hat{g}=2M_{W}c^{2}/E_{v}=2(80.369)/246.220=0.65282\), and \(\sqrt{\hat{g}^{2}+\hat{g}'^{2}}=2M_{Z}c^{2}/E_{v} =2(91.1880)/246.220=0.74070\), whence \(\hat{g}'^{2}=0.548643-0.426179=0.122464\) and \(\hat{g}'=0.34995\). Then \(\sin^{2}\theta_{W}=\hat{g}'^{2}/(\hat{g}^{2}+\hat{g}'^{2}) =0.122464/0.548643=0.22321\), which is also \(1-(80.369/91.1880)^{2}\) as Equation (104.40) requires. Finally \(\hat{e}=\hat{g}\sin\theta_{W} =0.65282\times0.47245=0.30843\) and \(\hat{e}^{2}/4\pi=0.095129/12.56637=0.0075699 =1/132.10\).
∎Proposition 104.32 is the chapter's first honest failure and it is instructive. Four measured quantities were fitted with three parameters, and the leftover is a \(3.7\) per cent discrepancy in \(\hat{e}^{2}\) — far outside every experimental error. Nothing is wrong: the relations are tree-level, and a one-loop electroweak theory is not a tree-level one. The discrepancy is conventionally packaged as a single correction \(\Delta r\) defined by replacing Equation (104.41) with
and its value follows from the measured masses. Solving Equation (104.43) for \(M_{W}\) with \(\Delta r=0\) and \(\alpha=\alpha(0)\) gives
that is \(M_{W}c^{2}=80.94\,\mathrm{GeV}\), against the measured \(80.369\,\mathrm{GeV}\); matching them requires \(\Delta r=0.036\). Two effects account for most of it, and they act in opposite directions. The running of the electromagnetic coupling from zero momentum transfer to \(M_{Z}c^{2}\) (Phenomenon 100.62), with \(\alpha^{-1}(M_{Z}c^{2})=128.94\), contributes \(\Delta\alpha=1-\alpha(0)/\alpha(M_{Z}c^{2})\approx0.059\), and the top quark contributes \(-\left(\cos^{2}\theta_{W}/\sin^{2}\theta_{W}\right)\Delta\rho \approx-3.48\times0.0093=-0.032\) through Equation (104.52) below; the sum, \(0.027\), lands within about \(0.009\) of the required \(0.036\), the remainder coming from the scalar-mass logarithm, the remaining bosonic self-energies and higher orders. That these two one-loop terms capture the bulk of a \(3.7\) per cent effect — and that the rest is calculable — is the reason the electroweak fit of Section 104.6.4 can measure the top and scalar masses without producing either particle.
Proposition 104.32 defines \(\sin^{2}\theta_{W}\) by the mass ratio. Other definitions are in use and they do not agree, because beyond tree level the angle is a renormalized quantity and a renormalized quantity needs a scheme (Definition 100.54). Three are standard:
-
the on-shell angle, \(s_{W}^{2}:=1-M_{W}^{2}/M_{Z}^{2}=0.2232\), defined to all orders by the physical masses;
-
the \(\overline{\mathrm{MS}}\) angle at the \(Z\) scale, \(\hat{s}^{2}_{Z}=0.23129(4)\), defined by the ratio of running couplings in the scheme of Equation (100.75);
-
the effective leptonic angle \(\sin^{2}\theta^{\ell}_{\mathrm{eff}}=0.23155(4)\), defined so that the measured \(Z\)-pole asymmetries take their tree-level form,
the last two as quoted in the electroweak review of [Navas:2024]. The differences, about one per cent, are hundreds of times the experimental errors, so a number quoted without its scheme is not a measurement of anything. This is why the honest statement of the unification test is Phenomenon 104.55: not that some number equals \(0.23\), but that one running, scheme-defined angle fits every process at every momentum transfer.
The first determination of the angle away from the \(Z\) was Prescott's measurement of the parity-violating asymmetry in the scattering of longitudinally polarized electrons from deuterium at SLAC [Prescott:1978], which gave \(\sin^{2}\theta_{W}=0.20(3)\), sharpened to \(0.224(20)\) by the follow-up scan over beam energies [Prescott:1979]; the measurement belongs to Experiment: Parity Violation. Its significance was not the number but that the number agreed with the one Gargamelle had extracted from neutrino scattering in an entirely different process.
The measured masses of the two weak bosons,
[Navas:2024] determine a mixing angle through \(\cos\theta_{W}=m_{W}/m_{Z}\), and that angle agrees with the one determined independently from the neutral-current couplings of Phenomenon 104.55. Equivalently the parameter
is measured to be unity to within the small radiative corrections [Schael:2006] [Navas:2024]. Three quantities measured in unrelated ways — two masses and a coupling ratio — satisfy one relation that nothing obliged them to satisfy. The statement has content only when the same renormalization scheme defines the angle on both sides: taken as a definition of \(\theta_{W}\) from the masses, \(\rho=1\) says nothing, and what is tested is the agreement of that angle with the one the couplings give. Rests on Equations (104.36) and (104.39).
Derivation. Derives Phenomenon 104.35. The masses are Equation (104.36), so \(m_{W}^{2}/m_{Z}^{2}=\hat{g}^{2}/(\hat{g}^{2}+\hat{g}'^{2}) =\cos^{2}\theta_{W}\) by Equation (104.39), and Equation (104.45) gives \(\rho=1\) identically. Two features of that one-line computation carry the physics.
First, \(E_{v}\) has cancelled. The relation is therefore a statement about the representation the breaking field occupies and not about how strongly the symmetry is broken; it survives any number of doublets and, as Theorem 104.41 shows, fails for almost any other multiplet.
Second, the angle in Equation (104.45) is defined by Equation (104.39) from the couplings, while the ratio on the left is measured from masses. What is tested is that these two determinations agree, and the test is quantitative only when a scheme is fixed, per Remark 104.34. In the on-shell scheme \(\rho=1\) is an identity and the content moves into the agreement of the mass-derived angle with the asymmetry-derived one; in \(\overline{\mathrm{MS}}\) the content stays in \(\rho\) itself, whose measured deviation from unity is the few-per-mille effect predicted by Equation (104.52). The measured \(\rho\) is in either case evidence that the electroweak symmetry is broken by doublets, which is a fact about the vacuum, not a convention.
∎Fermion masses and the Yukawa couplings
Equation (104.28) contains no fermion mass term, and that was not an omission.
For any fermion of Table 104.1, the Dirac mass term \(-mc^{2}\bar{\psi}\psi =-mc^{2}\left(\bar{\psi}_{L}\psi_{R}+\bar{\psi}_{R}\psi_{L}\right)\) is not invariant under \(\SU(2)_{L}\times\U(1)_{Y}\). Rests on Definition 104.10 and Table 104.1.
Derives Proposition 104.36. Insert \(\psi=\psi_{L}+\psi_{R}\) with \(\psi_{L,R}=\tfrac{1}{2}(\identity\mp\gamma^{5})\psi\) and use \(\bar{\psi}_{L}\psi_{L}=\bar{\psi}_{R}\psi_{R}=0\), which follows from \(\gamma^{5}\gamma^{0}=-\gamma^{0}\gamma^{5}\): the mass term connects the two chiralities and nothing else. But \(\psi_{L}\) is a component of an \(\SU(2)_{L}\) doublet while \(\psi_{R}\) is a singlet, so \(\bar{\psi}_{L}\psi_{R}\) transforms as a doublet, not a singlet, and cannot be a term in an invariant Lagrangian. Even setting the \(\SU(2)_{L}\) aside, the hypercharges do not match: for the electron \(Y_{L}=-1\) and \(Y_{R}=-2\), so \(\bar{\psi}_{L}\psi_{R}\) carries \(Y=-2-(-1)=-1\neq0\) and is not \(\U(1)_{Y}\) invariant either. The statement is not that the electron is massless; it is that its mass cannot be a parameter of the symmetric Lagrangian, and must come from the vacuum.
∎Let \(Q_{L}\), \(L_{L}\) be the left-chiral doublets and \(u_{R}\), \(d_{R}\), \(e_{R}\) the singlets, each carrying a generation index \(i=1,2,3\). The most general gauge-invariant, renormalizable coupling to \(\Phi\) is
where
is the conjugate doublet, which transforms as a doublet with \(Y=-1\). The couplings \(y_{ij}\) are arbitrary complex matrices of dimension \(\mathrm{J}^{1/2}\,\mathrm{m}^{1/2}\), so that \(\hat{y}:=y/\sqrt{\hbar c}\) is dimensionless.
Derivation of the dimension and of Equation (104.47). Derives Equation (104.47). A fermion bilinear \(\bar{\psi}\psi\) carries \(/\mathrm{m}^{3}\) by Definition 100.2 and \(\Phi\) carries \(\mathrm{J}^{1/2}\,\mathrm{m}^{-1/2}\), so for Equation (104.46) to be an energy density, \([y]=\mathrm{J}/\mathrm{m}^{3}\times\mathrm{m}^{3} \times\mathrm{J}^{-1/2}\,\mathrm{m}^{1/2} =\mathrm{J}^{1/2}\,\mathrm{m}^{1/2}\), and dividing by \(\sqrt{\hbar c}=\mathrm{J}^{1/2}\,\mathrm{m}^{1/2}\) leaves a pure number. As for \(\tilde{\Phi}\): under \(\Phi\mapsto U\Phi\) with \(U\in\SU(2)\), \(\Phi^{*}\mapsto U^{*}\Phi^{*}\), and the identity \(\sigma^{2}U^{*}\sigma^{2}=U\), valid for every \(\SU(2)\) matrix because the doublet representation is pseudoreal — unitarily equivalent to its own complex conjugate — gives \(\ii\sigma^{2}\Phi^{*}\mapsto U\left(\ii\sigma^{2}\Phi^{*}\right)\). Complex conjugation flips the sign of \(Y\), so \(\tilde{\Phi}\) has \(Y=-1\), and \(\bar{Q}_{L}\tilde{\Phi}u_{R}\) carries \(Y=-\tfrac{1}{3}-1+\tfrac{4}{3}=0\) as required, while \(\bar{Q}_{L}\Phi d_{R}\) carries \(Y=-\tfrac{1}{3}+1-\tfrac{2}{3}=0\). Without \(\tilde{\Phi}\) the up-type quarks could not be given mass by this doublet at all.
∎Substituting Equation (104.29) into Equation (104.46) produces mass matrices \(m^{f}_{ij}c^{2}=\hat{y}^{f}_{ij}E_{v}/\sqrt{2}\). Each is a general complex matrix and is brought to non-negative diagonal form by a biunitary transformation,
with \(U_{f}\) acting on left-chiral and \(V_{f}\) on right-chiral fields. In the mass basis so defined,
-
the charged current acquires the matrix \(V_{\mathrm{CKM}}=U_{u}^{\dagger}U_{d}\), which is unitary but not in general diagonal;
-
the neutral currents — both the photon's and the \(Z\)'s — remain exactly diagonal in flavour;
-
the couplings of \(h\) to fermions remain diagonal, and proportional to the masses.
Rests on Equation (104.29), Equation (104.46) and Theorem 101.59.
Derives Theorem 104.38. With \(\Phi\to\Phi_{0}\) the term \(y^{d}_{ij}\bar{Q}_{Li}\Phi d_{Rj}\) becomes \(\left(y^{d}_{ij}v/\sqrt{2}\right)\bar{d}_{Li}d_{Rj}\), which compared with \(-m c^{2}\bar{\psi}_{L}\psi_{R}\) identifies \(m^{d}_{ij}c^{2}=\hat{y}^{d}_{ij}\sqrt{\hbar c}\,v/\sqrt{2} =\hat{y}^{d}_{ij}E_{v}/\sqrt{2}\); likewise for \(u\) and \(e\), the conjugate doublet supplying \(\Phi^{0*}\to v/\sqrt{2}\) for the up type.
For Equation (104.48): given any complex \(m\), the matrix \(mm^{\dagger}\) is Hermitian positive semi-definite, hence \(mm^{\dagger}=U\Sigma^{2}U^{\dagger}\) with \(U\) unitary and \(\Sigma\geq0\) diagonal, by the spectral theorem for normal operators. Define \(V:=m^{\dagger}U\Sigma^{-1}\) where \(\Sigma\) is invertible; then \(V^{\dagger}V=\Sigma^{-1}U^{\dagger}mm^{\dagger}U\Sigma^{-1} =\Sigma^{-1}\Sigma^{2}\Sigma^{-1}=\identity\), so \(V\) is unitary, and \(U^{\dagger}mV=U^{\dagger}mm^{\dagger}U\Sigma^{-1} =\Sigma\). (If \(\Sigma\) has zero entries, restrict to its support and extend \(V\) arbitrarily unitarily on the complement.) This is the singular-value decomposition, and the two unitaries are independent precisely because the left- and right-chiral fields are independent — which is possible only because the theory is chiral.
(i) The charged current is \(\bar{u}_{Li}\gamma^{\mu}d_{Li}\) in the gauge basis. Substituting \(u_{L}=U_{u}u_{L}'\) and \(d_{L}=U_{d}d_{L}'\) with primes denoting mass eigenstates gives \(\bar{u}'_{L}U_{u}^{\dagger}\gamma^{\mu}U_{d}d'_{L} =\bar{u}'_{L}\gamma^{\mu}\left(U_{u}^{\dagger}U_{d}\right)d'_{L}\), which is Equation (101.67) with \(V_{\mathrm{CKM}}=U_{u}^{\dagger}U_{d}\), manifestly unitary as the product of two unitaries — which is Theorem 101.62 and is what Linear Algebra and Representation Theory licenses.
(ii) A neutral current is \(\bar{f}_{Li}\gamma^{\mu}\left(\text{coefficient}\right)f_{Li}\) with a coefficient depending only on \(T^{3}\) and \(Q\) (Equation (101.75)), hence the same for every generation; it is therefore proportional to the identity matrix in generation space, and \(U_{f}^{\dagger}\identity U_{f}=\identity\). Nothing survives the rotation. This is the Glashow–Iliopoulos–Maiani mechanism [Glashow:1970] in its general form, proved as Theorem 101.59; note that the argument needs the coefficient to depend only on the gauge quantum numbers, which fails as soon as two fields with the same \(T^{3}\) and \(Q\) but different masses can mix — the reason a second scalar doublet, or an extra quark that is not part of a complete doublet, generically produces flavour-changing neutral currents and is excluded by the kaon and \(B\)-meson data of Experiment: CP Violation.
(iii) The Yukawa term with \(\Phi=\frac{1}{\sqrt{2}}(0,v+h)\transpose\) is the mass term multiplied by \((1+h/v)\), so the coupling matrix of \(h\) is \(m^{f}/v\) — the same matrix, diagonalized by the same \(U_{f}\), \(V_{f}\). Hence \(h\) has no flavour-changing couplings and its coupling to each fermion is \(m_{f}c^{2}/E_{v}\) per unit field, which is the statement measured in Phenomenon 104.62.
∎| Fermion | $m_{f}c^{2}$ [\(\mathrm{GeV}\)] | $\hat{y}_{f}$ |
|---|---|---|
| $t$ | \(172.57\) | \(0.991\) |
| $b$ | \(4.183\) | \(0.0240\) |
| $\tau$ | \(1.77693\) | \(0.0102\) |
| $c$ | \(1.273\) | \(0.00731\) |
| $\mu$ | \(0.1056584\) | \(6.07\times 10^{-4}\) |
| $s$ | \(0.0935\) | \(5.37\times 10^{-4}\) |
| $d$ | \(0.00470\) | \(2.70\times 10^{-5}\) |
| $u$ | \(0.00216\) | \(1.24\times 10^{-5}\) |
| $e$ | \(5.109990\times 10^{-4}\) | \(2.93\times 10^{-6}\) |
Table 104.2 spans a factor of \(3.4\times 10^{5}\) from top to electron. Exactly one entry is of order one, and it is the top's: \(\hat{y}_{t}=0.991\), within one per cent of unity. Whether that is a coincidence is not known. What is known is that it makes the top quark the dominant virtual particle in every electroweak loop, which is why \(\Delta\rho\) in Equation (104.52) goes as \(m_{t}^{2}\), why the vacuum stability analysis of Section 104.8.1 turns on the top mass at the \(1\,\mathrm{GeV}\) level, and why the electroweak fit could locate the top before it was produced. The spread itself is an input, not a prediction: Table 104.2 lists numbers the theory takes from experiment, and What We Observe but Do Not Understand records that no evidence-based explanation of the pattern exists. Neutrino masses are not in the table at all, because Equation (104.46) cannot generate them without a right-chiral neutrino; that the masses are nonetheless non-zero (Flavour Physics and Neutrinos and Experiment: Neutrino Oscillations) is the one laboratory fact this Lagrangian does not accommodate.
Custodial symmetry and the $\rho$ parameter
\(\rho=1\) came out of Phenomenon 104.35 as an accident of the doublet. It is not an accident: it follows from a global symmetry of the scalar sector that the gauge couplings do not respect, which is why the corrections to it are small and calculable.
Assemble the doublet and its conjugate into the \(2\times2\) matrix
Then \(\Phi^{\dagger}\Phi=\tfrac{1}{2}\tr \left(\mathcal{M}^{\dagger}\mathcal{M}\right)\), so the scalar potential in Equation (104.27) is invariant under the global group
larger than the gauged \(\SU(2)_{L}\times\U(1)_{Y}\). The vacuum \(\mathcal{M}_{0}=\left(v/\sqrt{2}\right)\identity\) is invariant under the diagonal subgroup \(L=R\), the custodial \(\SU(2)_{V}\), under which \((W^{1},W^{2},W^{3})\) transforms as a triplet. Hence in the limit \(g'\to0\) the three fields have equal masses and \(\rho=1\) exactly. Rests on Equations (104.27), (104.35) and (104.45).
Derives Theorem 104.40. That \(\Phi^{\dagger}\Phi=\tfrac{1}{2}\tr(\mathcal{M}^{\dagger} \mathcal{M})\) is a computation: \(\tr(\mathcal{M}^{\dagger}\mathcal{M}) =2\left(\abs{\Phi^{0}}^{2}+\abs{\Phi^{+}}^{2}\right)\), since the two columns of \(\mathcal{M}\) have the same norm. A trace of the form \(\tr(\mathcal{M}^{\dagger}\mathcal{M})\) is invariant under Equation (104.50) for any unitary \(L\), \(R\), and the potential depends on \(\mathcal{M}\) only through it. The kinetic term is
the relative sign in the last term coming from the opposite hypercharge of \(\tilde{\Phi}\). The \(\SU(2)_{L}\) term is invariant under \(\SU(2)_{R}\), since \(R\) acts on the right; the \(\U(1)_{Y}\) term is not, because only \(T^{3}\) of \(\SU(2)_{R}\) is gauged. So \(\SU(2)_{R}\) is a symmetry of the whole Lagrangian when \(g'=0\) and of the scalar potential always.
At the vacuum, \(\Phi^{+}=0\) and \(\Phi^{0}=v/\sqrt{2}\), so \(\mathcal{M}_{0}=\left(v/\sqrt{2}\right)\identity\) and \(L\mathcal{M}_{0}R^{\dagger}=\mathcal{M}_{0}\) if and only if \(LR^{\dagger}=\identity\), that is \(L=R\): the unbroken global group is the diagonal \(\SU(2)_{V}\). Under it the gauge fields \(W^{a}_{\mu}\), which carry an adjoint index of \(\SU(2)_{L}\), transform as an \(\SU(2)_{V}\) triplet, and a mass matrix invariant under a group that acts irreducibly on a triplet must be proportional to the identity by Schur's lemma (Linear Algebra and Representation Theory). Hence \(M_{W^{1}}=M_{W^{2}}=M_{W^{3}}\) when \(g'=0\). Reinstating \(g'\) mixes \(W^{3}\) with \(B\) but changes no mass in the \(W\) sector, so \(M_{W}=M_{W^{3}}=M_{Z}\cos\theta_{W}\) by Equation (104.35), and Equation (104.45) gives \(\rho=1\).
∎Let the symmetry be broken by several complex scalar multiplets labelled \(i\), the \(i\)-th having weak isospin \(T_{i}\) and acquiring an expectation value \(v_{i}/\sqrt{2}\) in the component of third component \(T^{3}_{i}\), that component being electrically neutral. Then at tree level
For doublets (\(T=\tfrac{1}{2}\), \(T^{3}=-\tfrac{1}{2}\)) this is \(1\) for any number of them; for a complex triplet with \(T^{3}=\pm1\) it is \(\tfrac{1}{2}\); for a triplet breaking through its \(T^{3}=0\) component the denominator vanishes. The measurement \(\rho=1\) is thus evidence about the multiplet structure of the breaking sector [Ross:1975] [Sikivie:1980]. Rests on Theorem 104.28 and Equation (104.45).
Derives Theorem 104.41. Repeat the computation of Theorem 104.28 for a general multiplet. Writing \(T^{\pm}=T^{1}\pm\ii T^{2}\) and \(W^{\pm}_{\mu}=(W^{1}_{\mu}\mp\ii W^{2}_{\mu})/\sqrt{2}\), one has \(T^{1}W^{1}_{\mu}+T^{2}W^{2}_{\mu} =\tfrac{1}{\sqrt{2}}\left(T^{+}W^{+}_{\mu}+T^{-}W^{-}_{\mu}\right)\), so the part of \(\hbar^{-2}\abs{gW^{a}_{\mu}T^{a}\Phi_{0}}^{2}\) containing \(W^{+}_{\mu}W^{-\mu}\) is
using \(\acomm{T^{+}}{T^{-}}=2\left[(T^{1})^{2}+(T^{2})^{2}\right] =2\left[\vect{T}^{2}-(T^{3})^{2}\right]\) and the eigenvalue \(T(T+1)\) of the Casimir (Lie Groups, Lie Algebras, and Fibre Bundles). With \(\Phi_{0}^{\dagger}\Phi_{0}=v_{i}^{2}/2\) in the \(i\)-th multiplet, summing over multiplets and matching to the Proca form gives
For the neutral sector, the coefficient of the vacuum in the neutral covariant derivative is \(gT^{3}W^{3}_{\mu}+g'\left(Y/2\right)B_{\mu}\) with \(Y/2=Q-T^{3}\), and \(Q\Phi_{0}=0\) because the component acquiring the expectation value is neutral; hence the coefficient is \(\left(gW^{3}_{\mu}-g'B_{\mu}\right)T^{3}\), and
Dividing, with \(\cos^{2}\theta_{W}=g^{2}/\left(g^{2}+g'^{2}\right)\), every coupling cancels and Equation (104.51) remains. Substituting \(T=\tfrac{1}{2}\), \(T^{3}=-\tfrac{1}{2}\) gives numerator \(\left(\tfrac{3}{4}-\tfrac{1}{4}\right)v^{2}\) and denominator \(2\cdot\tfrac{1}{4}v^{2}\), equal; substituting \(T=1\), \(T^{3}=-1\) gives \(\left(2-1\right)v^{2}\) over \(2v^{2}\), that is \(\tfrac{1}{2}\); and \(T=1\), \(T^{3}=0\) gives \(2v^{2}\) over \(0\).
∎Custodial symmetry is broken by two things in the real theory, and each produces a calculable shift.
\(\SU(2)_{V}\) is violated (i) by the hypercharge coupling \(g'\), which gauges only \(T^{3}\) of \(\SU(2)_{R}\), and (ii) by the Yukawa matrices, which are \(\SU(2)_{R}\) symmetric only if \(y^{u}=y^{d}\) within each doublet. Since \(\hat{y}_{t}=0.991\) and \(\hat{y}_{b}=0.0240\) by Table 104.2, the third generation violates it maximally, and it dominates. Rests on Theorem 104.40 and Equation (104.46).
Derives Proposition 104.42. (i) is proved inside Theorem 104.40: the \(B_{\mu}\) term carries \(\mathcal{M}T^{3}\) on the right and is invariant only under the \(\U(1)\) subgroup of \(\SU(2)_{R}\) generated by \(T^{3}\). (ii) follows by writing Equation (104.46) in the \(\mathcal{M}\) notation: \(\bar{Q}_{L}\mathcal{M}\, \diag(y^{u},y^{d})\,Q_{R}\) with \(Q_{R}=(u_{R},d_{R})\transpose\) is invariant under \(\mathcal{M}\mapsto\mathcal{M}R^{\dagger}\), \(Q_{R}\mapsto RQ_{R}\) if and only if \(\diag(y^{u},y^{d})\) commutes with every \(R\in\SU(2)\), that is if and only if \(y^{u}=y^{d}\).
∎At one loop the leading custodial violation from the third generation is
for \(m_{t}c^{2}=172.57(29)\,\mathrm{GeV}\) [Navas:2024], the omitted terms being the bottom mass (negligible) and the logarithmic scalar contribution. The dependence is quadratic in the top mass, and only logarithmic in \(m_{h}\) — the screening theorem of Veltman [Veltman:1977]. Rests on Proposition 104.42 and Equation (104.32).
Derivation of the parametric form. Derives Proposition 104.43. \(\Delta\rho\) measures the difference between the \(W\) and \(Z\) self-energies at zero momentum, \(\Delta\rho=\Pi_{WW}(0)/M_{W}^{2}c^{2}-\Pi_{ZZ}(0)/M_{Z}^{2}c^{2}\), each computed from a quark loop. By Proposition 104.42 the difference vanishes when \(m_{t}=m_{b}\), so it is proportional to the isospin splitting. Each self-energy carries two Yukawa or gauge vertices and one loop factor \(1/16\pi^{2}\) (Equation (100.82) shows the same factor in QED), and the only dimensionful quantities available are the quark masses and \(E_{v}\); since \(\Delta\rho\) is dimensionless and the vertices are the Yukawa couplings \(\hat{y}\propto m_{f}c^{2}/E_{v}\), the leading term must be \(N_{c}\left(m_{t}c^{2}\right)^{2}/16\pi^{2}E_{v}^{2}\) times a pure number, with \(N_{c}=3\) for colour. Writing \(E_{v}^{2}=(\hbar c)^{3}/\sqrt{2}G_{F}\) turns this into Equation (104.52) up to the pure number, which the explicit loop integration of [Veltman:1977] fixes to exactly \(1\): Equation (104.52) is \(N_{c}\left(m_{t}c^{2}\right)^{2}/16\pi^{2}E_{v}^{2}\). Numerically, \(3\times1.1663787\times 10^{-5}\times(172.57)^{2} =1.0421\) and \(8\sqrt{2}\pi^{2}=111.66\), giving \(9.33\times 10^{-3}\). That the answer grows with \(m_{t}^{2}\) rather than \(\ln m_{t}\) is what made the top mass predictable from LEP data before the Tevatron produced it [Abachi:1995] [Abe:1995].
∎The one-loop coefficient in the top contribution to the rho parameter: the \(W\) and \(Z\) vacuum-polarization integrals with a \(t\)–\(b\) doublet in the loop, regulated dimensionally, whose difference at zero momentum transfer gives the factor \(3/8\sqrt{2}\pi^{2}\) quoted from Veltman's paper above; and the accompanying logarithmic dependence on the scalar mass, which is what the screening theorem asserts. Belongs in Appendix A.
Corrections that enter only through the gauge-boson self-energies — which is where any heavy new state that does not couple directly to light fermions must show up — are parametrized by three numbers \(S\), \(T\), \(U\) [Peskin:1990], normalized to vanish for the Standard Model at a reference point. \(T\) is custodial violation, \(\alpha T=\Delta\rho\); \(S\) measures the difference between the self-energies at \(q^{2}=M_{Z}^{2}c^{2}\) and at \(q^{2}=0\); \(U\) is rarely significant and is usually fixed to zero. Their measured values are consistent with zero at the level of a few hundredths [Navas:2024], which is a quantitative statement that no heavy particle with electroweak quantum numbers has yet made itself felt indirectly.
Renormalizability
't Hooft's proofs
Between 1967 and 1971 the Weinberg–Salam model was cited about as often as a paper is cited when nobody believes it matters. What changed was not a measurement: 't Hooft proved that the theory makes finite predictions at every order. He did it in two papers, and the citation keys used in this book do not follow the order of publication — [tHooft:1971b] is the earlier, on massless Yang–Mills fields, and [tHooft:1971a] is the later, on the spontaneously broken case; both entries in the bibliography record the fact.
The massless paper establishes that a pure Yang–Mills theory, gauge fixed by the Faddeev–Popov procedure of Theorem 106.52 with its ghosts, has divergences that are absorbed into a redefinition of the finitely many parameters already present. The broken paper is the one this chapter needs, and its instrument is a family of gauges interpolating between two descriptions each of which makes one property manifest and the other obscure.
Take the abelian Higgs model of Theorem 104.19 and, instead of the unitary gauge, write \(\phi=\frac{1}{\sqrt{2}}\left(v+h+\ii\chi\right)\) with \(h\) and \(\chi\) both retained. Then
-
the quadratic Lagrangian contains the mixing term \(\left(qv/\hbar\right)A^{\mu}\pp_{\mu}\chi\), which prevents the propagators from being read off;
-
adding the gauge-fixing term
\begin{equation}\tag{104.53} \Lag_{\mathrm{gf}}=-\frac{1}{2\mu_{0}\xi} \left(\pp^{\mu}A_{\mu}-\xi\mu_{0}\frac{qv}{\hbar}\chi\right)^{2} \end{equation}cancels the mixing exactly and gives \(\chi\) the mass \(\sqrt{\xi}\,m_{A}\);
-
the gauge propagator becomes
\begin{equation}\tag{104.54} \widetilde{D}_{\mu\nu}(k) =\frac{-\ii\mu_{0}\hbar^{3}c}{k^{2}-m_{A}^{2}c^{2}+\ii\epsilon} \left[\eta_{\mu\nu}-\left(1-\xi\right) \frac{k_{\mu}k_{\nu}}{k^{2}-\xi m_{A}^{2}c^{2}}\right]\ec \end{equation}which falls as \(k^{-2}\) for every finite \(\xi\);
-
the Faddeev–Popov ghost acquires the same mass \(\sqrt{\xi}\,m_{A}\) and does not decouple.
The limit \(\xi\to\infty\) returns the unitary gauge and Equation (101.77); \(\xi=1\) is the 't Hooft–Feynman gauge, in which the propagator is simply \(-\ii\mu_{0}\hbar^{3}c\,\eta_{\mu\nu}/(k^{2}-m_{A}^{2}c^{2})\). Rests on Theorem 104.19, Equation (104.22) and Equation (101.77).
Derives Theorem 104.45. (i) Expanding \(\abs{D_{\mu}\phi}^{2}\) with \(D_{\mu}=\pp_{\mu}+\ii qA_{\mu}/\hbar\) and \(\phi=\frac{1}{\sqrt{2}}(v+h+\ii\chi)\),
whose last bracket contains, at quadratic order, exactly \((qv/\hbar)A^{\mu}\pp_{\mu}\chi\). A term linear in \(A\) and linear in \(\chi\) makes the quadratic form non-diagonal in field space.
(ii) Expand Equation (104.53):
The middle term integrates by parts to \(-(qv/\hbar)A^{\mu} \pp_{\mu}\chi\) up to a total derivative, cancelling (i). The last term is a mass term for \(\chi\); comparing with \(-\tfrac{1}{2}(m_{\chi}c/\hbar)^{2}\chi^{2}\) and using \(\mu_{0}q^{2}v^{2}/\hbar^{2}=(m_{A}c/\hbar)^{2}\) from Equation (104.22) gives \(m_{\chi}=\sqrt{\xi}\,m_{A}\). The would-be Goldstone boson thus has a gauge-dependent mass, which is the sharpest possible statement that it is not a particle.
(iii) With the mixing gone, the gauge kernel in momentum space (Fourier convention Equation (100.14)) is
which on the transverse projector \(P^{\mathrm{T}}_{\mu\nu}=\eta_{\mu\nu}-k_{\mu}k_{\nu}/k^{2}\) is \(-(k^{2}-m_{A}^{2}c^{2})/\mu_{0}\hbar^{2}\) and on the longitudinal projector \(P^{\mathrm{L}}\) is \(-(k^{2}-\xi m_{A}^{2}c^{2})/\mu_{0}\hbar^{2}\xi\). Inverting on each eigenspace and multiplying by \(\ii\hbar c\), exactly as in Proposition 100.8, gives
and combining the two fractions over the common denominator gives Equation (104.54), since \(-\left(k^{2}-\xi m_{A}^{2}c^{2}\right)+\xi\left(k^{2} -m_{A}^{2}c^{2}\right)=k^{2}\left(\xi-1\right)\). As \(\xi\to\infty\) the bracket tends to \(\eta_{\mu\nu}-k_{\mu}k_{\nu}/m_{A}^{2}c^{2}\), which is Equation (101.77): the unitary gauge is the infinite-\(\xi\) member of the family, and \(\chi\) becomes infinitely heavy and drops out, which is why it was absent there.
(iv) The Faddeev–Popov determinant is that of \(\delta G/\delta\Lambda\) for \(G\) the bracket in Equation (104.53). Under a gauge transformation of parameter \(\Lambda\), \(\delta A_{\mu}=\pp_{\mu}\Lambda\) and, from \(\phi\mapsto\ee^{-\ii q\Lambda/\hbar}\phi\) to first order, \(\delta\chi=-q\Lambda\left(v+h\right)/\hbar\). Hence
so the ghost has mass \(\sqrt{\xi}\,m_{A}\) and a coupling to \(h\). In the abelian unbroken theory this operator is field-independent and the ghosts may be discarded (Corollary 106.53); here they may not.
∎Theorem 104.45 is the technical heart of the renormalizability proof. At \(\xi\to\infty\) the field content is manifestly physical — one massive vector with three polarizations, one scalar, no ghosts, no Goldstone — so unitarity is evident and power counting fails (Theorem 104.4(ii)). At finite \(\xi\) every propagator falls as \(k^{-2}\), so power counting works exactly as in Theorem 100.17 and the theory is renormalizable by inspection, but the spectrum now contains a Goldstone of mass \(\sqrt{\xi}m_{A}\) and a ghost of the same mass, neither of which is a particle. Since \(\xi\) is an arbitrary parameter, no observable can depend on it, and Theorem 106.61 proves that none does — so the two descriptions compute the same numbers, and the theory is unitary and renormalizable although no single gauge makes both obvious. The masses coinciding is not a coincidence: the Goldstone and the ghost are two members of a BRST quartet (Proposition 106.60) and cancel each other in every physical cut.
Let \(\mathcal{M}^{\mu_{1}\dots\mu_{n}}\) be an amplitude with \(n\) external gauge bosons of mass \(M\) and momenta \(k_{i}\), and let \(\mathcal{M}_{\chi}\) be the corresponding amplitude with each gauge boson replaced by its would-be Goldstone boson. Then at energies \(E\gg Mc^{2}\),
Rests on Theorem 104.45, Equation (106.76) and Equation (104.6).
Derivation. Derives Proposition 104.47. The Slavnov–Taylor identity Equation (106.76), applied to the broken theory in an \(R_{\xi}\) gauge, says that the BRST variation of any correlator vanishes. Applied to a Green function with one gauge leg and the ghost of that leg, it gives the relation between the longitudinal part of the gauge amplitude and the Goldstone amplitude,
which is the broken-theory counterpart of the transversality \(k_{\mu}\mathcal{M}^{\mu}=0\) of Equation (100.18) — the right-hand side is zero only when \(M=0\). Now use Equation (104.6): the longitudinal polarization is \(\varepsilon_{L}^{\mu}=k^{\mu}/Mc+v^{\mu}\) with \(v^{\mu}\) of order \(Mc/\abs{\vect{k}}\) and bounded. Contracting,
and iterating over the \(n\) legs gives Equation (104.55). (The identity holds with \(\mathcal{M}_{\chi}\) computed in the same \(R_{\xi}\) gauge, and the corrections are uniform in the scattering angle away from the forward region.)
∎An amplitude with longitudinal gauge bosons appeared, in the unitary-gauge computation, to grow like \((E/Mc^{2})^{n}\) because each polarization vector grows. Proposition 104.47 says that the whole amplitude equals a scalar amplitude with no growing polarization vectors at all. The individual diagrams therefore must cancel among themselves down to the size of the scalar amplitude, and that cancellation is not optional — it is forced by the Slavnov– Taylor identity, which is forced by the gauge symmetry. It will be used twice: in Theorem 104.51 to bound the scalar mass, and in Phenomenon 104.57 to read the non-abelian group structure off a measured cross-section.
The identities that organize the cancellations to all orders are the non-abelian generalization of the Ward–Takahashi identity of Theorem 100.48, found by Slavnov [Slavnov:1972] and Taylor [Taylor:1971], and their modern algebraic form is the BRST master equation Equation (106.76) of Section 106.4.2.
No new measurement stood between the general indifference to [Weinberg:1967] in 1968 and its status as the consensus theory in 1972. What changed was 't Hooft's proof [tHooft:1971a] [tHooft:1972]. This is worth recording plainly in a book organized around evidence: the criterion applied here was internal consistency, and it was applied before the decisive experiments — neutral currents in 1973, the bosons in 1983 — which the theory then survived. Consistency is not evidence, and the chapter's evidential claims all rest on Section 104.6; but a theory that cannot be computed cannot be tested, and renormalizability is what made the tests possible.
Dimensional regularization and gauge cancellations
The regulator used throughout is the one 't Hooft and Veltman introduced for this problem [tHooft:1972] and which Definition 100.30 states: continue the loop integral to \(d=4-2\varepsilon\) dimensions. Its virtue over a momentum cutoff is precisely that it preserves gauge invariance — a cutoff \(\abs{\ell}<\Lambda\) is not invariant under a gauge transformation that shifts the loop momentum, and the resulting non-invariant counterterms would include exactly the photon mass term that Theorem 104.8 forbids. The systematic treatment of the scheme dependence belongs to The Renormalization Group.
Dimensional regularization has a defect specific to a chiral theory. \(\gamma^{5}\) is defined by \(\gamma^{5}=\ii\gamma^{0}\gamma^{1}\gamma^{2}\gamma^{3}\), a four-dimensional construction, and there is no continuation to \(d\neq4\) that is simultaneously anticommuting with all \(\gamma^{\mu}\) and consistent with the cyclicity of the trace. The standard resolution splits the Dirac algebra into four-dimensional and \((d-4)\)-dimensional parts, at the cost of extra non-invariant counterterms that must be fixed by imposing the Slavnov–Taylor identities by hand. This is not a hidden inconsistency — the identities can be restored order by order — but it is real work, and it is the reason the anomaly of Section 104.5.3 shows up in this scheme as an unavoidable obstruction rather than as an artefact one could regulate away.
The most consequential gauge cancellation in the theory is the one that made the scalar mandatory before it was found.
Consider \(W_{L}^{+}W_{L}^{-}\to W_{L}^{+}W_{L}^{-}\) at \(\sqrt{s}\,c\gg M_{W}c^{2}\). Without the scalar, the \(J=0\) partial wave grows as \(s\) and violates unitarity. With the scalar of mass \(m_{h}\) included, the growth cancels and the surviving constant amplitude is
so that the requirement \(\abs{\Re a_{0}}\leq\tfrac{1}{2}\) gives
improved to \(713\,\mathrm{GeV}\) when the coupled channels \(W_{L}W_{L}\), \(Z_{L}Z_{L}\), \(hh\), \(hZ_{L}\) are diagonalized [Lee:1977b] (the result was announced in a companion letter the same year). The measured \(m_{h}c^{2}=125.20(11)\,\mathrm{GeV}\) satisfies it with room to spare. Rests on Proposition 104.47, Equation (104.37) and Equation (101.89).
Derives Theorem 104.51. By Proposition 104.47 the high-energy amplitude equals that for the eaten Goldstone bosons, \(\chi^{+}\chi^{-}\to\chi^{+}\chi^{-}\), computed from the scalar sector alone: the gauge couplings drop out of the leading behaviour entirely. Write the doublet as
and expand \(-\lambda(\Phi^{\dagger}\Phi)^{2}\). Two vertices matter: the contact term \(-\lambda\abs{\chi^{+}}^{4}\) and the coupling \(-2\lambda v\,h\abs{\chi^{+}}^{2}\), in which \(2\lambda v=\left(m_{h}c/\hbar\right)^{2}/v\) by Equation (104.37). Summing the contact term with \(h\) exchange in the \(s\) and \(t\) channels, and using the Mandelstam conventions of Notation 100.1 in which \(s\) and \(t\) are squared momenta,
the factor \(\hbar^{2}\) being the dimension an amplitude carries in the normalization of Proposition 100.24, and \((m_{h}c^{2})^{2}/E_{v}^{2}=2\hat{\lambda}\) by Equation (104.37). The two propagator poles have absorbed the contact term: each channel contributes the vertex \(2\lambda v\) twice with a propagator between them, and the leftover constant is exactly the quartic. For \(s\gg m_{h}^{2}c^{2}\), and away from the forward region where \(\abs{t}\) is also large, both brackets tend to \(1\) and the amplitude tends to the constant \(-2\hbar^{2}(m_{h}c^{2})^{2}/E_{v}^{2}\). That it is constant rather than growing is the cancellation: had the \(h\) been absent, only the contact term would survive and the bracket would be \(-s/m_{h}^{2}c^{2}\) in disguise — growth without limit.
The \(J=0\) projection is \(a_{0}=\frac{1}{32\pi\hbar^{2}}\int_{-1}^{1}\mathcal{M}\, \dd\!\left(\cos\theta\right)\), the normalization being the one in which \(\sigma_{J}=16\pi\hbar^{2}(2J+1)\abs{a_{J}}^{2}/s\) reproduces Equation (101.89). For a constant amplitude the integral is \(2\mathcal{M}\), so \(a_{0}=\mathcal{M}/16\pi\hbar^{2}\), that is
manifestly dimensionless; and \(E_{v}^{2}=(\hbar c)^{3}/\sqrt{2}G_{F}\) turns it into Equation (104.56). Imposing \(\abs{a_{0}}\leq\tfrac{1}{2}\),
whose square root is \(873\,\mathrm{GeV}\). The coupled-channel analysis replaces the coefficient \(1\) by \(3/2\) in the largest eigenvalue, giving \(m_{h}c^{2}\leq\left(4\sqrt{2}\pi(\hbar c)^{3}/3G_{F}\right)^{1/2} =713\,\mathrm{GeV}\).
∎Theorem 104.51 is a prediction of a very unusual kind: it does not say what exists, it says that something must, below about a \(\mathrm{TeV}\), or the theory is inconsistent there. Either a scalar with the couplings the mechanism requires, or a strongly interacting sector in which \(W_{L}W_{L}\) scattering saturates the bound and the perturbative description fails. The two alternatives are distinguishable, and the LHC was built with enough energy to settle it. It found the first (Experiment: The Higgs Boson Discovery), and the couplings match (Phenomenon 104.62). Had it found the second, this chapter would look very different; had it found neither, the Standard Model would have been falsified as a description of energies above a \(\mathrm{TeV}\). That is the sense in which Equation (104.57) was a genuine risk.
Anomaly cancellation
A symmetry of the classical action need not survive quantization. The axial current of a massless Dirac field is conserved classically and not quantum mechanically, by the triangle anomaly of Adler [Adler:1969] and of Bell and Jackiw [Bell:1969], derived in Theorem 106.87 and measured in the \(\pi^{0}\) lifetime (Section 106.7.2). In a global symmetry that is a prediction. In a gauge symmetry it is fatal, because the Slavnov–Taylor identities on which Theorem 104.45 and Remark 104.46 depend then fail: the unphysical polarizations stop decoupling, the \(S\)-matrix stops being unitary and the divergences stop organizing themselves. That is Theorem 106.94.
The electroweak theory is chiral by construction — \(T^{a}\) acts on \(\psi_{L}\) and not on \(\psi_{R}\) — so it is exactly the case in which the anomaly can occur. The condition for it not to is Theorem 106.95: the symmetrized trace \(\mathsf{A}^{abc}=\tr(\acomm{T^{a}}{T^{b}}T^{c})_{L} -\tr(\acomm{T^{a}}{T^{b}}T^{c})_{R}\) must vanish for every triple of gauge generators. Bouchiat, Iliopoulos and Meyer [Bouchiat:1972] checked it for this theory and found that it holds — but only for a complete generation, and only with three colours.
With the hypercharges of Table 104.1, in the convention \(Q=T^{3}+Y/2\), the four independent conditions are satisfied:
where the multiplicities count colour and \(\SU(2)\) components and the first bracket in each of the last two is the left-chiral contribution, the second the right-chiral one. Rests on Table 104.1 and Theorem 106.95.
Derives Proposition 104.53. \(\left[\SU(2)\right]^{3}\) vanishes identically because the symmetrized trace of three \(\SU(2)\) generators is zero in every representation, \(\SU(2)\) having only real and pseudoreal representations; and \(\left[\SU(3)\right]^{3}\) vanishes because the quarks appear in triplet and antitriplet with respect to colour once the two chiralities are combined. What remains is the four displayed conditions.
For Equation (104.58): only \(\SU(2)_{L}\) doublets contribute, each with \(\tr(\acomm{T^{a}}{T^{b}})=\delta^{ab}\), so the condition is \(\sum_{\text{doublets}}Y=0\) with colour multiplicity; the quark doublet contributes \(N_{c}\times\tfrac{1}{3}\) and the lepton doublet \(-1\). This is where the number of colours enters: the condition fixes \(N_{c}=3\).
For Equation (104.59): only coloured fields contribute, with \(\tr(\acomm{T^{a}}{T^{b}})=\delta^{ab}\) for the fundamental, so the condition is \(\sum_{L}Y-\sum_{R}Y=0\) over the quark \(\SU(2)\) components: the left doublet supplies \(2\times\tfrac{1}{3}\) and the right singlets \(\tfrac{4}{3}\) and \(-\tfrac{2}{3}\).
For Equation (104.60): the trace is \(\sum Y^{3}\) with all multiplicities. Left: \(6\) quark states (three colours, two isospin) of \(Y=\tfrac{1}{3}\), giving \(6/27\), plus \(2\) lepton states of \(Y=-1\), giving \(-2\); total \(-\tfrac{16}{9}\). Right: \(3(\tfrac{4}{3})^{3} =\tfrac{64}{9}\), \(3(-\tfrac{2}{3})^{3}=-\tfrac{8}{9}\), and \((-2)^{3}=-8\); total \(\tfrac{56}{9}-8=-\tfrac{16}{9}\). Equal.
For Equation (104.61), the triangle with two insertions of the energy–momentum tensor requires \(\sum Y=0\) separately for each chirality summed with sign: left gives \(2-2=0\), right gives \(4-2-2=0\). This is a statement about the electroweak currents in a fixed classical background metric and not a claim about quantum gravity.
In every case the quarks alone fail and the leptons alone fail; only the complete generation cancels. Proposition 106.96 does the same arithmetic in the halved-hypercharge convention of Remark 104.11, with identical conclusions.
∎Proposition 104.53 is usually presented as a consistency check. Read the other way it is a prediction. When the \(b\) quark was found in 1977 [Herb:1977], the third generation had a charged lepton, a neutrino and a charge \(-\tfrac{1}{3}\) quark, and Equation (104.58) then reads \(N_{c}Y_{Q}+Y_{L}\neq0\) unless the \(b\) has an \(\SU(2)_{L}\) partner of charge \(+\tfrac{2}{3}\). The theory therefore required a top quark — not as an aesthetic preference but because without it the theory does not exist — and it was found eighteen years later [Abachi:1995] [Abe:1995] with the mass the electroweak fit had by then indicated. The same argument forbids a fourth generation consisting of leptons alone, and Phenomenon 101.76 independently excludes a fourth light neutrino.
Experimental confirmation
Weak neutral currents
The neutral current is the model's first genuinely new prediction. Every weak process known in 1970 changed the charge of the fermions involved; Proposition 101.67 says a fourth current exists which changes neither charge nor flavour, with a strength fixed relative to the charged current by the single angle \(\theta_{W}\). A theory of this kind is falsifiable in the strongest sense: it predicted a class of events nobody had seen.
Gargamelle was a heavy-liquid bubble chamber of about \(12\,\mathrm{m}^{3}\), filled with freon, exposed at CERN to a neutrino beam. The signature of a neutral-current event is what is not there: hadrons produced with no outgoing muon. Two publications in 1973 carried the discovery. The leptonic one [Hasert:1973a] rests on a single event, \(\bar{\nu}_{\mu}e^{-}\to\bar{\nu}_{\mu}e^{-}\) — one electron recoiling with nothing else in the chamber — for which the expected background was a small fraction of an event; the hadronic one [Hasert:1973b] reports a sample of muonless hadronic events at a rate of about \(0.2\) of the charged-current rate. The credibility of the second rested on an argument about backgrounds rather than on the events themselves: neutrons produced by neutrino interactions in the surrounding material could enter the chamber and produce hadrons with no muon. The collaboration showed that such events would cluster near the chamber walls and fall off with depth, whereas the observed muonless events were distributed through the fiducial volume exactly as the charged-current events were. Phenomenon 101.72 states the result and Proposition 101.71 computes the cross-sections it is compared with.
Confirmation in a different process came from SLAC experiment E122 [Prescott:1978]: longitudinally polarized electrons scattered from deuterium show an asymmetry between the two helicities, of order \(10^{-4}\), which arises only from the interference of the electromagnetic amplitude with a parity-violating neutral-current one. The measurement gave \(\sin^{2}\theta_{W}=0.20(3)\) [Prescott:1978], refined to \(0.224(20)\) by the 1979 scan over beam energies [Prescott:1979], and, crucially, excluded the alternative models then in play — several of which reproduced the Gargamelle neutrino data with a parity-conserving neutral current and were killed by this single number. The apparatus belongs to Experiment: Parity Violation.
The weak neutral current is not the electromagnetic one. It couples to neutrinos, which carry no electric charge [Hasert:1973a] [Hasert:1973b], and it violates parity, as the asymmetry between the scattering of right- and left-handed polarized electrons from deuterium first showed [Prescott:1978]. One parameter, the weak mixing angle, fixes its strength relative to electromagnetism in every process, and the value extracted from neutrino scattering, from polarized electron scattering, from atomic parity violation and from the asymmetries on the \(Z\) resonance is the same,
in the conventions of [Navas:2024], over momentum transfers spanning several orders of magnitude. That one number should serve for all of them is the empirical content of unification. Rests on Equation (104.35), Equation (104.40) and Proposition 101.67.
Derivation. Derives Phenomenon 104.55. Three steps: the rotation that produces the current, the couplings it gives each fermion, and the reason a scheme must be attached to any number quoted.
The rotation. The neutral part of Equation (104.11) is \(\hbar^{-1}\left(gT^{3}W^{3}_{\mu}+g'\tfrac{Y}{2}B_{\mu}\right)\). Substituting Equation (104.35) and demanding that the massless combination \(A_{\mu}\) couple to \(Q\) and to nothing else forces \(\hat{e}=\hat{g}\sin\theta_{W}=\hat{g}'\cos\theta_{W}\), which is Equation (104.40); the orthogonal combination then couples to \(T^{3}-\sin^{2}\theta_{W}Q\) with strength \(\hat{e}/\sin\theta_{W}\cos\theta_{W}\). This is Proposition 101.67, proved there in full.
The couplings. Projecting onto chiralities gives Equation (101.75): for each fermion
with \(T^{3}\) that of the left-chiral component. The whole predictive content follows: the neutrino has \(Q=0\), so \(g_{V}=g_{A}=\tfrac{1}{2}\) and its neutral-current coupling is independent of \(\theta_{W}\) — which is why the ratio of neutral- to charged-current neutrino cross-sections measures the angle cleanly. The electron has \(g_{A}=-\tfrac{1}{2}\) and \(g_{V}=-\tfrac{1}{2}+2\sin^{2}\theta_{W}\), which is small and changes sign near \(\sin^{2}\theta_{W}=\tfrac{1}{4}\): a parity-violating electron asymmetry is therefore a sensitive, and sign-carrying, probe. The \(Z\)-pole asymmetries measure the combination \(\mathcal{A}_{f}=2g_{V}g_{A}/(g_{V}^{2}+g_{A}^{2})\), which for the electron is a steep function of \(\sin^{2}\theta_{W}\) for exactly the same reason. Four independent kinds of measurement thus depend on one number in four different functional ways, and that is what makes their agreement a test rather than a fit.
The scheme. Beyond tree level, \(\theta_{W}\) is not a parameter of the Lagrangian but a renormalized quantity, and different definitions differ by finite amounts of order \(\alpha/\pi\) — the three standard ones are listed in Remark 104.34 and differ by about one per cent, which is hundreds of times the experimental errors. Worse, in the \(\overline{\mathrm{MS}}\) scheme the angle runs: the same definition gives about \(0.2313\) at \(\mu=M_{Z}c^{2}\) and about \(0.2387\) at \(\mu\to0\), the difference being real physics — vacuum polarization by charged particles, Section 100.7 — and not a change of convention. Table 104.3 lists the determinations across momentum transfer, and the content of Equation (104.62) is that one running curve passes through all of them. A number quoted without its scheme and its scale is not a measurement.
∎| Measurement | Momentum transfer | $\sin^{2}\theta_{W}$ | Reference |
|---|---|---|---|
| Cesium atomic parity violation | $\approx0$ | $\approx0.24$ | [Wood:1997] |
| Polarized Møller scattering, SLAC E158 | $Q^{2}=0.026\,\mathrm{GeV}^{2}/c^{2}$ | \(0.2397(13)\) | [Anthony:2005] |
| Polarized $ep$ scattering, Qweak | $Q=0.157\,\mathrm{GeV}/c$ | \(0.2383(11)\) | [Androic:2018] |
| Polarized $ed$ scattering, SLAC E122 | $Q^{2}\approx1.5\,\mathrm{GeV}^{2}/c^{2}$ | $\approx0.20(3)$ | [Prescott:1978] |
| $Z$-pole asymmetries, LEP and SLD | $Q=M_{Z}c$ | \(0.23153(16)\) | [Schael:2006] |
Discovery of the $W$ and $Z$ bosons
By 1980 the masses were predicted. Putting \(\sin^{2}\theta_{W}=0.23\) from Table 104.3 into Equation (104.41),
at tree level, and about \(80\,\mathrm{GeV}\) and \(91\,\mathrm{GeV}\) once \(\Delta r\) of Remark 104.33 is included — numbers derived from a neutrino cross-section and a muon lifetime, with no collider involved.
Producing them required a beam of a few hundred \(\mathrm{GeV}\) in the centre of mass, and CERN got one by converting the SPS into a proton–antiproton collider. The enabling invention was stochastic cooling [vanderMeer:1985]: a pickup senses the deviation of a sample of the circulating beam from the mean orbit, and a kicker one straight section downstream applies a correcting impulse, reducing the phase-space volume of the antiproton beam by many orders of magnitude over hours. Without it there is no way to accumulate enough antiprotons.
The signature is a lepton of high transverse momentum with nothing balancing it. UA1 [Arnison:1983a] and UA2 [Banner:1983] reported in 1983 events with an isolated electron of \(p_{T}\gtrsim20\,\mathrm{GeV}/c\) and large missing transverse energy — the neutrino — from \(W\to e\nu\). The mass is read from the transverse-mass distribution rather than from the momenta directly, because the neutrino's longitudinal momentum is unmeasured; the observable is Equation (101.84) and the reason it works is a kinematic singularity.
Let a boson of mass \(M\) at rest decay to two massless particles. The transverse momentum of either with respect to any axis is \(p_{T}=\tfrac{1}{2}Mc\sin\theta^{*}\), and therefore the distribution in \(p_{T}\) has an integrable singularity at the endpoint,
which diverges as \(p_{T}\to\tfrac{1}{2}Mc\) however smooth the angular distribution is. The position of the edge measures \(M\). Rests on Equation (101.1).
Derives Proposition 104.56. In the rest frame each decay product has momentum \(\tfrac{1}{2}Mc\), so \(p_{T}=\tfrac{1}{2}Mc\sin\theta^{*}\) with \(\theta^{*}\) the polar angle to the chosen axis. Then \(\dd p_{T}/\dd(\cos\theta^{*}) =-\tfrac{1}{2}Mc\,\cos\theta^{*}/\sin\theta^{*}\), and inverting and substituting \(\sin\theta^{*}=2p_{T}/Mc\), \(\cos\theta^{*}=\sqrt{1-(2p_{T}/Mc)^{2}}\) gives Equation (104.64). The Jacobian, not the dynamics, produces the peak; this is why the method is robust against the unknown production mechanism, and why smearing it by the boson's transverse motion and by the detector resolution is the dominant systematic. Applying the same construction to the pair, with the neutrino's transverse momentum inferred from the missing energy, gives the transverse mass Equation (101.84), whose endpoint is \(Mc^{2}\) itself and which is less sensitive to the boson's motion.
∎Later in 1983 both experiments found the neutral partner in the cleaner channel \(Z\to e^{+}e^{-}\) and \(\mu^{+}\mu^{-}\), where both leptons are measured and the invariant mass is direct [Arnison:1983b] [Bagnaia:1983]: a handful of events at an invariant mass near \(95\,\mathrm{GeV}\)/\(c^{2}\), on essentially no background. Phenomenon 101.73 records the measurement and Equation (101.83) the comparison. The masses agreed with Equation (104.63) at the accuracy then available, which is the point: two numbers measured in the 1970s in low-energy experiments predicted where two new particles would be found, and they were found there.
LEP, SLD and the $Z$ resonance
LEP turned the \(Z\) from a discovery into an instrument. Four experiments — ALEPH, DELPHI, L3 and OPAL — recorded some seventeen million \(Z\) decays between 1989 and 1995, and SLD added a smaller sample with a longitudinally polarized electron beam, which gives access to asymmetries the unpolarized machine cannot reach. The combined analysis [Schael:2006] is the most precise test the electroweak theory has had.
Three classes of result matter here.
The lineshape. Scanning the beam energy across the resonance determines \(M_{Z}\) and the total width \(\Gamma_{Z}\), and the peak cross-section determines the product of partial widths. The masses quoted throughout this chapter come from this analysis. The energy calibration was itself a small physics programme: the beam energy was measured by resonant depolarization to a precision of about \(1\,\mathrm{MeV}\), at which level the circumference of the ring responds to the tides raised by the Moon and the Sun, to the water level in Lake Geneva, and to leakage currents from a nearby railway line. All three had to be modelled to reach the quoted \(M_{Z}\).
The invisible width. Subtracting the visible partial widths from the total leaves the width into channels the detector cannot see, and dividing by the calculated width per neutrino species counts them: \(N_{\nu}=2.984\pm0.008\), which is Phenomenon 101.76. That a resonance curve counts the generations of matter is among the most economical measurements in physics.
The asymmetries. Forward–backward asymmetries in \(e^{+}e^{-}\to f\bar{f}\) and, at SLD, the left–right asymmetry with polarized beams, measure \(\sin^{2}\theta^{\ell}_{\mathrm{eff}}\) to better than one part in a thousand (Table 104.3). Proposition 101.75 derives the partial widths that these are compared with.
LEP's second phase ran above the \(W\)-pair threshold and measured something no fixed-target experiment can: that the gauge bosons couple to each other.
The cross-section for \(e^{+}e^{-}\to W^{+}W^{-}\), measured at LEP2 from threshold up to \(209\,\mathrm{GeV}\), rises to a maximum and then falls. Three amplitudes contribute — neutrino exchange in the \(t\) channel, and annihilation through a photon and through a \(Z\) — and each grows with energy on its own; the data are reproduced only when all three are present and interfere destructively, the neutrino-exchange term alone exceeding the measurement by a large factor at the highest energies [Schael:2013]. The two annihilation diagrams exist only if three gauge bosons meet at a vertex. Rests on Proposition 104.7, Proposition 104.47 and Equation (104.40).
Derivation. Derives Phenomenon 104.57. The claim to be established is that the three amplitudes must cancel their leading high-energy growth, and that the cancellation fixes the triple-gauge coupling to the Yang–Mills value.
Take the final \(W\)s longitudinal, which is where the growth lives by Equation (104.6), and apply Proposition 104.47: at \(\sqrt{s}\,c\gg M_{W}c^{2}\) the amplitude equals that for producing the two eaten Goldstone bosons, \(e^{+}e^{-}\to\chi^{+}\chi^{-}\). Now enumerate the diagrams for that process.
-
The \(s\)-channel photon and \(Z\) survive, with the couplings of a charged scalar read off Equation (104.11): the \(\chi^{\pm}\) sits in the doublet with \(T^{3}=\pm\tfrac{1}{2}\) and \(Q=\pm1\), so it couples to the photon with charge \(\hat{e}\) and to the \(Z\) with \(\hat{e}\left(1-2\sin^{2}\theta_{W}\right)/ 2\sin\theta_{W}\cos\theta_{W}\), by Equation (101.74) evaluated at those quantum numbers.
-
The \(t\)-channel neutrino exchange survives only through the Yukawa coupling of \(\chi^{\pm}\) to \(e\nu\), which by Theorem 104.38 is \(\hat{y}_{e}/\sqrt{2}\), of order \(2\times 10^{-6}\) from Table 104.2, and is therefore negligible.
The high-energy amplitude is thus the production of a pair of charged scalars through two \(s\)-channel vector exchanges. Such an amplitude does not grow: an \(s\)-channel propagator \(1/(s-M^{2}c^{2})\) falls as \(1/s\) while the vertices supply at most one power of momentum each, so the amplitude tends to a constant in the \(J=1\) partial wave and the cross-section falls as \(1/s\), which is the observed shape.
The consequence is the cancellation. Computed instead in unitary gauge, each of the three diagrams carries two longitudinal polarization vectors and so grows as \(s/M_{W}^{2}c^{2}\); but their sum must equal the non-growing scalar amplitude just obtained. Hence the leading terms cancel identically. The cancellation uses the couplings in exactly the combinations Equation (104.40) supplies — the \(\gamma W^{+}W^{-}\) vertex equal to \(\hat{e}\) and the \(ZW^{+}W^{-}\) vertex equal to \(\hat{g}\cos\theta_{W}\), both of which are forced by Proposition 104.7 because the \(WW\gamma\) and \(WWZ\) couplings come from the same \(\varepsilon^{abc}\) in Equation (104.8). Change either coupling by a fraction \(\delta\) and a term growing as \(\delta s/M_{W}^{2}c^{2}\) survives; at \(\sqrt{s}\,c=200\,\mathrm{GeV}\) that factor is \(s/M_{W}^{2}c^{2}\approx6\), so a few-per-cent change in the vertex is visible in the cross-section. That is why the LEP2 measurement of a cross-section is a measurement of the group structure: the agreement in [Schael:2013] bounds the anomalous triple-gauge couplings at the per-cent level, and the neutrino-exchange diagram alone — with the two annihilation diagrams removed — overshoots the data at the highest energies by a large factor, exactly as an uncancelled \(s/M_{W}^{2}c^{2}\) requires.
∎Precision fits and the $W$ mass
Remark 104.33 showed that four measured quantities are described by three parameters, so one relation is predicted. The electroweak fit generalizes this: several dozen measurements — \(M_{Z}\), \(\Gamma_{Z}\), the partial widths, the asymmetries, \(M_{W}\), \(m_{t}\), \(\alpha\), \(\alpha_{s}\), \(G_{F}\) — are described by a handful of parameters, so the system is heavily over-constrained and its consistency is a test at every point [Baak:2014].
The historically decisive property is that the one-loop corrections depend on particles too heavy to produce. \(\Delta\rho\) grows as \(m_{t}^{2}\) by Equation (104.52), so the LEP data constrained the top mass before the Tevatron produced it, and the indirect value agreed with the direct measurement of [Abachi:1995] [Abe:1995]. The scalar enters only logarithmically — Veltman's screening theorem again — so the constraint on \(m_{h}\) was much weaker, but it was real: the fit before 2012 preferred a scalar mass of order \(100\,\mathrm{GeV}\) and excluded values above a few hundred \(\mathrm{GeV}\) [Baak:2014], and the particle was found at \(125\,\mathrm{GeV}\)/\(c^{2}\) inside that range. Two particles were thus located by loop effects before either was seen.
One entry does not fit, and it is stated here as what it is.
The CDF collaboration's final analysis of its Tevatron data [Aaltonen:2022] reports \(M_{W}c^{2}=80433.5(94)\,\mathrm{MeV}\), an uncertainty of \(9.4\,\mathrm{MeV}\) — the most precise single measurement of the quantity. It disagrees with the world average \(80.369(13)\,\mathrm{GeV}\) [Navas:2024] by about four standard deviations, and with the indirect prediction of the electroweak fit, \(M_{W}c^{2}=80357(6)\,\mathrm{MeV}\) [Navas:2024], by about seven: the difference of \(76.5\,\mathrm{MeV}\) is \(6.9\) times the combined uncertainty of \(11.2\,\mathrm{MeV}\) obtained by adding the two errors in quadrature. It also disagrees with the LHC measurements, which are individually less precise but mutually consistent and consistent with the fit.
This is recorded as an unresolved experimental tension and not as evidence for anything. Two measurements of the same quantity that disagree by seven standard deviations mean that at least one uncertainty is underestimated; which one is not known. The relevant systematics — parton distributions, the modelling of the \(W\) transverse-momentum spectrum, the momentum-scale calibration — are of the same order as the quoted total error, and no independent reanalysis has reproduced the CDF value. Until it is resolved, the honest statement is that the \(W\) mass is known to about \(13\,\mathrm{MeV}\) with one outlier at four sigma, and no shift in any Standard Model parameter is warranted.
The Higgs boson
Search and discovery
Everything about the scalar is predicted except its mass. Given \(m_{h}\), Equation (104.38) fixes \(\lambda\), and Theorem 104.38(iii) and Equation (104.23) fix every coupling, hence every production cross-section and every branching fraction. The search was therefore a scan in one variable with everything else nailed down — which is why a negative result at any mass is a genuine exclusion, and why the positive result is a test rather than a fit.
LEP searched in \(e^{+}e^{-}\to Zh\) up to a centre-of-mass energy of \(209\,\mathrm{GeV}\) and set the limit
[Barate:2003], which is about as far as that machine could reach. The Tevatron and then the LHC took the range upward. In July 2012 ATLAS [Aad:2012] and CMS [Chatrchyan:2012] independently reported a resonance near \(125\,\mathrm{GeV}\)/\(c^{2}\), each with a local significance of about five standard deviations, in the two channels with the cleanest signature: \(h\to\gamma\gamma\), where the mass is reconstructed from two photons, and \(h\to ZZ^{*}\to4\ell\), where it is reconstructed from four charged leptons. Their first joint mass combination [Aad:2015] gave \(125.09(24)\,\mathrm{GeV}\), and the present world average is \(125.20(11)\,\mathrm{GeV}\) [Navas:2024]. The apparatus, the trigger, the look-elsewhere effect and the statistical machinery belong to Experiment: The Higgs Boson Discovery; the statistical foundations belong to Probability and Statistics.
A narrow resonance is produced in proton–proton collisions and decays to two photons and to four charged leptons. It was observed independently by ATLAS [Aad:2012] and CMS [Chatrchyan:2012] in 2012, at a mass
in their first combination [Aad:2015], now \(125.20(11)\,\mathrm{GeV}\) [Navas:2024]. That it decays to two photons excludes spin one [Landau:1948] [Yang:1950], and the angular distributions of the four-lepton final state select \(J^{P}=0^{+}\) over the remaining alternatives. Direct searches at LEP had already excluded a Standard Model scalar below \(114.4\,\mathrm{GeV}/c^{2}\) [Barate:2003]. The full discovery account belongs to Experiment: The Higgs Boson Discovery. Rests on Theorem 104.38, Table 104.2 and Equation (101.86).
Derivation of the decay pattern. Derives Phenomenon 104.59. Once \(m_{h}\) is given, the tree-level partial widths follow from the couplings of Theorem 104.38(iii) with no freedom left. For a fermion pair,
\(N_{c}\) being \(3\) for quarks and \(1\) for leptons, and \(\hat{y}_{f}=\sqrt{2}m_{f}c^{2}/E_{v}\) from Table 104.2. The structure is that of Equation (101.86) — the width of a particle at rest into two massless ones is a coupling squared times the mass over \(16\pi\), times a phase-space factor — and the exponent \(3/2\) rather than \(1/2\) is the signature of a scalar: the \(\beta^{3}\) arises because a \(J=0\) state decaying to a fermion pair through a scalar coupling produces them in a \(P\) wave, contributing \(\beta^{2}\) beyond the \(\beta\) of two-body phase space. A pseudoscalar coupling \(\ii\bar{f}\gamma^{5}f\) would give \(\beta^{1}\) instead, which is one of the ways the parity of the state is tested.
Two checks with no strong-interaction corrections. For \(\tau^{+}\tau^{-}\), Table 104.2 gives \(\hat{y}_{\tau}=0.01021\), so \(\hbar\Gamma=1.0416\times 10^{-4}\times125.20\,\mathrm{GeV}/16\pi =0.259\,\mathrm{MeV}\); for \(\mu^{+}\mu^{-}\), \(\hat{y}_{\mu}=6.07\times 10^{-4}\) gives \(0.92\,\mathrm{keV}\). Against the total width of \(4.1\,\mathrm{MeV}\) these are branching fractions of \(6.3\times 10^{-2}\) and \(2.2\times 10^{-4}\), which are the values in Table 104.4.
The two heaviest available channels need a further remark. At \(m_{h}c^{2}=125.20\,\mathrm{GeV}\) both \(W^{+}W^{-}\) and \(ZZ\) are below threshold, since \(2M_{W}c^{2}=160.7\,\mathrm{GeV}\); they proceed with one boson off shell, \(h\to VV^{*}\to Vf\bar{f}'\), a three-body decay whose rate is suppressed by an extra factor of order \(\hat{g}^{2}/16\pi^{2}\) relative to the on-shell formula and which is not obtained by putting a negative number under a square root. That this suppression is only partial — \(WW^{*}\) is still the second largest channel — is because the coupling grows as the squared vector mass, Equation (104.69) below.
The total width predicted by summing all channels is \(4.1\,\mathrm{MeV}\), some \(3\times10^{-5}\) of the mass; the measured value is \(3.7\,\mathrm{MeV}\) with an uncertainty of about \(1.7\,\mathrm{MeV}\) [Navas:2024], in agreement. A width this narrow is itself a prediction: it follows from the couplings being small, which follows from most fermions being light.
∎| Channel | Mechanism | Branching fraction |
|---|---|---|
| $b\bar{b}$ | tree, Yukawa | \(0.58\) |
| $W^{+}W^{-*}$ | tree, gauge | \(0.21\) |
| $gg$ | loop (top) | \(0.082\) |
| $\tau^{+}\tau^{-}$ | tree, Yukawa | \(0.063\) |
| $c\bar{c}$ | tree, Yukawa | \(0.029\) |
| $ZZ^{*}$ | tree, gauge | \(0.026\) |
| $\gamma\gamma$ | loop ($W$, top) | \(0.0023\) |
| $Z\gamma$ | loop ($W$, top) | \(0.0015\) |
| $\mu^{+}\mu^{-}$ | tree, Yukawa | \(2.2\times 10^{-4}\) |
The loop-induced scalar vertices: the amplitudes for \(h\to\gamma\gamma\) and \(h\to gg\), from the \(W\) and top triangles, which have no tree-level counterpart and which supply the discovery channel and the dominant production mechanism (gluon fusion) respectively; together with the resulting production cross-sections at the LHC. These require the full one-loop tensor reduction and belong in Appendix A. Also owed are the three-body off-shell rates \(h\to WW^{*}\) and \(h\to ZZ^{*}\), whose branching fractions the table quotes without derivation. What is derived above is the two-body tree-level fermionic width formula and the tau and muon entries computed from it; every other entry of the branching table quotes the standard values.
Quantum numbers and couplings
A massive particle of spin \(1\) cannot decay into two real photons [Landau:1948] [Yang:1950]. Hence the observation of \(h\to\gamma\gamma\) excludes \(J=1\) for the resonance. Rests on Theorem 99.46.
Derives Theorem 104.60. Work in the rest frame of the decaying particle, where the two photons travel along \(\pm\hat{\vect{k}}\) with polarization vectors \(\vect{\varepsilon}_{1}\), \(\vect{\varepsilon}_{2}\), both transverse: \(\vect{\varepsilon}_{i}\cdot\hat{\vect{k}}=0\). The amplitude is a scalar function linear in each \(\vect{\varepsilon}_{i}\), and for \(J=1\) it must be a vector \(\vect{A}\) contracted with the decaying particle's polarization. Bose symmetry requires \(\vect{A}\) to be invariant under exchanging the two photons, which means \(\vect{\varepsilon}_{1}\leftrightarrow\vect{\varepsilon}_{2}\) together with \(\hat{\vect{k}}\to-\hat{\vect{k}}\).
Enumerate the vectors that are linear in each polarization and built from \(\vect{\varepsilon}_{1}\), \(\vect{\varepsilon}_{2}\), \(\hat{\vect{k}}\):
the last vanishing by transversality. \(\vect{A}_{1}\) and \(\vect{A}_{2}\) are antisymmetric under the exchange and are therefore forbidden for two identical bosons, and no other structure exists. The one remaining candidate, \(\left[\hat{\vect{k}}\cdot\left(\vect{\varepsilon}_{1}\times \vect{\varepsilon}_{2}\right)\right]\hat{\vect{k}}\), is not independent: transversality makes \(\vect{\varepsilon}_{1}\times\vect{\varepsilon}_{2}\) parallel to \(\hat{\vect{k}}\), so it is \(\vect{A}_{1}\), and it is antisymmetric for the same reason. Anything else contains either \(\hat{\vect{k}}\cdot\vect{\varepsilon}_{i}\), which vanishes, or a second power of some \(\vect{\varepsilon}_{i}\), which the linearity forbids. Hence the \(J=1\) amplitude vanishes identically.
The argument uses only rotational invariance, Bose statistics and transversality; it does not use parity, so it excludes \(1^{+}\) and \(1^{-}\) alike, and it fails for a massive vector in the final state (which has a longitudinal polarization) and for two virtual photons. Together with the four-lepton angular distributions, which distinguish \(0^{+}\) from \(0^{-}\) and from \(2^{+}\), this fixes \(J^{PC}=0^{++}\).
∎The decisive test of the mechanism is not that a scalar exists but that its couplings track mass, because that is what makes it the field that supplies mass rather than merely another particle.
In unitary gauge the interactions of \(h\) with the massive fields are
that is: every coupling is the corresponding mass term divided by \(v\), twice over for the vectors. The vertex strengths are therefore
and the variable that puts both families on one axis — the reduced coupling of the experimental figures — is \(g_{hff}\) for a fermion and \(\sqrt{g_{hVV}/2E_{v}}=M_{V}c^{2}/E_{v}\) for a vector. Plotted against mass, every Standard Model particle then lies on the single straight line of slope \(1/E_{v}\) through the origin. Rests on Theorem 104.38, Equation (104.46) and Equation (104.34).
Derives Proposition 104.61. For the fermions, Theorem 104.38(iii): substituting \(\Phi=\frac{1}{\sqrt{2}}(0,v+h)\transpose\) into Equation (104.46) multiplies the mass term by \(\left(1+h/v\right)\), whose linear part is the first term of Equation (104.69). Dividing by \(\sqrt{\hbar c}\) to make it dimensionless, and using \(E_{v}=v\sqrt{\hbar c}\), gives \(g_{hff}=m_{f}c^{2}/E_{v}\).
For the vectors, Equation (104.34) was computed at \(\Phi=\Phi_{0}\); restoring the fluctuation replaces \(v\) by \(v+h\) throughout, so the mass terms are multiplied by \(\left(1+h/v\right)^{2}=1+2h/v+h^{2}/v^{2}\). The linear part is twice the mass term over \(v\), which is the bracket in Equation (104.69); the factor of two is the whole difference between the fermion and vector cases and it comes from the mass term being quadratic in the field's coupling to \(\Phi\) rather than linear. Hence \(g_{hVV}=2(M_{V}c^{2})^{2}/E_{v}\), quadratic in the mass and carrying the dimension of an energy; dividing by \(2E_{v}\) and taking the square root returns \(M_{V}c^{2}/E_{v}\), the same linear law the fermions obey, which is why the two families can be drawn on one plot.
∎The strength with which the scalar couples to each particle is measured channel by channel, and it grows with the mass of the particle it couples to: linearly for the fermions and quadratically for the weak bosons. The ten-year coupling maps of ATLAS [Aad:2022] and CMS [Tumasyan:2022] cover \(W\), \(Z\), \(t\), \(b\), \(\tau\) and, with weaker significance, \(\mu\) — some three orders of magnitude in mass — and follow that pattern throughout. This, rather than the mere existence of a scalar, is the test of the mechanism: a scalar whose couplings did not track mass would not be the field that supplies mass. Rests on Equation (104.70), Proposition 104.61 and Proposition 104.26.
Derivation. Derives Phenomenon 104.62. The prediction is Equation (104.70), derived in Proposition 104.61: a one-parameter family with the parameter, \(E_{v}\), already fixed by the muon lifetime through Proposition 104.26. There is nothing to adjust.
The measurement is conventionally reported as coupling modifiers \(\kappa_{i}\), defined so that \(\kappa_{i}=1\) reproduces Equation (104.70) and rates scale as \(\kappa^{2}\). What is actually measured is a rate, that is a product of a production cross-section and a branching fraction, so a global fit over many final states is required to separate the couplings; the experiments report the results of that fit [Aad:2022] [Tumasyan:2022], and every \(\kappa_{i}\) comes out consistent with unity, with uncertainties of roughly a tenth for \(W\), \(Z\), \(t\), \(b\) and \(\tau\) and roughly a fifth for \(\mu\). Since \(m_{\mu}c^{2}=0.1057\,\mathrm{GeV}\) and \(m_{t}c^{2}=172.57\,\mathrm{GeV}\), the verified range is a factor of \(1630\) in mass, over which Equation (104.70) varies by the same factor — so the agreement is a statement about a three-decade lever arm and not about one point.
Two features make the test sharper than a fit to five numbers. The \(\gamma\gamma\) and \(gg\) vertices have no tree-level counterpart (Table 104.4), so their measured rates test the same couplings at one loop, through a different combination: the \(\gamma\gamma\) amplitude is dominated by the \(W\) loop with the top loop interfering destructively, and the \(gg\) amplitude by the top alone, so agreement in both simultaneously constrains the relative sign of \(\kappa_{W}\) and \(\kappa_{t}\) — something no tree-level measurement reaches. And any heavy coloured or charged particle would contribute to those same loops whether or not it can be produced directly, so the agreement bounds such particles indirectly.
∎The potential and the self-coupling
Everything verified so far probes the theory at the minimum of the potential. The shape of the potential away from the minimum is a separate claim and it is largely unmeasured.
Expanding Equation (104.27) in unitary gauge,
so both the trilinear and the quartic self-coupling are determined by \(m_{h}\) and \(v\), which are already measured. Writing the trilinear coupling as \(\kappa_{\lambda}\) times its Standard Model value, the mechanism predicts \(\kappa_{\lambda}=1\) with no freedom. Rests on Equations (104.16), (104.27) and (104.37).
Derives Proposition 104.63. From Equation (104.16) with the Goldstone modes gauged away, \(V=\lambda v h^{3}+\tfrac{\lambda}{4}h^{4}+\lambda v^{2}h^{2}\), and \(\lambda v^{2}=\tfrac{1}{2}(m_{h}c/\hbar)^{2}\) by Equation (104.37). Substituting \(\lambda=(m_{h}c)^{2}/2\hbar^{2}v^{2}\) into each term gives Equation (104.71). The two higher couplings are \(m_{h}^{2}/v\) and \(m_{h}^{2}/v^{2}\) times pure numbers: one parameter, \(\lambda\), generates the mass and both self-interactions, which is the content of the claim that the potential is the quartic one and not merely some function with a minimum in the right place.
∎Measuring \(\kappa_{\lambda}\) requires producing two scalars at once, whose cross-section at the LHC is some three orders of magnitude below single production and which additionally interferes with a box diagram that does not involve the self-coupling. The present constraints [Aad:2022] allow \(\kappa_{\lambda}\) over a range of order \(-1\) to \(+7\) — consistent with the predicted \(1\), and far too loose to be called a test. Until it narrows, the statement that the vacuum sits at the minimum of a quartic potential is an assumption supported by the mass and the couplings, not a measurement of the potential's shape.
The tree-level \(V(h)\) is not the object whose minimum determines the vacuum. That object is the effective potential of Equation (106.59), the Legendre transform of the generating functional evaluated at a constant field, whose one-loop correction Coleman and Weinberg computed [Coleman:1973] and Proposition 106.45 reproduces: each field circulating in the loop contributes a term proportional to \(\pm M^{4}(h)\ln M^{2}(h)\), with a plus sign for bosons and a minus for fermions, \(M(h)\) being that field's \(h\)-dependent mass. In the Standard Model the top quark's contribution is the largest and it enters with the fermionic minus sign. That single fact is the origin of everything in Section 104.8.1.
At high temperature the effective potential acquires thermal corrections and the minimum returns to \(\Phi=0\): the symmetry is restored, and the early universe passed through the transition at a temperature of order \(100\,\mathrm{GeV}/k_{B}\). Whether the transition is first order — with bubbles, latent heat and departure from equilibrium — or a smooth crossover depends on the scalar mass, and lattice computations place the endpoint of the first-order line at \(m_{h}c^{2}\approx72\,\mathrm{GeV}\) [Kajantie:1996]. The measured \(125.20\,\mathrm{GeV}\) is far above it, so the Standard Model transition is a crossover. The consequence is not academic: Sakharov's third condition for generating a baryon asymmetry [Sakharov:1967] is a departure from thermal equilibrium (Theorem 105.41), and a crossover supplies none. The Standard Model therefore cannot account for the observed baryon asymmetry, and Section 105.6.2 quantifies the shortfall. This is a failure of the theory derived from measured parameters, and it is stated here because this chapter measured them.
What the model does not explain
Vacuum stability
\(\lambda\) is a coupling, so it runs (The Renormalization Group). Its beta function is dominated by the top Yukawa, and the top's contribution is negative.
At one loop, in the conventions of Definition 104.24 and with \(\hat{y}_{t}\), \(\hat{g}\), \(\hat{g}'\) as in Notation 104.1,
[Degrassi:2012]. Evaluated at \(\mu=m_{t}c^{2}\), where the running couplings are \(\hat{\lambda}=0.126\), \(\hat{y}_{t}=0.94\), \(\hat{g}=0.64\), \(\hat{g}'=0.35\) — each a little below its value at the electroweak scale in Equation (104.38), Table 104.2 and Equation (104.42), because all of them run — the five terms are \(0.381\), \(1.336\), \(-4.684\), \(-0.511\) and \(0.232\), so
negative, and dominated by the single term \(-6\hat{y}_{t}^{4}\). Rests on Definition 104.24, Equation (104.38) and Corollary 106.39.
Derivation of the sign. Derives Proposition 104.66. The coefficients are quoted from [Degrassi:2012]; what is derived here is why the dominant one is negative and why it is large. By Remark 104.64 the one-loop effective potential receives \(-N_{c}M_{t}^{4}(h)\ln M_{t}^{2}(h)\) from the top, with the minus sign of a fermion loop — the same sign that makes a closed fermion loop carry \((-1)\) in Corollary 106.39. Since \(M_{t}(h)\propto\hat{y}_{t}h\), this is a term \(\propto-\hat{y}_{t}^{4}h^{4}\ln h\), that is a negative contribution to the running quartic, with the colour factor \(N_{c}=3\) and the fourth power of the coupling. Numerically \(\hat{y}_{t}^{4}=0.78\) while \(\hat{\lambda}^{2}=0.016\), so the fermion term beats the bosonic ones by a factor of order fifty, and Equation (104.73) follows by arithmetic. The one number that matters is the sign.
∎Integrating Equation (104.72) together with the running of \(\hat{y}_{t}\) and \(\hat{g}_{3}\) — which cannot be done by holding the coefficients fixed, because \(\hat{y}_{t}\) itself falls with scale and \(\abs{\dd\hat{\lambda}/\dd\ln\mu}\) falls with it — gives \(\hat{\lambda}(\mu)\) crossing zero at \(\mu\approx10^{10}\)–\(10^{11}\,\mathrm{GeV}\) [Degrassi:2012] [Buttazzo:2013]. Above that scale the potential turns over and a deeper minimum exists at large field values: the electroweak vacuum is not the global minimum but a metastable one. The computed tunnelling lifetime exceeds the age of the universe by hundreds of orders of magnitude, so nothing observable follows.
Three caveats, all of which the cited analyses state and none of which may be dropped:
-
the conclusion sits on a knife edge in \(m_{t}\). Shifting the top mass by about \(1\,\mathrm{GeV}\), or the strong coupling by about one standard deviation, moves the answer between absolute stability, metastability and instability. The measured \(m_{t}\) has an uncertainty of \(0.29\,\mathrm{GeV}\) [Navas:2024], but the relation between the quantity a hadron collider reconstructs and the short-distance mass entering Equation (104.72) carries a further theoretical uncertainty of comparable size;
-
the extrapolation assumes no new particle couples to the scalar anywhere between \(10^{3}\,\mathrm{GeV}\) and \(10^{10}\,\mathrm{GeV}\) — sixteen decades in which nothing has been looked at. A single new state changes the running;
-
the running is a statement about a coupling in a scheme, not about an observable. Whether “the vacuum is metastable” is a physical statement at all depends on the extrapolation being valid, which is the assumption in question.
The honest summary is that the measured parameters put the Standard Model, extrapolated with no new physics, marginally on the metastable side, and that the margin is smaller than the combined theoretical and experimental uncertainty on the top mass. It is an intriguing near-coincidence, and it is not evidence for anything.
Vacuum stability: the coupled integration of the one-loop renormalization-group equations for the quartic coupling, the top Yukawa and the three gauge couplings from the electroweak scale upward, which is what locates the zero of the quartic near \(10^{10}\,\mathrm{GeV}\); and the bounce action governing the tunnelling rate out of the electroweak minimum. The beta function and its sign are derived in the text above; the integration and the bounce belong in Appendix A.
Open questions, stated honestly
The electroweak theory explains a great deal from very little: two gauge couplings, one scalar mass and one vacuum expectation value generate the \(W\), the \(Z\), the photon's exact masslessness, the Fermi constant, the neutral current and the whole pattern of boson self-interactions. What it does not do is explain its own inputs.
The Yukawa hierarchy. Table 104.2 is a list of nine numbers spanning \(3.4\times 10^{5}\), taken from experiment. Nothing in the Lagrangian relates them, and no evidence-based mechanism producing them exists. The four parameters of the CKM matrix (Proposition 101.63) are in the same position.
The number of generations. Three, by Phenomenon 101.76 for the light neutrinos and by Proposition 104.53 for the requirement that each be complete. Why three is not known; the theory works for any number.
The scale. \(E_{v}=246.22\,\mathrm{GeV}\) is one input. Expressed against the only other scale built from the constants of Nature, the Planck energy \(\sqrt{\hbar c^{5}/G}\approx1.22\times 10^{19}\,\mathrm{GeV}\), it is smaller by a factor \(2\times10^{-17}\), and Remark 104.30 records a related and much larger discrepancy in the vacuum energy.
Sensitivity of \(m_{h}^{2}\) to heavy thresholds. If the theory is embedded in a larger one containing a particle of mass \(M\) coupled to \(\Phi\), the scalar's squared mass receives a correction of order \(M^{2}\) times a loop factor, so reproducing the measured \(125.20\,\mathrm{GeV}\) requires a cancellation of relative size \((m_{h}/M)^{2}\). It should be said plainly what this is and is not. It is an argument about the form a more fundamental theory would have to take; it is not an observation, because no such particle is known to exist and the correction is not a measurable quantity. It has motivated a great many proposed extensions — supersymmetry, technicolour, grand unification, extra dimensions — and direct searches at the LHC have so far returned null results for every one of them. None of those programmes has observational support, and by editorial rule 1 they are named here only to record that fact.
Neutrino mass. Equation (104.46) gives the neutrino no mass, because there is no \(\nu_{R}\) in Table 104.1. Oscillations show the masses are not zero (Experiment: Neutrino Oscillations). This is the one place where a laboratory measurement contradicts the Lagrangian of Equation (104.28) rather than merely failing to be explained by it, and Flavour Physics and Neutrinos takes it up.
Dark matter, dark energy, the baryon asymmetry. None has an explanation here. Remark 104.65 showed that the measured scalar mass positively forbids the Standard Model from generating the baryon asymmetry, which is a sharper statement than silence. The cosmological evidence belongs to The Dark Sector: Evidence Without Explanation, and the full inventory of what is observed and unexplained is What We Observe but Do Not Understand.