Flavour Physics and Neutrinos
Flavour is the part of the Standard Model that nothing explains. The gauge structures of Quantum Electrodynamics and Renormalization, Weak Interactions and Quantum Chromodynamics follow from one principle applied to three groups; the existence of three copies of every fermion, their wildly disparate masses, and the two mixing matrices that relate the states of definite mass to the states that couple to the \(W\) are pure measurement. This chapter assembles that measurement. It is also where the treatise's editorial rule bites hardest: flavour is the richest source of proposals without evidence, and what is recorded below is what has been seen — quark mixing, \(CP\) violation in three meson systems, and neutrino oscillation — together with an explicit statement of what has not.
The chapter follows Quantum Chromodynamics because the hadronic states whose decays carry the information are QCD bound states, and it feeds Experiment: CP Violation and Experiment: Neutrino Oscillations, the two chapters that carry the decisive experimental programmes; the first is still a skeleton of reserved headings and the second is written in part, so where a measurement is owed to one of them this chapter says so. Neutrino oscillation is the only laboratory result that the Standard Model as originally written does not accommodate, since it requires neutrino mass; that makes this chapter the natural bridge to What We Observe but Do Not Understand, and its cosmological consequences — the matter–antimatter asymmetry — connect to Evidence-Based Cosmology. Standard treatments are [Weinberg:1996] [Halzen:1984] [Giunti:2007]; all quoted values are from [Navas:2024] unless another source is named, and every constant is the CODATA 2022 recommendation [Mohr:2025].
Two pieces of scaffolding are inherited rather than rebuilt. The kinematic and dimensional conventions — metric signature, the mass-shell relation \(p^{2}=m^{2}c^{2}\), the SI dimension carried by a field and by a coupling — are those fixed in Section 100.1.1, and are used here without restatement. The origin of the quark mixing matrix, as the residual misalignment between the basis in which the Yukawa couplings are diagonal and the basis in which the charged current is diagonal, is derived in Theorem 104.38 from the Higgs couplings of Electroweak Unification and the Higgs Boson; Equation (104.48) is the biunitary diagonalization that produces it. This chapter starts from that result and asks what the matrix actually is.
Every reference cited below writes its formulae with \(\hbar=c=1\), and this book does not: editorial rule 3 requires SI throughout, so \(\hbar\) and \(c\) appear in every equation and in every intermediate step here. The dictionary is mechanical, and is given once so that a reader can move between the two. A quantity quoted in the literature as an energy to the power \(n\) becomes, in SI, that energy divided by the appropriate power of \(\hbar c\): a length is \(\hbar c/E\), a time is \(\hbar/E\), a squared mass is \(E^{2}/c^{4}\). The numerical bridge is
Two conventions deserve naming because they recur. The Fermi constant is universally quoted as \(G_{F}/(\hbar c)^{3} =1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2}\), which is not \(G_{F}\); the constant itself is \(G_{F}=1.43585\times 10^{-62}\,\mathrm{J}\,\mathrm{m}^{3}\), as Definition 101.8 sets out. And a neutrino squared-mass splitting is quoted as \(\Delta m^{2}\) in \(\mathrm{eV}^{2}\) when the quantity meant is \(\Delta m^{2}c^{4}\), an energy squared. Both usages are retained below with the powers of \(\hbar\) and \(c\) written out, because the derivations are carried out in SI and not translated into it afterwards. No derivation in this chapter takes place inside this remark.
The discovery of new quantum numbers
V particles
In October 1946 Rochester and Butler exposed a cloud chamber in a magnetic field to the cosmic radiation at sea level and, in some five thousand photographs, found two in which a track forked without any incoming charged particle at the vertex [Rochester:1947]. The first showed a neutral object decaying into two oppositely charged particles — the forked “V” that named the class — and the second a charged object changing direction with the emission of something neutral. Neither could be a scattering: momentum conservation at the vertex failed unless an unseen neutral particle was created, and the reconstructed mass of the decaying object, about \(500\,\mathrm{MeV}/c^{2}\) in the neutral case, matched nothing then known. The cosmic-ray context, and the instrumental history that made such an exposure possible, belong to Cosmic Rays and Astroparticle Physics.
What made the new particles a problem rather than merely a discovery was a mismatch of time scales, and it is worth putting numbers on it because the whole of flavour physics grows out of that mismatch.
The new particles are produced in nucleon–nucleon and pion–nucleon collisions with cross sections of hadronic size, of order \(10^{-30}\,\mathrm{m}^{2}\), which is the strength of the strong interaction; yet they decay with lifetimes of order \(10^{-10}\,\mathrm{s}\), which is thirteen to fourteen orders of magnitude longer than a strong decay of the same energy release. The \(\Lambda^{0}\) is the cleanest case: its measured width is \(2.515\times 10^{-15}\,\mathrm{GeV}\), corresponding to
against \(\tau_{\Delta}=\hbar/\Gamma_{\Delta}=5.6\times 10^{-24}\,\mathrm{s}\) for the \(\Delta(1232)\), a genuine strong decay of comparable energy release [Navas:2024]. Rests on Equation (101.53).
Derivation. Derives Phenomenon 103.2. The natural time scale for a decay driven by the interaction that also binds the hadron is the time light takes to cross it, \(R/c=10^{-15}\,\mathrm{m}/c=3.3\times 10^{-24}\,\mathrm{s}\), and \(\tau_{\Delta}\) is within a factor of two of it: the \(\Delta\) decays as fast as it can. The ratio \(\tau_{\Lambda}/\tau_{\Delta}=4.7\times 10^{13}\) therefore cannot be absorbed into phase space, which varies as a power of the energy release and can account for two or three orders of magnitude at most, nor into a centrifugal barrier, which for the observed spins is absent. The only available conclusion is that the interaction responsible for the decay is not the interaction responsible for the production. Comparing \(\tau_{\Lambda}\) with the muon lifetime \(2.197\times 10^{-6}\,\mathrm{s}\) and allowing for the fifth-power dependence of Equation (101.53) on the energy release identifies the decay interaction as the weak one.
The inference is then sharp. Something is conserved by the interaction that makes these particles and violated by the interaction that destroys them. That is the whole content of the next subsection, and it is the template for every conserved flavour quantum number in this chapter: exact for two of the three forces, broken by the third.
∎A second regularity sharpened the puzzle: the new particles are never produced singly. A \(\pi^{-}\) striking a proton yields \(\Lambda^{0}+K^{0}\) but never \(\Lambda^{0}+\pi^{0}\), although the latter is energetically far easier. Pais named this associated production, and it is exactly what a conserved additive quantum number, carried with opposite sign by the two partners, would produce.
Strangeness
Strangeness \(S\) is an additive quantum number, integer valued, assigned to each hadron so that the electric charge satisfies
with \(I_{3}\) the third component of isospin and \(B\) the baryon number. It is conserved by the strong and electromagnetic interactions and violated by the weak interaction, which changes it by at most one unit in a single transition.
The assignment was made independently by Gell-Mann [GellMann:1953b] and by Nakano and Nishijima [Nakano:1953], from the observed production and decay patterns alone, several years before any dynamical picture existed. Equation (103.3) is the same relation that reappears in Equation (102.1) as a statement about quark content and, in a wider form, in Equation (104.13) as a statement about the hypercharge of an electroweak multiplet. Its three appearances are one relation seen at three depths.
If \(S\) is conserved by the strong interaction and the initial state of a collision of ordinary matter has \(S=0\), then strange particles can only be produced in pairs of opposite strangeness, and the lightest strange hadron of each sign is stable against strong decay. Rests on Definition 103.3 and Phenomenon 103.2.
Derives Proposition 103.4. Additivity and conservation give \(\sum_{\text{final}}S=0\), so a final state containing one particle with \(S=-1\) must contain another with \(S=+1\); a final state \(\Lambda^{0}\pi^{0}\) with \(S=-1\) is forbidden however much phase space favours it, which is what is seen. For the second statement, let \(h\) be the lightest hadron with \(S=-1\). A strong decay must conserve \(S\), so its products together carry \(S=-1\); but every hadron with \(S=-1\) is at least as heavy as \(h\), and any accompanying non-strange hadron has positive mass, so no such final state is open. The decay must therefore proceed through an interaction that violates \(S\), and by Phenomenon 103.2 that interaction is weak. Applying the argument to \(\Lambda^{0}\) (\(S=-1\)) and to \(K^{+}\) (\(S=+1\)) accounts for the long lifetimes of both.
∎The \(K^{+}\) is an isodoublet member with \(I_{3}=+\tfrac{1}{2}\), \(B=0\); Equation (103.3) then requires \(S=+1\). Its antiparticle \(K^{-}\) has \(S=-1\), so \(K^{+}\) and \(K^{-}\) are not members of one isodoublet, and the neutral kaons \(K^{0}\) (\(S=+1\)) and \(\bar{K}^{0}\) (\(S=-1\)) are distinct particles rather than one self-conjugate state — the fact that makes the whole of Section 103.5 possible. The \(\Lambda^{0}\) is an isosinglet, \(I_{3}=0\), \(B=1\), \(Q=0\), hence \(S=-1\). The \(\Omega^{-}\), with \(I_{3}=0\), \(B=1\), \(Q=-e\), requires \(S=-3\): three units of a quantum number that the weak interaction can shed only one at a time, which is why it decays in a cascade and why its lifetime is long even by strange standards. That prediction, and its confirmation, are Phenomenon 102.3.
In quark language, settled two decades later, \(S\) counts the strange quarks with a sign: \(S=-\left(n_{s}-n_{\bar{s}}\right)\), the minus sign an accident of the convention fixed before quarks were known. The same construction repeats for every heavy flavour: charm \(C\), bottom \(\tilde{B}\) and top \(T\) are defined so that Equation (103.3) generalizes to
and each is conserved by the strong and electromagnetic interactions and violated by the weak. Strangeness is not a special quantum number; it is the first member of a family, and the family is the subject of this chapter.
Definition 103.3 is a classification, not a mechanism. It says that some quantity is conserved by two forces and violated by the third; it does not say why, nor how much the violation is, nor why the violating amplitude should be the same size as the one that governs nuclear beta decay. Those questions were answered — the first two by Section 103.2.1, the third only partly, and the deepest of them not at all — and the honest summary is that a selection rule was found in 1953 and the number controlling it was measured in 1963.
Quark mixing
The Cabibbo angle
By 1963 the weak interaction had a universal strength: the coupling extracted from muon decay and from nuclear beta decay agreed to within a few per cent, which is Phenomenon 101.43 — stated there over muon decay, superallowed nuclear beta decay and \(\tau\) decay, the third of which came much later. Strangeness-changing decays did not fit. The semileptonic decay \(\Lambda^{0}\to p\,e^{-}\bar{\nu}_{e}\) proceeds at some five per cent of the rate that the same universal coupling predicts, and \(K^{+}\to\pi^{0}e^{+}\nu_{e}\) likewise. There were two ways to read this: either universality is approximate and the weak interaction has a separate, weaker, strangeness-changing piece, or universality is exact and the strength is shared between the two channels.
Cabibbo took the second reading [Cabibbo:1963], and it is the conceptual step on which everything below rests: the state that couples to the \(W\) is not a state of definite mass.
Semileptonic decays that change strangeness proceed at roughly one twentieth of the rate of their strangeness-conserving counterparts, and a single rotation angle accounts for the whole of the deficit. The measured magnitudes of the first-row mixing elements are
obtained from superallowed nuclear beta decay, from kaon decays and from semileptonic \(B\) decays respectively [Navas:2024]. They satisfy \(\abs{V_{ud}}^{2}+\abs{V_{us}}^{2}+\abs{V_{ub}}^{2}=1\) to about a part in \(10^{3}\), with a residual deficit of a few standard deviations that is at present the sharpest unexplained departure from the Standard Model anywhere in flavour physics. Rests on Phenomenon 101.43 and Definition 103.3.
Derivation. Derives Phenomenon 103.7. Cabibbo's hypothesis is that the down-type state entering the charged current is a rotation of the mass eigenstates, \(d'=d\cos\theta_{C}+s\sin\theta_{C}\), carrying the same overall coupling \(G_{F}\) as the leptonic current [Cabibbo:1963]. A strangeness-conserving amplitude then acquires a factor \(\cos\theta_{C}\) and a strangeness-changing one a factor \(\sin\theta_{C}\), so their rates stand in the ratio \(\tan^{2}\theta_{C}\) and the two effects are not independent: the suppression of the second is precisely the deficit of the first against the leptonic strength, which is what restores universality. Reading \(\sin\theta_{C}=\abs{V_{us}}\) from Equation (103.5) gives \(\theta_{C}=12.96^\circ\) and hence \(\cos\theta_{C}=0.9745\), which is \(\abs{V_{ud}}\) to within one part in a thousand — one angle, measured twice over in unrelated processes. The global fit quoted later as \(\theta_{12}=13.00^\circ\) in Equation (103.22) is the same angle: the \(0.04^\circ\) between the two is the difference between the direct kaon determination of \(\abs{V_{us}}\) used here and the fitted \(\lambda=0.22501\) used there, and nothing else. The rate ratio it predicts, \(\tan^{2}\theta_{C}=0.0530\), is the observed factor of about twenty. That the squares sum to unity is the two-generation case of the unitarity discussed in Section 103.4.3; that they sum to unity only to a part in \(10^{3}\) is a measurement, not an approximation, and it is recorded here rather than smoothed away.
∎The arithmetic behind the last sentence is worth doing explicitly, because the number it produces is the one genuinely open experimental anomaly in quark flavour physics today.
With the values of Equation (103.5),
which is below unity by \(2.3\) standard deviations. Rests on Equation (103.5).
Derives Proposition 103.8. \(\abs{V_{ud}}^{2}=0.948033\), \(\abs{V_{us}}^{2}=0.050310\) and \(\abs{V_{ub}}^{2}=1.46\times 10^{-5}\), summing to \(0.998358\). The uncertainty propagates as \(\delta\Delta=2\left[\left(\abs{V_{ud}}\delta\abs{V_{ud}}\right)^{2} +\left(\abs{V_{us}}\delta\abs{V_{us}}\right)^{2}\right]^{1/2}\), the \(V_{ub}\) term being negligible, which gives \(2\left[(3.12\times 10^{-4})^{2}+(1.79\times 10^{-4})^{2}\right]^{1/2} =7.2\times 10^{-4}\). Hence \(1-\Delta_{\mathrm{CKM}} =1.64\times 10^{-3}\pm7.2\times 10^{-4}\), a deficit of \(2.3\sigma\).
∎The significance of Equation (103.6) depends on which determination of \(\abs{V_{us}}\) is used — the value from \(K\to\pi e\nu\) and the value from the ratio of \(K\to\mu\nu\) to \(\pi\to\mu\nu\) differ by about \(2\sigma\) between themselves — and on the radiative corrections applied to superallowed nuclear beta decay, which were revised in the late 2010s and are what moved this quantity from agreement to tension. Quoted significances between \(2\sigma\) and \(3\sigma\) therefore all appear in the literature and all are defensible. What is not defensible is either extreme: calling it a discovery, or suppressing it. Unitarity of the first row is a prediction with no free parameters, the measurement is at the level of one part in \(10^{3}\), and it does not currently come out right. This treatise records that, and records equally that the most likely resolution is a systematic effect in one of the two inputs.
GIM and the prediction of charm
Cabibbo's rotation solves the charged-current problem and creates a neutral-current one, which Proposition 101.58 states: the same rotated field \(d'\) appears bilinearly in any neutral current, and its cross terms change strangeness with a coefficient \(\sin\theta_{C}\cos\theta_{C}=0.219\), of ordinary size. Nature says otherwise.
The neutral weak current does not change quark flavour. A flavour-changing neutral coupling of ordinary strength would make \(K_{L}\to\mu^{+}\mu^{-}\) comparable in rate with \(K^{+}\to\mu^{+}\nu_{\mu}\); the measured branching fractions are \(6.84\times 10^{-9}\) and \(0.636\) respectively [Navas:2024], a suppression by eight orders of magnitude, and the mass differences of the neutral kaon and \(B\) systems are correspondingly tiny. The suppression is not a numerical accident of one channel: no tree-level neutral flavour transition has ever been observed in any system, in any generation. Rests on Theorems 101.59 and 104.38.
Derivation. Derives Phenomenon 103.10. The tree-level statement is Theorem 101.59, proved there for two generations and extended to three in Theorem 104.38(ii): the neutral-current coefficient depends only on \(T^{3}\) and \(Q\), hence is the same for every generation, hence is proportional to the identity in generation space, and \(U^{\dagger}\identity U=\identity\) for any unitary \(U\). Nothing survives the rotation to the mass basis. This is the Glashow–Iliopoulos–Maiani mechanism [Glashow:1970], and it is a consequence of unitarity together with the requirement that every fermion of a given charge and chirality sit in a complete weak multiplet. The two hypotheses are what a search for flavour-changing neutral currents actually tests.
At tree level the cancellation is exact and the observed rate would be zero, which is also wrong. The measured \(6.84\times 10^{-9}\) is small but not absent, and it is produced at one loop, where the cancellation is only partial. That calculation is Theorem 103.11 below, and it is where the mechanism stops being an algebraic identity and starts predicting a number.
∎The one-loop statement is the important one historically, because it converts a null result into a mass.
Let \(A(m_{i})\) be the amplitude for a flavour-changing neutral process, computed with a single up-type quark of mass \(m_{i}\) circulating in the loop, and let \(\lambda_{i}=V_{id}^{\phantom{*}}V_{is}^{*}\) be the associated product of mixing factors. Then the physical amplitude is
so any part of \(A(m_{i})\) that is independent of \(m_{i}\) cancels identically. Expanding \(A\) in the small parameter \(x_{i}=m_{i}^{2}c^{4}/\left(M_{W}c^{2}\right)^{2}\) leaves
so the surviving amplitude is suppressed by the ratio of a squared mass difference of the internal quarks to the squared \(W\) mass. A flavour-changing neutral amplitude therefore measures the mass of the heaviest quark circulating in the loop whose mixing factor \(\lambda_{i}\) is not itself negligible. Rests on Theorem 101.62.
Derives Theorem 103.11. The sum rule \(\sum_{i}\lambda_{i} =\sum_{i}V_{id}^{\phantom{*}}V_{is}^{*}=\delta_{ds}=0\) is the \((d,s)\) column relation of Theorem 101.62. Write \(A(m_{i})=A(0)+A'(0)x_{i}+O(x_{i}^{2})\), where the expansion is legitimate for \(m_{i}c^{2}\ll M_{W}c^{2}\) and where \(A(0)\) is by construction independent of \(i\). Then
which is Equation (103.8). Two remarks fix its meaning. First, the cancellation is not a fine tuning: it is the statement that the loop, evaluated with degenerate internal masses, is a flavour-diagonal operator, and the mixing matrix cannot rotate a multiple of the identity into anything else. Second, the surviving term is a difference, since \(\sum_{i}\lambda_{i}=0\) allows \(m_{i}^{2}\) to be shifted by any constant; writing the sum over the first two generations alone gives \(\lambda_{c}\left(m_{c}^{2}-m_{u}^{2}\right)c^{4}\), and the mass of the light quark drops out of everything.
∎The \(K_{L}\)–\(K_{S}\) mass difference is generated by the box diagram in which two \(W\) bosons and two internal up-type quarks convert \(K^{0}\) into \(\bar{K}^{0}\). Keeping the charm term of Equation (103.8), which dominates for a system built from \(d\) and \(s\) quarks, the short-distance contribution is
where \(f_{K}=155.7\,\mathrm{MeV}/c^{2}\) is the kaon decay constant and \(B_{K}=0.717\) the lattice matrix element of the four-quark operator [Navas:2024], computed by the methods of Section 102.7.1. The dimensions balance as follows: \(G_{F}/(\hbar c)^{3}\) carries \(/\mathrm{J}^{2}\), so its square carries \(\mathrm{J}^{-4}\), and the five remaining factors — \(f_{K}^{2}c^{4}\), \(m_{K}c^{2}\) and \(\left(m_{c}c^{2}\right)^{2}\) — supply \(\mathrm{J}^{5}\), leaving an energy, which is what \(\Delta m_{K}c^{2}\) is. Inserting \(G_{F}/(\hbar c)^{3}=1.1663787\times 10^{-5}\,/\mathrm{GeV}^{2}\), \(m_{K}c^{2}=497.611\,\mathrm{MeV}\), \(\abs{V_{cd}V_{cs}}=\sin\theta_{C}\cos\theta_{C}=0.2186\) and \(m_{c}c^{2}=1.273\,\mathrm{GeV}\) gives
against the measured \(\Delta m_{K}c^{2}=3.484\times 10^{-6}\,\mathrm{eV}\), a fractional mass splitting of \(7.0\times 10^{-15}\) and one of the smallest quantities ever measured [Navas:2024]. The box accounts for \(44\%\) of it; the short-distance QCD correction raises this to about \(60\%\) and the remainder is long-distance, that is, two-pion and other hadronic intermediate states that no short-distance expansion reaches.
Run backwards, before the charm quark was known, the same formula bounds its mass: saturating the observed \(\Delta m_{K}\) with the box requires \(m_{c}c^{2}\approx1.9\,\mathrm{GeV}\), and requiring merely that the box not exceed the measurement gives an upper bound of the same order. Gaillard and Lee obtained a few \(\mathrm{GeV}/c^{2}\) by this argument in 1974 [Gaillard:1974]. The particle was found that November at \(3.097\,\mathrm{GeV}/c^{2}\), which for a bound state of two charm quarks is exactly where a \(1.5\,\mathrm{GeV}/c^{2}\) quark puts it.
The short-distance two-\(W\) box amplitude for the kaon mass difference. What is carried out inline is the GIM structure of the internal sum, the dimensional check of every factor, and the numerical evaluation; the loop integral itself — the box function in its limit of internal mass small against the \(W\) mass, together with the reduction of the four-quark operator to the decay constant and the bag parameter — is not derived here and belongs in Appendix A.
Remark 101.60 records the structure of the inference: the input is the absence of a process and the output is the existence of a particle. Example 103.12 is what makes it a measurement rather than a hunch. The absence of \(K_{L}\to\mu^{+}\mu^{-}\) at ordinary strength requires the mechanism; the smallness of \(\Delta m_{K}\), which is not zero, then fixes the scale of the particle the mechanism needs. Glashow, Iliopoulos and Maiani proposed the fourth quark in 1970 on the first fact [Glashow:1970]; the second gave its mass four years later, months before it was found. The same logic, applied to \(B^{0}\)–\(\bar{B}^{0}\) mixing in 1987, gave the top mass a decade before the top was observed: see Section 103.5.3.
The November revolution
On 11 November 1974 two groups announced the same particle. At Brookhaven, Ting's group measured \(e^{+}e^{-}\) pairs from \(p+\mathrm{Be}\) collisions and found a peak at \(3.1\,\mathrm{GeV}/c^{2}\) whose width was consistent with the apparatus resolution and nothing more [Aubert:1974]. At SPEAR, Richter's group scanned the \(e^{+}e^{-}\) annihilation cross section and found it rise by two orders of magnitude over a few \(\mathrm{MeV}\) of beam energy [Augustin:1974]. The particle is now written \(J/\psi\), with
[Navas:2024]. It is the width that mattered.
The \(J/\psi\) lies far above the threshold for decay into ordinary hadrons and carries the quantum numbers of the photon, so nothing forbids it from decaying strongly. Yet its width is \(92.6\,\mathrm{keV}\), against \(149\,\mathrm{MeV}\) for the \(\rho(770)\) and \(4.25\,\mathrm{MeV}\) for the \(\phi(1020)\) [Navas:2024] — narrower than the \(\rho\) by a factor \(1.6\times 10^{3}\) and than the \(\phi\) by a factor \(46\). Rests on Theorem 105.9 and Phenomenon 102.41.
Derivation. Derives Phenomenon 103.14. The three resonances differ in one respect: how their valence quarks can reach the final state. The \(\rho\) is \(u\bar{u}\) and \(d\bar{d}\), and its decay \(\rho\to\pi\pi\) can proceed with both valence quarks carried into the final-state hadrons, one \(q\bar{q}\) pair being created from the vacuum. The \(\phi\) is \(s\bar{s}\); its dominant decay \(\phi\to K\bar{K}\) likewise carries both valence quarks through, but only just, the phase space being \(32\,\mathrm{MeV}\), and its decay to three pions — which requires the \(s\bar{s}\) pair to annihilate — is suppressed despite the far larger phase space. The \(J/\psi\) is \(c\bar{c}\) below the threshold for a pair of charmed mesons, so every decay to light hadrons requires the \(c\) and \(\bar{c}\) to annihilate.
That is the rule of Okubo [Okubo:1963], Zweig [Zweig:1964] and Iizuka [Iizuka:1966]: a decay whose quark-line diagram falls apart into two disconnected pieces is suppressed. Stated as an empirical rule it explains the \(\phi\); stated dynamically it explains the \(J/\psi\) and the size of the suppression. The disconnected diagram requires at least three gluons, because one gluon carries colour and cannot make a colour singlet, and two gluons cannot make a state of odd charge conjugation (Theorem 105.9 is the same argument for photons). The rate is therefore of order \(\alpha_{s}^{6}\) evaluated at a scale set by the charm mass, and by Phenomenon 102.41 that coupling is already small: \(\alpha_{s}\left(m_{c}c^{2}\right)\approx0.35\), so \(\alpha_{s}^{6}\approx1.8\times 10^{-3}\) against a light-quark process running at \(\alpha_{s}\approx1\). Asymptotic freedom, discovered the year before, is what turns an empirical rule into a number of the right order.
∎The interpretation was settled within months by three further facts. The system has radial and orbital excitations with the spacings of a non-relativistic two-body bound state, treatable by the methods of Approximation Methods and testing the static potential of Equation (102.69) directly, as Remark 102.63 sets out. Above \(3.73\,\mathrm{GeV}\) the annihilation cross section rises to a new plateau, consistent with the opening of a channel with one unit of a new quantum number in each of two final-state hadrons — open charm. And the ratio \(R\) of Phenomenon 102.15 steps up by \(4/3\), which is \(3\times(2/3)^{2}\): three colours of a quark of charge \(+\tfrac{2}{3}e\). The fourth flavour predicted by Example 103.12 was there, with the charge the GIM mechanism required.
Three generations
Kobayashi and Maskawa
In 1973, with three quarks known and a fourth conjectural, Kobayashi and Maskawa asked what it would take for the weak interaction to violate \(CP\) [Kobayashi:1973]. The answer is a counting argument and nothing else, which is what makes it remarkable.
The counting is carried out twice in this book, in Proposition 101.63 and in Theorem 105.25: an \(N\times N\) unitary matrix carries \(N(N-1)/2\) angles and \((N-1)(N-2)/2\) phases that survive rephasing of the quark fields. It is not repeated here. What belongs here is the consequence and the inference.
For \(N=2\) the mixing matrix can be brought to real orthogonal form by rephasing, and the charged-current interaction is then invariant under \(CP\); for \(N=3\) one phase survives and \(CP\) invariance holds only if that phase is \(0\) or \(\pi\). Rests on Equations (101.67), (101.68) and (105.16).
Derives Proposition 103.15. Under \(CP\) the charged-current term of Equation (101.67) maps to itself with \(V\mapsto V^{*}\) (Equation (105.16) gives the transformation of the bilinear). If \(V\) can be made real by a rephasing that leaves the rest of the Lagrangian invariant, then \(V=V^{*}\) in that basis and the interaction is \(CP\) symmetric. Equation (101.68) with \(N=2\) gives \((N-1)(N-2)/2=0\) surviving phases, so the Cabibbo matrix is real and \(CP\) is conserved however \(\theta_{C}\) is chosen; with \(N=3\) one phase survives and cannot be removed, so \(V\neq V^{*}\) in any basis unless the phase is a multiple of \(\pi\).
∎The measurement that forced this was a single branching ratio of \(2\times 10^{-3}\) [Christenson:1964]. From it, and from the requirement that the theory be renormalizable and its gauge structure unchanged, Kobayashi and Maskawa concluded that two further quarks exist. Both were found: the bottom in 1977 [Herb:1977] and the top in 1995 [Abe:1995] [Abachi:1995]. It is worth being precise about what was and was not predicted. The number three was predicted, as a minimum. The masses were not, and are still not predicted by anything. What the counting delivers is that the observation of \(CP\) violation is, on the stated hypotheses, a measurement of the number of generations — and the completely independent measurement of that number from the \(Z\) width, Phenomenon 103.22, agrees with it.
Two alternatives existed in 1973 and both are now excluded. A right-handed charged current with a complex coupling could violate \(CP\) with two generations, and is excluded by the helicity measurements of Section 101.3.3 and by the direct violation of Section 103.5.2; a superweak interaction acting only in the mixing could violate \(CP\) with two generations, and is excluded by \(\varepsilon'/\varepsilon\neq0\). The Kobayashi–Maskawa mechanism is not merely the simplest surviving option, it is the only one of the three that survived.
The tau lepton
In 1975 Perl's group at SPEAR reported events of the form \(e^{+}e^{-}\to e^{\pm}\mu^{\mp}+\text{missing energy}\), with no other charged particle and no photon [Perl:1975]. Two features made them a discovery rather than a background. The final state violates electron and muon number separately unless at least two unobserved neutral particles are present; and the events appear only above a threshold near \(3.6\,\mathrm{GeV}\) in the total energy, which is the signature of pair production of a new particle of mass about \(1.8\,\mathrm{GeV}/c^{2}\).
There exists a charged lepton of mass \(m_{\tau}c^{2}=1776.93(9)\,\mathrm{MeV}\) and lifetime \(\tau_{\tau}=2.903\times 10^{-13}\,\mathrm{s}\) [Navas:2024], with its own neutrino, whose charged-current interaction has the same strength and the same \(V-A\) structure as the electron's and the muon's. Rests on Equation (101.53).
Derivation. Derives Phenomenon 103.17. The universality claim is testable to a fraction of a per cent by one number, and the test is the cleanest in the whole of flavour physics because it involves no hadrons at all. Equation (101.53) gives the rate of a purely leptonic charged-current decay,
whose fifth power of the mass is the dominant factor. Since the \(\tau\) has hadronic decay channels as well, the prediction is for the branching fraction
Inserting \(\tau_{\tau}=2.9034\times 10^{-13}\,\mathrm{s}\), \(\tau_{\mu}=2.1969811\times 10^{-6}\,\mathrm{s}\), \(m_{\tau}c^{2}=1776.93\,\mathrm{MeV}\) and \(m_{\mu}c^{2}=105.6583755\,\mathrm{MeV}\), so that \(\left(m_{\tau}/m_{\mu}\right)^{5}=1.3453\times 10^{6}\) and the phase-space ratio is \(1.00019\), gives \(\mathcal{B}=0.1778\). The measured value is \(\mathcal{B}\left(\tau\to e\nu\bar{\nu}\right)=0.1782(4)\) [Navas:2024]. Agreement to two parts in a thousand, over a factor \(1.3\times 10^{6}\) in rate, is what “the same coupling” means quantitatively; expressed as a coupling ratio, \(g_{\tau}/g_{\mu}\) differs from unity by less than \(2\times 10^{-3}\).
∎The phase-space function \(f(x)=1-8x+8x^{3}-x^{4}-12x^{2}\ln x\) for a purely leptonic charged-current decay with a massive daughter lepton. The muon-decay rate inherited from the weak-interactions chapter is derived there only for a massless daughter, where \(f(0)=1\); the massive case is the one used above, and the three-body phase-space integral that produces it belongs in Appendix A.
The \(\tau\) neutrino was assumed for a quarter of a century and observed directly only in 2000, by the DONUT experiment, which recorded four \(\nu_{\tau}\) charged-current interactions producing a short track with a kink — the \(\tau\) decaying within a millimetre of its production point [Kodama:2001]. Its appearance in an oscillated beam, which is a different measurement, came in 2015 [Agafonova:2015] and belongs to Experiment: Neutrino Oscillations.
The \(\tau\) is the only lepton heavy enough to decay into hadrons, and that makes it a laboratory for two things this book measures elsewhere. Its hadronic branching fraction, compared with the leptonic one, gives \(\alpha_{s}\) at a scale of \(1.78\,\mathrm{GeV}\) — the lowest scale at which the coupling is extracted — and the running of that value up to \(M_{Z}c^{2}\) is one of the entries in Table 102.2. Its decays into \(\pi\pi\) measure the same hadronic vacuum polarization that limits the Standard Model prediction of the muon anomaly, which belongs to Section 112.6.3. Neither use was anticipated when the particle was found.
The bottom quark
Lederman's group at Fermilab measured the mass spectrum of muon pairs produced in \(400\,\mathrm{GeV}\) proton–nucleus collisions and found a structure near \(9.5\,\mathrm{GeV}/c^{2}\) standing above a smooth continuum [Herb:1977]. It resolved into the \(\Upsilon(1S)\) and \(\Upsilon(2S)\), with \(m_{\Upsilon}c^{2}=9.46040(10)\,\mathrm{GeV}\) and \(\Gamma_{\Upsilon}=54.0(13)\,\mathrm{keV}\) [Navas:2024]: the same signature as the \(J/\psi\) — a resonance far too narrow for a strong decay — one flavour further up.
Three consequences organize the rest of this chapter.
First, bottomonium and charmonium together test the static potential of Equation (102.69) in a way neither can alone. The reduced masses differ by a factor of about three, yet the spacing between the ground state and the first radial excitation is nearly the same in the two systems, \(0.589\,\mathrm{GeV}/c^{2}\) against \(0.563\,\mathrm{GeV}/c^{2}\); Remark 102.63 draws out why only an intermediate, roughly logarithmic, potential does that.
Second, the bottom quark is the one heavy flavour that forms mesons with a long lifetime on the scale of a modern detector. A \(B\) meson lives \(1.52\times 10^{-12}\,\mathrm{s}\), so at the energies of an asymmetric-energy collider it travels a few hundred micrometres before decaying — far enough to separate its decay vertex from the production point with a silicon detector, and hence to tag which of the two mesons produced in an event decayed first. That single instrumental fact is what makes the time-dependent measurements of Section 103.5.3 possible; without it the \(CP\) asymmetry integrates to a much smaller number.
Third, the mixing of the third generation to the first two is small — the \(b\) quark is the third generation, and what is small is its coupling out of it. The direct determinations are \(\abs{V_{cb}}=0.0408(14)\) and \(\abs{V_{ub}}=0.00382(20)\) (Table 103.1), and that smallness is why \(B\) mesons live as long as they do. The lifetime and the mixing element are the same measurement seen twice.
The top quark
The top quark was expected from 1977, when the \(b\) was found with weak isospin \(-\tfrac{1}{2}\) and therefore in need of a partner, and was not observed until 1995. In between it was measured indirectly, which is the part of the story that belongs to a treatise of evidence.
Electroweak radiative corrections depend on the top mass quadratically, through the correction \(\Delta\rho\) to the \(\rho\) parameter of Equation (104.52), and the LEP precision data therefore determined \(m_{t}\) before any top quark was produced. The value inferred from the fit in 1994 was \(178\pm20\,\mathrm{GeV}/c^{2}\); the direct measurements published the following year gave \(176\pm13\,\mathrm{GeV}/c^{2}\) [Abe:1995] and \(199\pm30\,\mathrm{GeV}/c^{2}\) [Abachi:1995], and the current world average is \(m_{t}c^{2}=172.57(29)\,\mathrm{GeV}\) [Navas:2024]. Rests on Equation (104.52), Proposition 104.43 and Theorem 104.40.
Derivation. Derives Phenomenon 103.19. The mechanism is Proposition 104.43: a doublet whose two members have different masses breaks the custodial symmetry of Theorem 104.40, and the resulting shift in the ratio of \(W\) and \(Z\) self-energies grows as the squared mass splitting, \(\Delta\rho\propto\left(m_{t}^{2}-m_{b}^{2}\right)c^{4}\). Because the dependence is quadratic rather than logarithmic, a measurement of the \(Z\) lineshape and asymmetries at the per-mille level of Section 104.6.3 pins \(m_{t}\) to about ten per cent. That is a real prediction and it was confirmed: the indirect and direct values agree within their uncertainties, which is a test of the loop structure of the theory and not merely of its tree-level content. The same fit subsequently bounded the Higgs mass, whose contribution is only logarithmic and therefore far weaker — and that bound too was confirmed, in Phenomenon 104.59.
∎Two properties of the top are peculiar to it and both follow from its mass.
The top width is \(\Gamma_{t}=1.42\,\mathrm{GeV}\) [Navas:2024], so its lifetime is
which is shorter than the time \(\hbar/\left(\Lambda_{\mathrm{QCD}}\right) \approx3.3\times 10^{-24}\,\mathrm{s}\) that QCD needs to dress a colour charge into a hadron. The top therefore never forms a bound state. Rests on Equation (102.45), Equation (103.1) and Phenomenon 102.65.
Derives Proposition 103.20. \(\Lambda_{\mathrm{QCD}}\approx0.2\,\mathrm{GeV}\) by Equation (102.45), and the hadronization time is the inverse of that energy in the sense of Equation (103.1), \(\hbar/0.2\times 10^{9}\,\mathrm{eV}=3.3\times 10^{-24}\,\mathrm{s}\): seven times Equation (103.14). The width itself scales as \(m_{t}^{3}\), since \(t\to bW\) is a two-body decay into a real \(W\) with the top mass setting both the coupling's phase space and the longitudinal polarization enhancement, so the inequality gets stronger with mass and is a statement about the top specifically. The consequences are experimental and considerable: the top's spin information survives into its decay products, unrandomized by hadronization, and the object reconstructed in a detector is a bare colour triplet in flight — never, by Phenomenon 102.65, an asymptotic state.
∎By Theorem 104.38 the coupling of a fermion to the Higgs field is \(\hat{y}_{f}=\sqrt{2}\,m_{f}c^{2}/E_{v}\) with \(E_{v}=246.22\,\mathrm{GeV}\), so
a dimensionless number equal to one within its uncertainty, while the electron's is \(2.9\times 10^{-6}\). Table 104.2 lists the full range. No principle in the Standard Model relates them, and that the largest of them sits at unity is the kind of fact that looks like a clue and has never been cashed as one. It is listed as open in Section 133.1.
How many generations
The part of the width of the \(Z\) resonance that no observed final state accounts for counts the neutrino species light enough for the \(Z\) to decay into them. Combining the scans of the four LEP experiments,
[Schael:2006]: two standard deviations below three, and more than a hundred standard deviations below four. The count was already unambiguous in the first scan taken [Decamp:1989]. Any further neutrino must therefore be heavier than \(m_{Z}c^{2}/2\) or must not couple to the \(Z\) at all. Rests on Equation (101.75).
Derivation. Derives Phenomenon 103.22. The invisible width is the total width minus the sum of the widths into observed final states, and the species count is that quantity divided by the width into one neutrino pair. The latter is not measured; it is computed, and the computation is short.
By Equation (101.75) the \(Z\) couples to a fermion \(f\) through vector and axial couplings
and the partial width of a vector boson of mass \(M_{Z}\) into a pair of fermions light compared with it is
\(N_{c}^{f}\) being \(3\) for quarks and \(1\) for leptons. The dimensional check is worth doing once: \(G_{F}/(\hbar c)^{3}\) carries \(/\mathrm{J}^{2}\), multiplying it by an energy cubed leaves an energy, and the remaining \(\hbar^{-1}\) converts that energy into a rate — so Equation (103.17) is \(/\mathrm{s}\), and \(\hbar\Gamma\) is the width in energy units usually quoted.
For a neutrino, \(T^{3}=+\tfrac{1}{2}\) and \(Q=0\), so \(g_{V}=g_{A}=\tfrac{1}{2}\) and the bracket is \(\tfrac{1}{2}\):
the full one-loop value being \(167.2\,\mathrm{MeV}\). For a charged lepton, \(T^{3}=-\tfrac{1}{2}\) and \(Q=-1\), so with \(\sin^{2}\theta_{W}=0.2312\) — the \(\overline{\mathrm{MS}}\) value at the scale \(M_{Z}c^{2}\) [Navas:2024], the fourth digit being meaningless without that scheme and that scale, which is why Equation (104.62) quotes the same quantity to three figures — the bracket is \(\left(-0.0376\right)^{2}+\left(-\tfrac{1}{2}\right)^{2}=0.2514\) and \(\hbar\Gamma_{\ell\ell}=83.4\,\mathrm{MeV}\), against the measured \(83.985(86)\,\mathrm{MeV}\) [Schael:2006] — the \(0.7\%\) shortfall being exactly the radiative correction the tree-level formula omits. The count is then
the tree-level version of which is \(0.5/0.2514=1.9888\), differing from the full value by one part in a thousand. With the measured \(\Gamma_{\mathrm{inv}}=499.0(15)\,\mathrm{MeV}\) and \(\hbar\Gamma_{\ell\ell}=83.985(86)\,\mathrm{MeV}\) [Schael:2006], the ratio is \(499.0/83.985=5.942\), and dividing by \(1.9912\) gives \(N_{\nu}=2.984\), which is Equation (103.15).
The residual \(2\sigma\) deviation from three deserves a word, because it is not a hint of a fourth neutrino. It is comparable with the radiative corrections just displayed, and the largest of those enters not the widths but the luminosity: the \(t\)-channel photon exchange in small-angle Bhabha scattering, which is what LEP counted to normalize every cross section. A re-analysis of that effect published in 2019 — not in this book's bibliography, and therefore named without a citation rather than cited to a key that does not exist — raised the central value to \(2.996\pm0.007\). The measurement counts three, and the apparent tension was an artefact of the normalization, not of the neutrino sector.
∎The tree-level partial width of the \(Z\) into a fermion pair. The formula is stated above with its dimensional check and then used to count the neutrino species; its derivation from the neutral-current Lagrangian — the squared matrix element summed over final and averaged over initial polarizations, times the two-body phase space — is not carried out here and belongs in Appendix A.
Equation (103.15) counts neutrinos that are active — coupling to the \(Z\) with the standard strength — and light, meaning \(m_{\nu}c^{2}<M_{Z}c^{2}/2\). A fourth neutrino evading both conditions is not excluded by it, and the searches for one, with their null results, belong to Section 114.6.4, which reserves them. What the count does exclude is a fourth generation of the observed kind, since a fourth charged lepton would come with a light active partner. An independent bound of the same strength comes from primordial nucleosynthesis, where the expansion rate at freeze-out — and hence the neutron-to-proton ratio frozen into the light elements — depends on the number of relativistic species. The light-element floor itself is Phenomenon 48.10; that phenomenon states the abundances as a one-parameter family in the baryon-to-photon ratio and does not yet extract a species count from them, so the nucleosynthesis determination of the number of neutrinos is owed to Evidence-Based Cosmology rather than held there. It agrees with Equation (103.15), which is what makes the number three a fact about Nature rather than about the \(Z\).
And it is a fact without an explanation. Nothing in the Standard Model fixes the number of generations; three is measured, twice, by completely different means, and predicted by nothing. It is listed in Section 133.1.
The CKM matrix
Structure and parametrization
\(V_{\mathrm{CKM}}=U_{u}^{\dagger}U_{d}\) is the unitary \(3\times3\) matrix of Theorem 104.38 relating the down-type states that couple to the \(W\) to the down-type states of definite mass,
Its entries are dimensionless. Four of its nine complex entries are physically independent: three angles and one phase, by Equation (101.68).
The convention that the matrix acts on the down-type states rather than the up-type ones is arbitrary and universal; it makes \(V_{\mathrm{CKM}}\) the object appearing in Equation (101.67). Two parametrizations are in use, and each is right for a different purpose.
With \(c_{ij}=\cos\theta_{ij}\), \(s_{ij}=\sin\theta_{ij}\) and one phase \(\delta\),
The three angles may be taken in the first quadrant and \(\delta\) in \([0,2\pi)\) without loss of generality.
Derivation of the form. Derives Definition 103.25. Write \(V=R_{23}(\theta_{23})\,\Phi(\delta)\,R_{13}(\theta_{13})\, \Phi(-\delta)\,R_{12}(\theta_{12})\), where \(R_{ij}\) is a real rotation in the \(ij\) plane and \(\Phi(\delta)=\diag(1,1,\ee^{\ii\delta})\). Any unitary \(3\times3\) matrix can be brought to this form because such a matrix has \(9\) real parameters — the CKM matrix is unitary with no determinant condition imposed, so the count is that of \(\U(3)\) and not of \(\SU(3)\), which has \(8\) — of which \(3\) are the angles of the \(\SO(3)\) subgroup and \(6\) are phases, and \(5\) of the phases are removed by rephasing the quark fields, per Proposition 101.63. Multiplying out the five factors gives Equation (103.21). The placement of \(\delta\) is a convention: any rephasing moves it around the matrix without changing a single observable, since every observable is a rephasing invariant of the type of Equation (105.77). That the phase can be attached to \(V_{ub}\), the smallest element, is what makes the standard parametrization convenient and what makes \(CP\) violation small even though \(\delta\) is of order unity.
∎The measured angles are
[Navas:2024]: a strong hierarchy in the angles and a phase of ordinary size. These four numbers are not independent of the four quoted below — they are the same measurement in the other parametrization, and Remark 103.27 carries out the conversion, because a book that quoted both from separate compilations would sooner or later quote two values of one quantity. The hierarchy in the angles is the content of that second parametrization.
Set \(\lambda=s_{12}\), \(A\lambda^{2}=s_{23}\) and \(A\lambda^{3}\left(\rho-\ii\eta\right)=s_{13}\ee^{-\ii\delta}\). Then to third order in \(\lambda\),
with the measured values [Navas:2024]
where \(\bar{\rho}+\ii\bar{\eta} =\left(\rho+\ii\eta\right)\left(1-\tfrac{1}{2}\lambda^{2}\right) +O(\lambda^{4})\) are the convention-independent combinations that survive to higher order [Wolfenstein:1983].
Derivation. Derives Definition 103.26. Expand Equation (103.21) with \(s_{12}=\lambda\), \(s_{23}=A\lambda^{2}\), \(s_{13}\ee^{-\ii\delta}=A\lambda^{3}(\rho-\ii\eta)\) and \(c_{ij}=1-\tfrac{1}{2}s_{ij}^{2}+\ldots\), keeping terms to \(O(\lambda^{3})\) in each entry. The \((1,1)\) entry is \(c_{12}c_{13}=1-\tfrac{1}{2}\lambda^{2}+O(\lambda^{4})\), the \((1,2)\) entry is \(s_{12}c_{13}=\lambda+O(\lambda^{5})\), and the \((2,1)\) entry is \(-s_{12}c_{23}-c_{12}s_{23}s_{13}\ee^{\ii\delta} =-\lambda+O(\lambda^{6})\), so the upper-left \(2\times2\) block is a rotation through the Cabibbo angle up to \(O(\lambda^{4})\). The \((3,1)\) entry is \(s_{12}s_{23}-c_{12}c_{23}s_{13}\ee^{\ii\delta} =A\lambda^{3}-A\lambda^{3}(\rho+\ii\eta) =A\lambda^{3}\left(1-\rho-\ii\eta\right)\) to this order, which is the entry carrying the interesting phase. Note that the expansion is not a mathematical approximation imposed for convenience: \(\lambda^{2}\) is \(0.0506\) and \(\lambda^{4}\) is \(2.6\times 10^{-3}\), so the hierarchy is a property of the measured matrix.
∎Equation (103.22) and Equation (103.24) are one measurement in two coordinate systems, and this book takes the apex as primary because every closure test below is drawn in the \(\left(\bar{\rho},\bar{\eta}\right)\) plane. Inverting Definition 103.26 with \(\rho+\ii\eta=\left(\bar{\rho}+\ii\bar{\eta}\right) \left(1-\tfrac{1}{2}\lambda^{2}\right)^{-1}\) gives \(s_{12}=\lambda\), so \(\theta_{12}=13.003^\circ\); \(s_{23}=A\lambda^{2}=4.182\times 10^{-2}\), so \(\theta_{23}=2.397^\circ\); and \(s_{13}=A\lambda^{3}\sqrt{\rho^{2}+\eta^{2}}=3.732\times 10^{-3}\), so \(\theta_{13}=0.2138^\circ\). Note that this uses the unbarred pair; the third-order expansion of Example 103.28 uses the barred pair and gets \(3.64\times 10^{-3}\) for the same quantity, which is the \(2.5\%\) discussed there. The phase is fixed by the same relation: \(s_{13}\ee^{-\ii\delta}=A\lambda^{3}\left(\rho-\ii\eta\right)\) has the argument of \(\bar{\rho}-\ii\bar{\eta}\), the positive rescaling dropping out, so
That is the same arithmetic that produces the angle \(\gamma\) in Phenomenon 103.31, which is why the two agree to within a twentieth of a degree: reconstructing the exact unitary matrix from \(\left(\theta_{12},\theta_{23},\theta_{13},\delta\right)\) and evaluating \(\gamma=\arg\!\left( -V_{ud}V_{ub}^{*}\big/V_{cd}V_{cb}^{*}\right)\) as in Equation (103.31) gives \(65.66^\circ\), a difference of \(0.04^\circ\) which is the \(O(\lambda^{4})\) residue. A reader consulting an older compilation will meet \(\delta\approx68^\circ\); that number belongs with an earlier apex, and quoting it beside Equation (103.24) would put two mutually inconsistent phases in one chapter.
Four combinations follow from Equation (103.24) once the expansion is truncated at third order:
The corresponding numbers from the global fit [Navas:2024] — that is, from the exact parametrization of Equation (103.21), not from the expansion — are \(4.182\times 10^{-2}\), \(3.69\times 10^{-3}\), \(8.57\times 10^{-3}\) and \(3.08\times 10^{-5}\), the last being Equation (105.78).
It is essential to say what this does and does not test. \(\left(\lambda,A,\bar{\rho},\bar{\eta}\right)\) is a reparametrization of the same four degrees of freedom as \(\left(\theta_{12},\theta_{23},\theta_{13},\delta\right)\), so nothing here is a test of the CKM matrix, and Equation (103.26) is not a test at all: Definition 103.26 defines \(A\lambda^{2}=s_{23}\) and \(\abs{V_{cb}}=s_{23}c_{13}\) exactly, so the agreement of the two printed digit for digit is an identity, differing only by \(c_{13}=1-7\times 10^{-6}\). What the remaining three lines measure is the error of the truncation, and that error is predicted in advance. Writing \(\bar{\rho},\bar{\eta}\) in place of \(\rho,\eta\) changes a quantity by a relative \(\lambda^{2}/2=2.5\%\), which is exactly the size of the \(O(\lambda^{5})\) terms the third-order parametrization drops. In Equations (103.27) and (103.29) the apex coordinates enter directly and the full \(2.5\%\) is available; in Equation (103.28) they enter as \(1-\bar{\rho}\), where the same shift is only \(0.5\%\); and in Equation (103.26) they do not enter at all. The line that should be exact is exact, the line that should be accurate to \(0.5\%\) is accurate to \(0.1\%\), and the two that should be wrong by up to \(2.5\%\) are wrong by \(1.4\%\) and \(1.3\%\). The expansion is therefore usable wherever a per-cent statement is wanted, and must be abandoned wherever it is not — which is why every closure test below is stated in terms of the exact matrix.
The genuine test, against numbers that were not inputs to the fit, is against the direct determinations of Table 103.1: \(\abs{V_{cb}}=0.0408(14)\) against \(4.182\times 10^{-2}\) is \(0.7\sigma\), \(\abs{V_{ub}}=0.00382(20)\) against \(3.64\times 10^{-3}\) is \(0.9\sigma\), and \(\abs{V_{td}}=0.0086(2)\) against \(8.58\times 10^{-3}\) is \(0.1\sigma\). Those three comparisons carry the physics; the four lines above carry only the arithmetic of the expansion.
Measured magnitudes
Every element is measured, most of them more than once and by methods sharing no systematics. Table 103.1 collects them with the process each comes from.
| Element | Magnitude | Principal determination |
|---|---|---|
| $\abs{V_{ud}}$ | \(0.97367(32)\) | superallowed $0^{+}\to0^{+}$ nuclear beta decay; free neutron decay |
| $\abs{V_{us}}$ | \(0.2243(8)\) | $K\to\pi\ell\nu$ form factor; ratio of $K\to\mu\nu$ to $\pi\to\mu\nu$; hadronic $\tau$ decay |
| $\abs{V_{ub}}$ | \(0.00382(20)\) | inclusive and exclusive $B\to X_{u}\ell\nu$ |
| $\abs{V_{cd}}$ | \(0.221(4)\) | charm production in neutrino scattering; $D\to\pi\ell\nu$ |
| $\abs{V_{cs}}$ | \(0.975(6)\) | $D\to K\ell\nu$; $D_{s}\to\ell\nu$; $W$ decay to charm |
| $\abs{V_{cb}}$ | \(0.0408(14)\) | inclusive and exclusive $B\to X_{c}\ell\nu$ |
| $\abs{V_{td}}$ | \(0.0086(2)\) | $B^{0}$–$\bar{B}^{0}$ oscillation frequency (loop) |
| $\abs{V_{ts}}$ | \(0.0415(9)\) | $B_{s}$–$\bar{B}_{s}$ oscillation frequency (loop) |
| $\abs{V_{tb}}$ | \(1.014(29)\) | single top production cross section (tree) |
Three features of the table matter more than the numbers.
The third row is measured at one loop and at tree level, and the two agree. \(\abs{V_{td}}\) and \(\abs{V_{ts}}\) come from oscillation frequencies, which are box diagrams of the kind computed in Example 103.12, and therefore assume the loop content of the Standard Model. \(\abs{V_{tb}}\) comes from single top production, which is a tree-level charged-current process and assumes nothing of the sort. Consistency between them is a statement that no unknown particle is circulating in the box.
The determination of \(\abs{V_{cb}}\) and \(\abs{V_{ub}}\) is not settled. Two methods exist for each. The inclusive method sums over all charmed (or charmless) final states and uses an operator product expansion in \(1/m_{b}\); the exclusive method measures one channel, \(B\to D^{*}\ell\nu\) or \(B\to\pi\ell\nu\), and requires a lattice form factor at the kinematic point of interest. For \(\abs{V_{cb}}\) the two differ by about \(2\) to \(3\) standard deviations, inclusive high; for \(\abs{V_{ub}}\) by a similar amount, again inclusive high. The discrepancy has persisted for two decades, has narrowed as lattice calculations improved without disappearing, and is almost certainly a problem of hadronic theory rather than of the electroweak sector — but it is not resolved, and the enlarged uncertainties in Table 103.1 are an acknowledgement of that rather than a measurement.
Nothing in the table is predicted. Four numbers, and the overall scale of the quark masses, are inputs of the Standard Model. The tests of the next subsection are tests of the relations among the nine entries, not of their values.
Unitarity triangle
Unitarity of a \(3\times3\) matrix gives six orthogonality relations between distinct columns or rows. Each is a statement that three complex numbers sum to zero, hence that they close a triangle in the complex plane. Four of the six triangles are extremely flat — one side is smaller than the others by two powers of \(\lambda\) or more — and two have all three sides of the same order. Building the matrix from Equation (103.24) and taking the ratio of the shortest side to the longest gives, for the three column relations \((d,s)\), \((d,b)\), \((s,b)\), the values \(0.002\), \(0.387\), \(0.020\), and for the three row relations \((u,c)\), \((u,t)\), \((c,t)\), the values \(0.001\), \(0.403\), \(0.046\). The \((d,b)\) column triangle, whose sides are all of order \(A\lambda^{3}\), is the one universally called the unitarity triangle and is the one defined below; the \((u,t)\) row triangle is if anything marginally less degenerate and has the same area by Proposition 103.30. It is not drawn separately because it carries almost the same information: its angles, computed from the same matrix, are \(91.60^\circ\), \(64.63^\circ\) and \(23.77^\circ\) against \(91.60^\circ\), \(65.66^\circ\) and \(22.74^\circ\) for the column triangle, so the largest angle is common to both and the other two differ at \(O(\lambda^{2})\).
The relation between the first and third columns of Equation (103.20),
has three terms of order \(A\lambda^{3}\). Dividing through by \(V_{cd}V_{cb}^{*}\) normalizes the middle side to unit length and puts the triangle's base on the real axis from \((0,0)\) to \((1,0)\); the apex is then at \(\left(\bar{\rho},\bar{\eta}\right)\). Its interior angles are
and \(\alpha+\beta+\gamma=\pi\) identically.
The unnormalized triangle of Equation (103.30) has area \(J/2\), with \(J\) the Jarlskog invariant of Definition 105.26; all six unitarity triangles have the same area. \(CP\) is violated if and only if that area is non-zero. Rests on Equation (103.30), Equation (105.77) and Proposition 105.27.
Derives Proposition 103.30. For three complex numbers \(z_{1}+z_{2}+z_{3}=0\), the triangle they close has area \(\tfrac{1}{2}\abs{\mathrm{Im}\left(z_{1}z_{2}^{*}\right)}\), since the cross product of the two vectors \(z_{1}\) and \(-z_{3}=z_{1}+z_{2}\) is \(\mathrm{Im}\left(z_{1}\overline{z_{1}+z_{2}}\right) =\mathrm{Im}\left(z_{1}z_{2}^{*}\right)\). With \(z_{1}=V_{ud}V_{ub}^{*}\) and \(z_{2}=V_{cd}V_{cb}^{*}\) this is \(\mathrm{Im}\left(V_{ud}V_{ub}^{*}V_{cd}^{*}V_{cb}\right)\), which is a Jarlskog quartet and therefore equals \(\pm J\) by Equation (105.77). Since Equation (105.77) makes every quartet equal up to sign, and each of the six relations produces one, all six triangles have area \(J/2\). A degenerate triangle — three collinear complex numbers — has \(J=0\) and by Proposition 105.27 conserves \(CP\). That the area is a rephasing invariant while the individual angles are not is the reason \(J\), and not \(\delta\), is the physical measure of \(CP\) violation.
∎The triangle is now measured many times over, and the measurements are independent in a strong sense: they use different decays, different detectors, and different parts of the theory.
The three angles of Equation (103.31) are measured separately, each from a \(CP\) asymmetry in a different class of \(B\) decay, with results [Navas:2024]
and the two sides, from \(\abs{V_{ub}/V_{cb}}\) and from the ratio of \(B_{s}\) to \(B_{d}\) oscillation frequencies, are measured separately again. All of them meet, within uncertainties, at one apex. Rests on Definition 103.29, Equation (103.31) and Equation (103.24).
Derivation. Derives Phenomenon 103.31. Two checks, one of the angles and one relating angles to sides.
First the angles. Equation (103.32) sums to \(173^\circ\pm6^\circ\), consistent with \(180^\circ\) at \(1.2\) standard deviations; and the constraint \(\alpha+\beta+\gamma=\pi\) is an identity of Definition 103.29, not an input to any of the three measurements. That the three numbers, extracted from \(B\to\pi\pi\) and \(\rho\rho\), from \(B\to J/\psi K_{S}\), and from \(B\to DK\) respectively, sum correctly is a non-trivial test.
Now the apex. Definition 103.29 places it at \(\left(\bar{\rho},\bar{\eta}\right)\), so
Inserting Equation (103.24) gives \(\beta=22.7^\circ\), hence \(\sin2\beta=0.713\), and \(\gamma=65.7^\circ\). The directly measured values are \(\sin2\beta=0.70\pm0.02\), from the time-dependent asymmetry in \(B^{0}\to J/\psi K_{S}\) [Aubert:2001] [Abe:2001], and \(\gamma=66^\circ\pm3^\circ\), from interference between \(b\to c\) and \(b\to u\) amplitudes in \(B\to DK\). Both agree. The first of them is the sharper statement, because \(\sin2\beta\) is measured to three per cent in a tree-dominated decay whose theoretical uncertainty is below one per cent — and it agrees with a value obtained from \(\Delta m_{d}\), \(\Delta m_{s}\), \(\varepsilon_{K}\) and \(\abs{V_{ub}/V_{cb}}\), that is from three loop processes and one tree process, none of them a \(CP\) asymmetry.
What is being tested is the following. There are four parameters. The observables listed number well over a dozen. Any one of them could have landed outside the region the others define, and that would have falsified the hypothesis that a single phase in a unitary matrix accounts for all \(CP\) violation in the quark sector. None has. The global fits [Charles:2005] quantify the agreement, and the residual freedom in the apex is now a few per cent in each coordinate.
∎Three qualifications keep Phenomenon 103.31 honest. The angle \(\alpha\) carries a \(5^\circ\) uncertainty dominated by the hadronic “penguin” amplitudes that must be subtracted from \(B\to\pi\pi\), which is theory, not statistics. The side \(\abs{V_{ub}/V_{cb}}\) inherits the inclusive–exclusive tension of Section 103.4.2, and the two choices move the apex by about one standard deviation. And the first-row relation \(\Delta_{\mathrm{CKM}}\) of Equation (103.6), which is the most precisely measured unitarity statement of all, is the one that does not come out right. The correct summary is that the triangle closes at the few-per-cent level and the first row does not close at the \(10^{-3}\) level; both are measurements and both are recorded.
Neutral meson mixing and $CP$ violation
Neutral kaon mixing
A neutral meson carrying a flavour quantum number has an antiparticle distinct from itself, and the weak interaction, which does not conserve that quantum number, connects the two. The system is then a two-state quantum system with decay, and its behaviour is the most sensitive interferometer physics possesses.
Let \(P^{0}\) and \(\bar{P}^{0}\) be a neutral meson and its antiparticle, distinguished by a flavour quantum number conserved by the strong interaction. In the two-dimensional subspace they span, the effective time evolution is
with \(\mathsf{M}\) and \(\mathsf{\Gamma}\) Hermitian \(2\times2\) matrices, the first carrying \(\mathrm{kg}\) and the second \(/\mathrm{s}\). The eigenvalues of the non-Hermitian matrix in brackets define two states of definite mass and width, split by \(\Delta m\) and \(\Delta\Gamma\), and the two dimensionless ratios
decide whether mixing is observable: \(x\ll1\) and \(y\ll1\) means the meson decays before it can oscillate.
The derivation of Equation (103.33) from the Weisskopf–Wigner treatment of an unstable state, the constraints that \(CPT\) invariance and \(CP\) invariance impose on \(\mathsf{M}\) and \(\mathsf{\Gamma}\), and the explicit diagonalization are Proposition 105.22 and Proposition 105.23; they are not repeated here. What belongs here is the physics the four measured systems display.
| System | Quarks | $x$ | $y$ | Heaviest quark in the box |
|---|---|---|---|---|
| $K^{0}$–$\bar{K}^{0}$ | $d\bar{s}$ | 0.946 | 0.997 | charm (top suppressed by mixing) |
| $D^{0}$–$\bar{D}^{0}$ | $c\bar{u}$ | 0.0041 | 0.0065 | bottom |
| $B^{0}$–$\bar{B}^{0}$ | $d\bar{b}$ | 0.770 | 0.001 | top |
| $B_{s}$–$\bar{B}_{s}$ | $s\bar{b}$ | 27.0 | 0.062 | top |
Gell-Mann and Pais saw the essential point in 1955, before any of it was measured [GellMann:1955]. If \(CP\) were conserved, the states of definite mass and lifetime would be the \(CP\) eigenstates
with \(CP=+1\) and \(-1\) respectively in the phase convention of Equation (105.45). The two-pion final state has \(CP=+1\) and the three-pion state \(CP=-1\) (Equation (105.47)), and the three-pion channel has almost no phase space — \(m_{K}c^{2}-3m_{\pi}c^{2}=83\,\mathrm{MeV}\) against \(219\,\mathrm{MeV}\) for two pions. So \(K_{1}\) should decay fast and \(K_{2}\) slowly. The prediction was of a new particle, of the same mass and opposite \(CP\), and it was found: the measured lifetimes are
which is Equation (105.48).
A beam that is pure \(K^{0}\) at \(t=0\) contains, at proper time \(t\), a \(\bar{K}^{0}\) component of probability
so \(\Delta m_{K}\) is read off an interference pattern in time, not from a mass measurement. This is why \(\Delta m_{K}c^{2}=3.484(6)\times 10^{-6}\,\mathrm{eV}\) is known to three significant figures although it is a fraction \(7.0\times 10^{-15}\) of the kaon mass itself. Rests on Definition 103.33, Equation (103.35) and Proposition 101.56.
Derives Proposition 103.34. Write \(\ket{K^{0}}=\left(\ket{K_{L}}+\ket{K_{S}}\right)/\sqrt{2}\) and \(\ket{\bar{K}^{0}}=\left(\ket{K_{L}}-\ket{K_{S}}\right)/\sqrt{2}\), which inverts Equation (103.35) up to the \(CP\)-violating admixture of order \(2\times 10^{-3}\) that is neglected here and restored in Section 103.5.2. Each mass eigenstate evolves with its own phase and its own damping, \(\ket{K_{i}(t)}=\ee^{-\ii m_{i}c^{2}t/\hbar}\ee^{-\Gamma_{i}t/2} \ket{K_{i}}\), so
whose squared modulus is Equation (103.37). The strangeness of the beam at time \(t\) is measured directly, by the sign of the charged lepton in a semileptonic decay: the rule \(\Delta S=\Delta Q\) (Proposition 101.56) says a \(K^{0}\) gives \(\ell^{+}\) and a \(\bar{K}^{0}\) gives \(\ell^{-}\). Counting the two charges as a function of decay position along a beam line therefore traces Equation (103.37), and the cosine's period measures \(\Delta m_{K}\). With \(x=\Delta m_{K}c^{2}/\hbar\bar{\Gamma}=0.946\) from Table 103.2, roughly one oscillation occurs per lifetime, which is the optimum: much less and there is no interference, much more and it averages away.
∎A beam of \(K_{L}\), purified by letting the \(K_{S}\) component die away, recovers a \(K_{S}\) component on passing through matter. The mechanism is that the strong interaction sees \(K^{0}\) and \(\bar{K}^{0}\), not \(K_{L}\) and \(K_{S}\): the \(\bar{K}^{0}\) can convert a nucleon's quark content in ways the \(K^{0}\) cannot, so the forward scattering amplitudes \(f\) and \(\bar{f}\) differ. Writing \(\ket{K_{L}}\propto \ket{K^{0}}+\ket{\bar{K}^{0}}\), the transmitted beam is \(f\ket{K^{0}}+\bar{f}\ket{\bar{K}^{0}} \propto\left(f+\bar{f}\right)\ket{K_{L}} +\left(f-\bar{f}\right)\ket{K_{S}}\), so the regenerated amplitude is proportional to \(f-\bar{f}\) and vanishes only if the strong interaction is blind to strangeness, which it is not. The effect, Equation (105.64), is the single largest background to the measurement of Section 103.5.2 and had to be eliminated by evacuating the decay volume.
Discovery of $CP$ violation
The neutral kaon system contains a short-lived state decaying to two pions and a long-lived state, some six hundred times longer lived, decaying to three. In a beam allowed to travel far enough that the short-lived component has died away, a small fraction of the surviving decays are nevertheless to \(\pi^{+}\pi^{-}\): the branching fraction is about \(2\times 10^{-3}\) [Christenson:1964] [Navas:2024]. The experiment belongs to Experiment: CP Violation. Rests on Definition 103.33, Equation (103.35) and Equation (105.47).
Derivation. Derives Phenomenon 103.36. The pion has intrinsic parity \(-1\), so a state of two pions with relative orbital angular momentum \(L\) has parity \((-1)^{2}(-1)^{L}=(-1)^{L}\); the kaon has spin zero, so \(L=0\) and the parity is \(+1\). Charge conjugation exchanges \(\pi^{+}\) with \(\pi^{-}\), which for \(L=0\) is the identity, so the state also has \(C=+1\). The two-pion state is therefore an eigenstate of \(CP\) with eigenvalue \(+1\). The three-pion state available to the same system has, by the same counting together with the small energy release that suppresses non-zero orbital angular momenta, \(CP=-1\). If \(CP\) were a symmetry of the decay interaction the two final states would be reached from two orthogonal kaon states, and whichever state decays to three pions could never decay to two. That one and the same long-lived state does both is thus a direct proof that \(CP\) is not conserved [Christenson:1964]. What the argument does not settle is where the violation sits — in the mixing that defines the long-lived state, or in the decay amplitude itself. Separating them required the measurement of \(\epsilon'/\epsilon\) a generation later, and only that measurement excluded a superweak interaction acting solely on the mixing.
∎The separation is parametrized by two complex numbers. The \(CP\)-violating admixture in the mixing is \(\varepsilon\), defined by Equation (105.54); the \(CP\) violation in the decay amplitude itself is \(\varepsilon'\), defined through the two isospin amplitudes of the two-pion final state by Equation (105.72). The measured values are
the first from the ratio of \(K_{L}\) to \(K_{S}\) two-pion rates and the second from the double ratio of charged to neutral two-pion modes in the two beams — a quantity constructed precisely so that detector acceptance and beam normalization cancel (Proposition 105.24). The result was established independently by NA48 [Fanti:1999] and KTeV [AlaviHarati:1999] after two decades of measurements that disagreed with each other.
Wolfenstein's superweak hypothesis [Wolfenstein:1964] proposed a new \(\Delta S=2\) interaction acting only in the mass matrix, which would produce \(\varepsilon\neq0\) and \(\varepsilon'=0\) exactly. It fits the 1964 measurement perfectly, requires no third generation, and survived for thirty-five years. Equation (103.38) killed it: \(\varepsilon'/\varepsilon\) is small but non-zero at more than seven standard deviations, so \(CP\) violation occurs in the decay amplitude and not only in the mixing. This is the point at which the Kobayashi–Maskawa mechanism ceased to have a live competitor. The Standard Model prediction for \(\varepsilon'/\varepsilon\) is a difference of two large and partially cancelling hadronic contributions and remains uncertain at the level of tens of per cent; what the measurement establishes is the sign and existence of direct \(CP\) violation, not a precision test.
$B$ mixing and the $B$ factories
In 1987 the ARGUS collaboration at DESY, studying \(\Upsilon(4S)\) decays to \(B\bar{B}\) pairs, counted events in which both \(B\) mesons decayed semileptonically with leptons of the same sign [Albrecht:1987]. Same-sign dileptons require one of the two mesons to have changed into its own antiparticle. The measured ratio of same-sign to opposite-sign pairs was \(r=0.21\), and the consequence was immediate and unwelcome.
For a system with \(y\approx0\), the probability that a meson born as \(P^{0}\) decays as \(\bar{P}^{0}\), integrated over all decay times, is
ARGUS's \(r=0.21\) therefore gives \(x_{d}=0.73\). Rests on Equations (103.34) and (103.37).
Derives Proposition 103.38. From Equation (103.37) with \(\Gamma_{L}=\Gamma_{S}=\Gamma\), the mixed and unmixed probabilities are \(\tfrac{1}{2}\ee^{-\Gamma t} \left[1\mp\cos\left(\Delta m\,c^{2}t/\hbar\right)\right]\). Using \(\int_{0}^{\infty}\ee^{-\Gamma t}\dd t=\Gamma^{-1}\) and \(\int_{0}^{\infty}\ee^{-\Gamma t}\cos\omega t\,\dd t =\Gamma/\left(\Gamma^{2}+\omega^{2}\right)\) with \(\omega=\Delta m\,c^{2}/\hbar=x\Gamma\), and normalizing by the total \(\Gamma^{-1}\),
Solving \(r=\chi/(1-\chi)=0.21\) for \(x\) gives \(x^{2}=2r/(1-r)=0.532\), hence \(x=0.73\).
∎The box diagram generating \(\Delta m_{d}\) is the one of Example 103.12 with the external quarks \(d\) and \(b\), and by Theorem 103.11 it is dominated by the heaviest internal quark, which for a \(b\)-flavoured meson is the top. The amplitude is no longer in the regime \(m_{i}\ll M_{W}\), so the full loop function is needed:
with \(x_{t}=\left(m_{t}c^{2}\right)^{2} /\left(M_{W}c^{2}\right)^{2}\), \(\eta_{B}=0.55\) the short-distance QCD factor, and
which reduces to \(S_{0}(x)\to x\) for \(x\ll1\), recovering Equation (103.8). With \(m_{t}c^{2}=172.57\,\mathrm{GeV}\) and \(M_{W}c^{2}=80.369\,\mathrm{GeV}\) one has \(x_{t}=4.61\) and \(S_{0}=2.52\); inserting \(f_{B}\sqrt{B_{B}}\,c^{2}=210\,\mathrm{MeV}\), \(m_{B}c^{2}=5.2797\,\mathrm{GeV}\) and \(\abs{V_{td}}=8.57\times 10^{-3}\) gives
against the measured \(3.336\times 10^{-4}\,\mathrm{eV}\) — five per cent, which is the accuracy of the lattice input.
In the dimensionless variable of Equation (103.34), with \(\tau_{B}=1.519\times 10^{-12}\,\mathrm{s}\), Equation (103.41) corresponds to \(x_{d}=\Delta m_{d}c^{2}\tau_{B}/\hbar=0.81\), against the measured \(0.770\) of Table 103.2 — the same five per cent, now carried into the quantity ARGUS actually reported.
Now run it backwards to 1987. With \(m_{t}c^{2}=30\,\mathrm{GeV}\), the value then expected, \(S_{0}=0.129\) and Equation (103.41) falls by a factor \(19.5\), giving \(x_{d}=0.042\) against the \(0.73\) that ARGUS measured. The hadronic inputs were poorly known in 1987, but not by a factor of eighteen in a squared amplitude, and the conclusion drawn at once was that the top quark is far heavier than anyone had supposed.
Read with today's inputs, ARGUS's number becomes a determination of the top mass, and it is worth doing the inversion rather than asserting its result. Since \(\Delta m_{d}\propto S_{0}(x_{t})\) with everything else held fixed, ARGUS's \(x_{d}=0.73\) requires \(S_{0}=2.5226\times0.73/x_{d}^{\text{ref}}\), where \(x_{d}^{\text{ref}}\) is whatever the rest of the formula produces today. Anchoring on the formula itself, \(x_{d}^{\text{ref}}=0.81\), gives \(S_{0}=2.27\), hence \(x_{t}=4.03\) and \(m_{t}c^{2}=161\,\mathrm{GeV}\); anchoring instead on the measured \(x_{d}^{\text{ref}}=0.770\), which removes the five-per-cent lattice offset, gives \(S_{0}=2.39\), hence \(x_{t}=4.30\) and \(m_{t}c^{2}=167\,\mathrm{GeV}\). The two bracket the answer: ARGUS's mixing rate says \(m_{t}c^{2}\) is between about \(160\,\mathrm{GeV}\) and \(170\,\mathrm{GeV}\), against the \(172.57\,\mathrm{GeV}\) eventually measured, eight years before the top was produced — the second time in this chapter that a loop measured a quark that could not yet be made. The gap between the two anchors is the honest uncertainty of the inference, and it is far smaller than the uncertainty the hadronic inputs carried in 1987; the 1987 statement was correctly made as an inequality, and only today's lattice numbers turn it into a measurement.
The loop function \(S_{0}(x)\) and the \(B\)-meson box amplitude for the mass difference. Both are stated above and then used twice to weigh the top quark, so the derivation is load-bearing: the box integral for an internal mass comparable with the \(W\) mass, whose value is \(S_{0}\), is not carried out here and belongs in Appendix A.
Taking the ratio of Equation (103.40) for the \(B_{s}\) and \(B_{d}\) systems, the loop function, the QCD factor and the \(W\) mass cancel exactly, leaving
The lattice quantity \(\xi=1.208\) is a ratio of decay constants, in which most of the systematic uncertainty of the individual calculations cancels; this is what makes Equation (103.42) the best determination of \(\abs{V_{td}/V_{ts}}\). Rests on Equation (103.40).
Derives Proposition 103.40. Divide Equation (103.40) by its \(B_{s}\) counterpart; every factor independent of the light spectator quark cancels. Inverting with the measured oscillation frequencies \(\Delta m_{d}c^{2}/\hbar=5.069\times 10^{11}\,/\mathrm{s}\) and \(\Delta m_{s}c^{2}/\hbar=1.7765\times 10^{13}\,/\mathrm{s}\) [Navas:2024] [Aaij:2013], that is \(\Delta m_{d}c^{2}=3.336\times 10^{-4}\,\mathrm{eV}\) and \(\Delta m_{s}c^{2}=1.169\times 10^{-2}\,\mathrm{eV}\),
against \(8.6\times 10^{-3}/4.15\times 10^{-2}=0.207\) from the individual entries of Table 103.1. Note the units. The literature quotes these as angular frequencies \(\Delta mc^{2}/\hbar\) in \(/\mathrm{ps}\) — \(0.5069\,/\mathrm{ps}\) and \(17.765\,/\mathrm{ps}\) — while this book quotes the energy splitting \(\Delta mc^{2}\) in \(\mathrm{eV}\). The conversion is \(\hbar=6.5821196\times 10^{-16}\,\mathrm{eV}\,\mathrm{s}\), so \(1\,/\mathrm{ps}\) corresponds to \(6.582\times 10^{-4}\,\mathrm{eV}\), and the ratio the derivation uses is of course the same either way.
∎The asymmetric-energy \(B\) factories were built for one measurement. An \(\Upsilon(4S)\) produced at rest decays to a \(B\bar{B}\) pair almost at rest, and the two mesons travel some \(30\,\mu\mathrm{m}\) before decaying — too little to separate. Colliding beams of unequal energy gives the pair a Lorentz boost along the beam axis, stretching the separation to a few hundred micrometres, which a silicon vertex detector resolves. The time between the two decays is then measurable, and with it the time-dependent asymmetry of Equation (105.81).
In \(B^{0}\to J/\psi\,K_{S}\) the time-dependent decay-rate asymmetry between mesons tagged at production as \(B^{0}\) and as \(\bar{B}^{0}\) is an oscillation of amplitude
observed independently by BaBar [Aubert:2001] and Belle [Abe:2001] in 2001. The effect is of order unity, against \(2\times 10^{-3}\) in the kaon system, and it agrees with the value predicted from measurements containing no \(CP\) asymmetry at all. Rests on Proposition 105.28, Equation (103.40) and Equation (103.31).
Derivation. Derives Phenomenon 103.41. Proposition 105.28 gives the asymmetry as \(S_{f}\sin\left(\Delta m_{d}c^{2}t/\hbar\right) -C_{f}\cos\left(\Delta m_{d}c^{2}t/\hbar\right)\) with \(S_{f}=2\,\mathrm{Im}\,\lambda/\left(1+\abs{\lambda}^{2}\right)\) and \(\lambda=\left(q/p\right) \left(\bar{\mathcal{A}}_{f}/\mathcal{A}_{f}\right)\). For \(f=J/\psi K_{S}\) three things conspire to make this clean. The decay \(b\to c\bar{c}s\) is dominated by a single tree amplitude, so \(\abs{\bar{\mathcal{A}}/\mathcal{A}}=1\) to within a per cent and \(C_{f}\approx0\); the mixing phase \(q/p\) is, from the box of Equation (103.40), twice the phase of \(V_{tb}^{*}V_{td}\); and the \(K^{0}\)–\(\bar{K}^{0}\) mixing in the final state contributes twice the phase of \(V_{cs}^{*}V_{cd}\). Collecting,
which is exactly the angle of Equation (103.31); hence \(S=\sin2\beta\) with no hadronic quantity anywhere in the relation. That is the whole reason the measurement is decisive: the hadronic matrix elements cancel between the two interfering paths because both end in the same final state.
Why is the effect a thousand times larger than in the kaon? Because \(CP\) violation requires interference between two amplitudes of comparable size and different weak phase. In the kaon, the \(CP\)-violating amplitude is a small admixture in a state dominated by \(CP\)-conserving physics, so the asymmetry carries the small factor \(\varepsilon\). In \(B^{0}\to J/\psi K_{S}\) the two interfering paths — decay without mixing, and decay after mixing — have equal magnitude by construction, since \(\abs{q/p}=1\) and \(\abs{\lambda}=1\), so the asymmetry is the full \(\sin\) of the phase difference. Nothing is small except the phase, and the phase is not small.
∎Mixing in the \(B_{s}\) system, with \(x_{s}=27\), oscillates twenty-seven times per lifetime and required a detector with picosecond decay-time resolution; it was resolved at the Tevatron and measured precisely at LHCb, which also observed \(CP\) violation in \(B_{s}\) decays [Aaij:2013]. The value of \(\Delta m_{s}\) enters Equation (103.42) and is one of the constraints on the apex in Phenomenon 103.31.
Charm and the present picture
The charm system was the last to yield. By Theorem 103.11, a \(D^{0}\) box runs over \(d\), \(s\) and \(b\) quarks, all light compared with \(M_{W}\), so the GIM suppression is severe and \(x\) and \(y\) are of order \(10^{-3}\) (Table 103.2); the \(CP\)-violating part is suppressed further by the small mixing factor \(\abs{V_{cb}V_{ub}}\sim1.5\times 10^{-4}\).
The difference of \(CP\) asymmetries between \(D^{0}\to K^{+}K^{-}\) and \(D^{0}\to\pi^{+}\pi^{-}\) is
non-zero at \(5.3\) standard deviations [Aaij:2019]. With it, \(CP\) violation is established in all three quark systems in which it can occur. Rests on Theorem 103.11 and Definition 103.25.
Derivation. Derives Phenomenon 103.42. Taking the difference of two asymmetries rather than either alone is what makes the measurement possible: production asymmetries between \(D^{0}\) and \(\bar{D}^{0}\) in \(pp\) collisions, and detection asymmetries between positive and negative tracks, are common to the two final states and cancel in the difference, while the \(CP\)-odd parts do not, since \(U\)-spin relates the two channels with opposite sign. The residual is then a genuine asymmetry between a decay and its conjugate. Its magnitude, \(1.5\times 10^{-3}\), is of the order the Kobayashi–Maskawa phase predicts once the small mixing factor and the interference with the dominant Cabibbo-favoured amplitude are accounted for — but the hadronic matrix elements of charm decays are not calculable to better than a factor of a few, so the correct statement is that the observation is consistent with the Standard Model and does not test it sharply.
∎It is worth stating plainly what Section 103.5 has established. Every \(CP\)-violating observable measured in the quark sector — \(\varepsilon\) and \(\varepsilon'\) in the kaon, \(\sin2\beta\), \(\gamma\) and \(\alpha\) in the \(B\) mesons, the \(B_{s}\) asymmetries, and Equation (103.44) in charm — is described by the single phase of Definition 103.25, whose rephasing-invariant measure is \(J=3.08\times 10^{-5}\). Four parameters, well over a dozen independent measurements, one consistent solution. The places where the description is least sharply tested are those where hadronic matrix elements dominate the uncertainty: \(\varepsilon'/\varepsilon\), the charm asymmetry just quoted, and the angle \(\alpha\). The places where it is sharply tested — \(\sin2\beta\), \(\gamma\), the ratio \(\Delta m_{s}/\Delta m_{d}\) — are those where the hadronic physics cancels.
What $CP$ violation cannot yet explain
No astrophysical concentration of antimatter is observed, and the baryon density inferred from primordial nucleosynthesis agrees with the one inferred from the acoustic peaks of the microwave background on a baryon-to-photon ratio
[Aghanim:2020]. Generating such an asymmetry from a symmetric initial state requires baryon-number violation, violation of \(C\) and of \(CP\), and a departure from thermal equilibrium [Sakharov:1967]. The \(CP\) violation established in Phenomenon 103.36 is genuine but quantitatively far too small — by eleven powers of ten — to produce \(\eta\). Rests on Phenomenon 103.36, Theorem 105.41 and Proposition 105.42.
Derivation. Derives Phenomenon 103.44. The three Sakharov conditions are derived as Theorem 105.41 and the value of \(\eta\) from \(\Omega_{b}h^{2}\) as Proposition 105.42; neither is repeated. What is needed here is the size of the effect the CKM phase can produce, and the argument is dimensional.
By Proposition 105.27, all \(CP\) violation in the quark sector is proportional to the invariant \(\det\comm{M_{u}M_{u}^{\dagger}c^{4}}{M_{d}M_{d}^{\dagger}c^{4}}\), which carries \(J\) multiplied by the six squared-mass differences of Equation (105.79) and has SI dimension \(\mathrm{J}^{12}\). Any asymmetry generated at a temperature \(T\) is a pure number, so it can depend on that invariant only through the dimensionless ratio Equation (105.106), and evaluating it with the running quark masses at \(k_{B}T\approx100\,\mathrm{GeV}\), where the baryon-number-violating processes are active, gives Equation (105.107),
to be compared with \(\eta=6.1\times 10^{-10}\): the ratio \(\eta/d_{CP}=3\times 10^{11}\), a shortfall of eleven orders of magnitude, and detailed calculations [Gavela:1994] do not improve it. The reason is structural. \(CP\) violation in this theory requires all six quark masses to be distinct, and at \(100\,\mathrm{GeV}\) five of the six are negligible against the temperature, so the effect is suppressed by twelve powers of small mass ratios. The phase is of order unity; the masses kill it.
The third Sakharov condition fails as well, and independently. A departure from equilibrium requires the electroweak transition to be first order, and for the measured Higgs mass of Equation (104.66) it is not: it is a smooth crossover, as Remark 104.65 records. Two of the three conditions are therefore not met by the Standard Model with its measured parameters. The origin of Equation (103.45) is unknown, and this treatise says so; it is carried to Section 133.5.
∎Mechanisms that would generate the observed asymmetry exist in outline — heavy Majorana neutrinos decaying out of equilibrium, a first-order transition driven by additional scalars — and none has any experimental support. They are named in Section 133.5 and in Section 103.7.4 as hypotheses without evidence, which is what they are. What is measured is Equation (103.45), and what is calculated is Equation (103.46); the eleven orders of magnitude between them are one of the few places where the Standard Model is known, from laboratory numbers plus one cosmological observation, to be incomplete.
Lepton flavour and neutrino mixing
The PMNS matrix
Pontecorvo asked in 1957 whether the neutral-kaon phenomenon of Section 103.5.1 has a leptonic analogue [Pontecorvo:1957] [Pontecorvo:1958]. His original proposal was \(\nu\leftrightarrow\bar{\nu}\) oscillation, which is not what happens; but the idea that a neutral particle produced in a weak interaction need not be a state of definite mass, and can therefore convert into another, is exactly what happens. The correct form was written by Maki, Nakagawa and Sakata in 1962 [Maki:1962], the year the second neutrino flavour was established [Danby:1962]: flavour states are superpositions of mass states, and the superposition matrix plays for leptons the role \(V_{\mathrm{CKM}}\) plays for quarks.
In the basis in which the charged-lepton mass matrix is diagonal, the neutrino flavour states are
with \(U\) unitary and dimensionless. Written in the parametrization of Equation (103.21) with angles \(\theta_{12},\theta_{23},\theta_{13}\) and Dirac phase \(\delta_{CP}\), it carries in addition a diagonal factor
present if and only if the neutrinos are Majorana particles.
The two phases \(\alpha_{21}\), \(\alpha_{31}\) of Equation (103.48) cancel from every oscillation probability, and from every process in which lepton number is conserved. No oscillation experiment, however precise, can determine whether the neutrino is a Dirac or a Majorana particle. Rests on Definition 103.46 and Equation (103.48).
Derives Proposition 103.47. Every oscillation amplitude has the structure \(\sum_{i}U_{\beta i}^{\phantom{*}}\ee^{-\ii\varphi_{i}} U_{\alpha i}^{*}\), in which each index \(i\) appears once on \(U\) and once on \(U^{*}\). Under Equation (103.48), \(U_{\alpha i}\mapsto U_{\alpha i}\ee^{\ii\alpha_{i1}/2}\) for every \(\alpha\), so the product \(U_{\beta i}U_{\alpha i}^{*}\) acquires \(\ee^{\ii\alpha_{i1}/2}\ee^{-\ii\alpha_{i1}/2}=1\). The same cancellation occurs in any amplitude with one neutrino line entering and one leaving, that is, in any lepton-number-conserving process. It fails only where two neutrino lines are contracted with each other, which requires a Majorana propagator and violates lepton number by two units — the process of Section 103.7.2, and the reason that is the only practical test.
The counting is worth recording. For \(N\) Dirac neutrinos the phase count is \((N-1)(N-2)/2\), as for quarks, because the neutrino fields may be rephased. For \(N\) Majorana neutrinos a rephasing \(\nu_{i}\to\ee^{\ii\phi}\nu_{i}\) is forbidden — it would multiply the Majorana mass term \(\overline{\nu^{c}}\nu\) by \(\ee^{2\ii\phi}\) — so only the charged-lepton fields may be rephased, \(N-1\) of those phases being usable, and the surviving count is \(N(N-1)/2\): three phases for \(N=3\), one Dirac and two Majorana.
∎The magnitudes of the PMNS entries are, from the global fit to all oscillation data [Navas:2024] [Esteban:2020],
against the near-diagonal \(\abs{V_{\mathrm{CKM}}}\) of Table 103.1. Two of the three lepton mixing angles are of order tens of degrees; all three quark mixing angles are small. Rests on Definition 103.46, Equation (103.21) and Equation (103.22).
Derivation. Derives Phenomenon 103.48. Equation (103.49) follows by inserting the measured \(\sin^{2}\theta_{12}=0.307\), \(\sin^{2}\theta_{23}=0.546\) and \(\sin^{2}\theta_{13}=0.0220\) of Table 103.3 into Equation (103.21), with \(\delta_{CP}\) set to zero, which shifts the entries of the lower two rows by at most a few units in the second decimal. The corresponding angles are
to be compared with Equation (103.22), where the largest is \(13.0^\circ\) and the smallest \(0.21^\circ\). The comparison is not between two numbers but between two patterns: the quark matrix is a small perturbation of the identity, with a hierarchy \(1:\lambda:\lambda^{3}\); the lepton matrix has no hierarchy at all except that one of its nine entries is small.
∎The Standard Model requires both matrices to be unitary — the argument is Remark 5.152 in both cases — and fixes nothing else about either. Their entries are Yukawa couplings in disguise, by Theorem 104.38, and the Yukawa couplings are free parameters. That one matrix comes out nearly diagonal and the other nearly democratic is therefore a fact with no theoretical content in the Standard Model whatsoever. Attempts to relate them exist and none has evidence; the fact is recorded as open in Section 133.1. It should not be dressed up as a hint.
Oscillation in vacuum
A neutrino created with one flavour is detected, some distance later, with another. Three independent classes of observation establish it. The flux of atmospheric \(\nu_{\mu}\) depends on zenith angle, that is on the distance travelled, in a way no production model imitates [Fukuda:1998]. The solar \(\nu_{e}\) flux measured by charged-current reactions is about a third of the total active flux measured by the neutral-current reaction on deuterium, and the total agrees with the solar model [Ahmad:2002] — which resolves the deficit Davis reported thirty years earlier [Davis:1968] without appeal to solar physics at all. Reactor \(\bar{\nu}_{e}\) disappear and reappear as a function of \(L/E\) [Eguchi:2003]. Flavour is therefore not conserved in propagation, and the neutrino masses are not all zero. Lepton mixing is moreover large, two of its three angles being of order tens of degrees [Navas:2024], in flat contrast with the small angles of Section 103.4. The experiments are Experiment: Neutrino Oscillations. Rests on Definition 103.46 and Phenomenon 103.48.
Derivation. Derives Phenomenon 103.50. Take two flavour states related to two states of definite mass \(m_{1}\neq m_{2}\) by a single angle,
The mass states, not the flavour states, are the ones that propagate with a definite phase. For a neutrino of energy \(E\) far above its rest energy, \(E_{i}=\sqrt{p^{2}c^{2}+m_{i}^{2}c^{4}}\approx E+m_{i}^{2}c^{4}/2E\), so after travelling a distance \(L\) in a time \(L/c\) the component \(\ket{\nu_{i}}\) has acquired the phase
The first term is common to both components and unobservable. The amplitude to find the other flavour is the projection of the propagated state on \(\ket{\nu_{\beta}}\),
and with \(\abs{\ee^{-\ii\varphi_{2}}-\ee^{-\ii\varphi_{1}}}^{2} =2\left[1-\cos(\varphi_{1}-\varphi_{2})\right] =4\sin^{2}\tfrac{1}{2}(\varphi_{1}-\varphi_{2})\) its squared modulus is
Three consequences follow, and all three are what is seen. The effect is an interference between two mass components, so it measures a difference of squared masses and never a mass: the oscillation experiments cannot say how heavy a neutrino is, only that two of them differ. It vanishes if the masses are equal or if \(\theta=0\), so its observation proves both that the mass states are non-degenerate and that the flavour and mass bases are misaligned. And it depends on \(L\) and \(E\) only through their ratio, which is why the atmospheric, reactor and accelerator data — taken at baselines from tens of metres to thousands of kilometres — collapse onto one curve when plotted against \(L/E\).
∎Equation (103.52) was obtained by giving the two mass components the same momentum and letting them travel the same distance in the same time. Each of those is an approximation, and each has generated a literature. A neutrino is produced localized — in a nucleus, in a pion decay — so its state is a wave packet, and the components of different mass travel at slightly different group velocities and eventually cease to overlap. The correct statement is that Equation (103.52) is the limit of a wave-packet calculation under two conditions, both of which hold by enormous margins in every experiment discussed here.
Localization. The production and detection regions must be small compared with the oscillation length Equation (103.55): otherwise the phase varies over the source and the pattern washes out. What sets \(\sigma_{x}\) for a reactor antineutrino is not the size of the decaying nucleus, which is of order \(10^{-15}\,\mathrm{m}\), but the extent over which the decaying system is coherently localized — a fission fragment bound in the fuel lattice, whose position spread is of atomic order, \(\sigma_{x}\sim10^{-11}\,\mathrm{m}\). Against an oscillation length of \(99\,\mathrm{km}\) either number is negligible, so the localization condition is never in question. The coherence condition is the one that has to be checked, and there the distinction between the two matters by eight orders of magnitude.
Coherence. The packets must still overlap on arrival, which requires
the separation of the packets growing as \(L\,\Delta m^{2}c^{4}/2E^{2}\) because that is the difference of their group velocities. For the same reactor case, \(\sigma_{x}=10^{-11}\,\mathrm{m}\), \(E=3\,\mathrm{MeV}\) and \(\Delta m^{2}c^{4}=7.53\times 10^{-5}\,\mathrm{eV}^{2}\) give \(L_{\mathrm{coh}}=6.8\times 10^{3}\,\mathrm{km}\), against KamLAND's \(180\,\mathrm{km}\). The margin is a factor of \(38\), and it is the atomic \(\sigma_{x}\) that supplies it: inserting the nuclear \(10^{-15}\,\mathrm{m}\) instead would give \(L_{\mathrm{coh}}=0.68\,\mathrm{km}\), so that coherence would be lost some \(260\) times over before the detector and no oscillatory dependence on \(L/E\) could be seen at all. KamLAND sees one, so the relevant localization is the atomic one — which is a measurement of \(\sigma_{x}\), not an assumption about it. Note that Equation (103.53) is not the same statement as the averaging of Equation (103.52) over a spread in \(L\) or \(E\): that is a classical loss of resolution, this is a quantum loss of coherence, and in every existing experiment the first sets in long before the second.
The wave-packet coherence length for neutrino oscillation, including its factor \(4\sqrt{2}\). What is given inline is the group-velocity separation, which fixes the scaling with the packet width, the energy and the squared-mass splitting; the Gaussian overlap integral that fixes the numerical coefficient, and with it the precise meaning of the packet width, is not carried out here and belongs in Appendix A.
The chapter's SI discipline has a payoff here, because the number every experimental paper uses is a disguised statement about \(\hbar\) and \(c\) and is almost never derived.
With the squared-mass splitting expressed as an energy squared in \(\mathrm{eV}^{2}\), the baseline in kilometres and the energy in \(\mathrm{GeV}\),
and the oscillation length, the distance over which the phase advances by \(\pi\), is
Derives Proposition 103.52. Multiply numerator and denominator of the phase in Equation (103.52) by \(c\): the numerator becomes \(\Delta m^{2}c^{4}\), an energy squared, and the denominator \(4\hbar cE\), also an energy squared multiplied by a length. Hence
which is manifestly dimensionless when \(L\) is a length. Now insert \(\hbar c=1.9732698\times 10^{-7}\,\mathrm{eV}\,\mathrm{m}\) from Equation (103.1), and convert the units of the three inputs: \(L=\left(L/\mathrm{km}\right)\times10^{3}\,\mathrm{m}\) and \(E=\left(E/\mathrm{GeV}\right)\times10^{9}\,\mathrm{eV}\). Then
and the prefactor is the pure number
This is the constant that appears as “\(1.27\)” in every experimental paper on the subject, where it is quoted rather than derived; it is \(\left(4\hbar c\right)^{-1}\) in disguise, and nothing else. Setting \(\varphi=\pi\) and solving for \(L\) gives \(\pi/1.26693=2.4797\), which is Equation (103.55).
∎Equation (103.55) explains the design of every experiment treated in Experiment: Neutrino Oscillations. Only two of the three splittings are independent, since \(\Delta m^{2}_{31}=\Delta m^{2}_{32}+\Delta m^{2}_{21}\) identically; this book tabulates \(\Delta m^{2}_{21}\) and \(\Delta m^{2}_{32}\) in Table 103.3 and writes \(\Delta m^{2}_{31}\) for the sum, which exceeds \(\abs{\Delta m^{2}_{32}}\) by \(3\%\) in the normal ordering — immaterial wherever a two-figure statement is being made, which is everywhere below except Equation (103.75). For the atmospheric splitting \(\abs{\Delta m^{2}_{32}}c^{4}=2.455\times 10^{-3}\,\mathrm{eV}^{2}\) and \(E=1\,\mathrm{GeV}\), \(L_{\mathrm{osc}}=1010\,\mathrm{km}\): hence accelerator baselines of \(295\,\mathrm{km}\) (T2K) and \(810\,\mathrm{km}\) (NOvA), and hence the atmospheric signal appearing as a deficit of upward-going and not downward-going muon neutrinos, since the Earth's diameter is \(12742\,\mathrm{km}\) and its atmosphere \(15\,\mathrm{km}\) thick. For the solar splitting \(7.53\times 10^{-5}\,\mathrm{eV}^{2}\) at a reactor energy of \(3\,\mathrm{MeV}\), \(L_{\mathrm{osc}}=99\,\mathrm{km}\): hence KamLAND at \(180\,\mathrm{km}\). For the atmospheric splitting at the same reactor energy, \(L_{\mathrm{osc}}=3.0\,\mathrm{km}\): hence Daya Bay's far detectors at \(1.6\,\mathrm{km}\), which is where the \(\theta_{13}\) oscillation is near its first maximum while the solar term is still negligible. The separation of the two splittings by a factor \(33\) is what allows each to be measured almost independently of the other.
For three flavours mixed by Equation (103.47),
with \(\Delta m^{2}_{ij}=m_{i}^{2}-m_{j}^{2}\). The second sum is odd under \(\alpha\leftrightarrow\beta\) and changes sign between neutrinos and antineutrinos; it is the leptonic \(CP\)-violating term. Rests on Equations (103.47) and (103.51).
Derives Theorem 103.54. The amplitude is \(A_{\alpha\to\beta}=\sum_{i}U_{\beta i}^{\phantom{*}} \ee^{-\ii m_{i}^{2}c^{3}L/2\hbar E}U_{\alpha i}^{*}\), the common phase of Equation (103.51) having been dropped. Then
Split the double sum into \(i=j\), which gives \(\sum_{i}\abs{U_{\alpha i}}^{2}\abs{U_{\beta i}}^{2}\), and the off-diagonal terms, which pair \((i,j)\) with \((j,i)\). Writing \(W_{ij}=U_{\alpha i}^{*}U_{\beta i}U_{\alpha j}U_{\beta j}^{*}\), the coefficient standing in front of \(\ee^{-\ii\varphi_{ij}}\) in the display above is \(U_{\beta i}U_{\alpha i}^{*}U_{\beta j}^{*}U_{\alpha j} =\left(U_{\alpha i}^{*}U_{\beta i}\right) \left(U_{\alpha j}U_{\beta j}^{*}\right)=W_{ij}\) itself, and the \((j,i)\) term is its complex conjugate because \(W_{ji}=W_{ij}^{*}\) and \(\varphi_{ji}=-\varphi_{ij}\). The paired contribution is therefore \(2\,\mathrm{Re}\left(W_{ij}\ee^{-\ii\varphi_{ij}}\right) =2\,\mathrm{Re}\,W_{ij}\cos\varphi_{ij} +2\,\mathrm{Im}\,W_{ij}\sin\varphi_{ij}\) with \(\varphi_{ij}=\Delta m^{2}_{ij}c^{3}L/2\hbar E\) — the sign of the second term being the sign of the \(CP\)-violating term in Equation (103.56), so it is worth stating why it is a plus: \(\ee^{-\ii\varphi}\) contributes \(-\ii\sin\varphi\), and \(\mathrm{Re}(-\ii W)=+\,\mathrm{Im}\,W\). Using \(\cos\varphi=1-2\sin^{2}(\varphi/2)\) and the unitarity relation \(\sum_{i}U_{\alpha i}^{*}U_{\beta i}=\delta_{\alpha\beta}\), whose squared modulus supplies \(\sum_{i}\abs{U_{\alpha i}}^{2}\abs{U_{\beta i}}^{2} +2\sum_{i>j}\mathrm{Re}\,W_{ij}=\delta_{\alpha\beta}\), gives Equation (103.56). Replacing \(U\) by \(U^{*}\), which is what \(CP\) conjugation does, flips the sign of the imaginary part and leaves the real part alone, which is the last statement.
∎Because \(\abs{\Delta m^{2}_{31}}/\Delta m^{2}_{21}=34\) and \(\abs{U_{e3}}^{2}=0.022\) is small, Equation (103.56) reduces to a two-flavour formula in each of two regimes. At \(L/E\) tuned to the atmospheric splitting the solar phase is negligible and \(P\left(\nu_{\mu}\to\nu_{\mu}\right)\approx 1-\sin^{2}2\theta_{23}\sin^{2} \left(\Delta m^{2}_{31}c^{4}L/4\hbar cE\right)\) with \(\sin^{2}2\theta_{23}=0.99\); at \(L/E\) tuned to the solar splitting the atmospheric phase averages to one half and \(P\left(\bar{\nu}_{e}\to\bar{\nu}_{e}\right)\approx \cos^{4}\theta_{13}\left[1-\sin^{2}2\theta_{12}\sin^{2} \left(\Delta m^{2}_{21}c^{4}L/4\hbar cE\right)\right] +\sin^{4}\theta_{13}\). This factorization, and not any approximation made for convenience, is why the parameters could be measured one at a time over thirty years. Rests on Equation (103.56).
Matter effects
A neutrino crossing ordinary matter forward-scatters coherently off the electrons, protons and neutrons in it. Coherent forward scattering does not change the neutrino's momentum and produces no observable interaction; what it produces is a refractive index, exactly as for light in glass, and hence a shift in the propagation phase. The shift would be irrelevant if it were the same for all three flavours. It is not: matter contains electrons and contains no muons or taus, so \(\nu_{e}\) alone can forward-scatter through the charged current, and its potential differs from the others'.
In a medium of electron number density \(n_{e}\), an electron neutrino acquires the extra potential energy
of SI dimension \(\mathrm{J}\), while \(\nu_{\mu}\) and \(\nu_{\tau}\) acquire none from the charged current. An antineutrino acquires \(-V_{e}\). Numerically, at the centre of the Sun where \(n_{e}=6\times 10^{31}\,/\mathrm{m}^{3}\),
Rests on Equation (101.6), Definition 100.2 and Equation (101.8).
Derives Proposition 103.56. The relevant term of the effective Lagrangian of Equation (101.6), Fierz-rearranged so that the neutrino fields are paired with each other, is
of dimension \(\mathrm{J}/\mathrm{m}^{3}\), since each bilinear carries \(/\mathrm{m}^{3}\) by Definition 100.2 and \(G_{F}\) carries \(\mathrm{J}\,\mathrm{m}^{3}\). Average the electron bilinear over the medium. For an unpolarized, isotropic electron gas at rest, the spatial components of the vector current average to zero and the axial current averages to the net spin density, also zero; the surviving component is \(\avg{\bar{\psi}_{e}\gamma^{0}\psi_{e}}=n_{e}\), the electron number density. The neutrino bilinear with \(\gamma^{0}\) counts the neutrino number density, so the average Lagrangian is a potential energy per neutrino of \(V_{e}=\sqrt{2}G_{F}n_{e}\): one factor of \(2\) from the \(\left(\identity-\gamma^{5}\right)\) acting on a left-chiral neutrino, against the \(1/\sqrt{2}\) of the Fermi normalization. The neutral current contributes as well, but identically for all three flavours, so it adds a multiple of the identity to the Hamiltonian and cannot affect oscillation — the reason only the charged-current term appears in Equation (103.57). For an antineutrino, \(C\) conjugation reverses the sign of the vector current, giving \(-V_{e}\); this sign is the whole of the mass-ordering sensitivity discussed below. The numerical evaluation is Equation (103.58), using \(G_{F}\) in SI from Equation (101.8) — not \(G_{F}/(\hbar c)^{3}\), the confusion Definition 101.8 exists to prevent.
∎For two flavours with vacuum mixing angle \(\theta\) and splitting \(\Delta m^{2}\), propagation through matter of electron density \(n_{e}\) is governed by an effective mixing angle \(\theta_{m}\) and an effective splitting \(\Delta m^{2}_{m}\) given by
The mixing is maximal, \(\sin^{2}2\theta_{m}=1\), at the resonance
however small the vacuum angle \(\theta\), provided \(\cos2\theta>0\); the resonance occurs for neutrinos if \(\Delta m^{2}>0\) and for antineutrinos if \(\Delta m^{2}<0\) [Wolfenstein:1978] [Mikheyev:1985]. Rests on Equation (103.50) and Proposition 103.56.
Derives Theorem 103.57. The evolution of the flavour amplitudes with time is \(\ii\hbar\,\pp_{t}\Psi=\Ham\Psi\), and for an ultrarelativistic neutrino of fixed momentum the free part of \(\Ham\) is \(\sqrt{p^{2}c^{2}+m_{i}^{2}c^{4}}\approx pc+m_{i}^{2}c^{4}/2E\) in the mass basis. Dropping the term proportional to the identity and rotating to the flavour basis with Equation (103.50),
the matter term being diagonal in the flavour basis by Proposition 103.56. Subtracting \(\tfrac{1}{2}V_{e}\identity\), which shifts both levels equally and changes no probability,
A real symmetric traceless \(2\times2\) matrix of this form is diagonalized by a rotation through \(\theta_{m}\) with \(\tan2\theta_{m}=B/A\), which after multiplying numerator and denominator by \(2E\) is Equation (103.59); its eigenvalue difference is \(\sqrt{A^{2}+B^{2}}\), which multiplied by \(2E\) is Equation (103.60). The angle is \(\pi/4\) exactly when \(A=0\), which is Equation (103.61); and since \(V_{e}>0\) for matter, \(A\) can vanish only if \(\Delta m^{2}\cos2\theta>0\). For an antineutrino \(V_{e}\to-V_{e}\) and the condition becomes \(\Delta m^{2}\cos2\theta<0\). The resonance therefore occurs for neutrinos or for antineutrinos according to the sign of the splitting, and that is the only place in this chapter where a sign, rather than a magnitude, is observable.
∎Insert the solar values \(\Delta m^{2}_{21}c^{4}=7.53\times 10^{-5}\,\mathrm{eV}^{2}\), \(\sin^{2}\theta_{12}=0.307\) so that \(\cos2\theta_{12}=1-2\sin^{2}\theta_{12}=0.386\), and Equation (103.58), into Equation (103.61):
That number organizes the whole solar neutrino programme. Neutrinos produced well below \(1.9\,\mathrm{MeV}\) — the \(pp\) neutrinos, with endpoint \(420\,\mathrm{keV}\), which dominate the flux — never reach resonance and oscillate essentially as in vacuum, their survival probability averaging to
Neutrinos produced well above it — the \(^{8}\mathrm{B}\) neutrinos, extending to \(15\,\mathrm{MeV}\), which are the ones Homestake and SNO detected — are created above resonance, where the matter term dominates and \(\theta_{m}\to\pi/2\), so the neutrino leaves the Sun as the heavier mass eigenstate \(\nu_{2}\) and its survival probability is
Both numbers are measured, the second by SNO (Phenomenon 114.5) and the first by Borexino, which detected the \(pp\) neutrinos in real time [Bellini:2014] and later the CNO neutrinos [Agostini:2020a]. A single curve, rising from \(0.31\) to \(0.57\) as the energy falls through a few \(\mathrm{MeV}\), passes through every measured point. Both expressions are the two-flavour limit; restoring the third angle multiplies each by \(\cos^{4}\theta_{13}=0.957\) and adds \(\sin^{4}\theta_{13}=4.8\times 10^{-4}\), moving them to \(0.55\) and \(0.29\), which is the four-per-cent accuracy at which the two-flavour statement should be read. It is worth being clear about what this establishes that vacuum oscillation could not: the sign of \(\Delta m^{2}_{21}\), and hence that \(\nu_{2}\) is heavier than \(\nu_{1}\) — equivalently, that \(\theta_{12}\) lies in the first octant. Vacuum oscillation is even in \(\Delta m^{2}\) and blind to both.
The neutrino follows the instantaneous eigenstate through the resonance, rather than jumping between eigenstates, provided
and for the Sun, with \(L_{\rho}=6.6\times 10^{7}\,\mathrm{m}\) and the values of Example 103.58, \(\gamma=1.5\times 10^{4}\). Rests on Theorem 103.57 and Equation (103.60).
Derives Proposition 103.59. The adiabatic theorem states that a system in a non-degenerate instantaneous eigenstate of a slowly varying Hamiltonian stays in it, with a transition probability that vanishes as the variation is slowed; the general statement and its proof belong to Section 82.6.1, which reserves them. Its condition is that the rate of change of the Hamiltonian's eigenbasis be small compared with the level splitting divided by \(\hbar\). Here the eigenbasis is fixed by \(\theta_{m}\), and the fastest variation is at resonance. Differentiating \(\tan2\theta_{m}=B/A\) from the proof of Theorem 103.57, in which only \(A\) depends on \(r\),
the last step using \(V_{e}=\Delta m^{2}c^{4}\cos2\theta/2E\) at resonance; hence \(\dd\theta_{m}/\dd r=\cot2\theta/\left(2L_{\rho}\right)\) in magnitude. The two eigenvalues of the traceless Hamiltonian there are \(\pm\Delta m^{2}_{m}c^{4}/4E\) with \(\Delta m^{2}_{m}c^{4}=\Delta m^{2}c^{4}\sin2\theta\) by Equation (103.60), so the ratio of the phase rate to the rotation rate is
Numerically, with \(\Delta m^{2}c^{4}\sin^{2}2\theta =7.53\times 10^{-5}\,\mathrm{eV}^{2}\times0.851 =6.41\times 10^{-5}\,\mathrm{eV}^{2}\), \(2E\cos2\theta=2\times1.91\times 10^{6}\,\mathrm{eV}\times0.386 =1.47\times 10^{6}\,\mathrm{eV}\) and \(L_{\rho}/\hbar c=6.6\times 10^{7}\,\mathrm{m}/1.9733\times 10^{-7}\,\mathrm{eV}\,\mathrm{m} =3.34\times 10^{14}\,/\mathrm{eV}\), the product is \(1.5\times 10^{4}\). The solar density profile is exponential over the relevant region, which is what makes \(L_{\rho}\) a constant, and the margin is four orders of magnitude: the conversion is adiabatic and the survival probabilities of Example 103.58 are exact in that limit.
∎Two consequences of Theorem 103.57 are measured or being measured, and both belong to Section 114.5.2, which reserves them.
Day–night asymmetry. Solar neutrinos arriving at night cross the Earth, whose electron density partially regenerates the \(\nu_{e}\) component, so the night-time rate should exceed the day-time rate by a few per cent. The effect is at the edge of the sensitivity of existing detectors and is observed at about the \(2\sigma\) level; it is a prediction of the mechanism with no free parameter beyond the already-measured ones.
The mass ordering. By the sign statement of Theorem 103.57, matter enhances \(\nu_{\mu}\to\nu_{e}\) and suppresses \(\bar{\nu}_{\mu}\to\bar{\nu}_{e}\) if \(\Delta m^{2}_{31}>0\), and does the opposite if it is negative. A long-baseline beam run alternately in neutrino and antineutrino mode therefore determines the ordering — but the same comparison is also where \(\delta_{CP}\) shows up, by Theorem 103.54, and the two effects are degenerate at a single baseline. Breaking the degeneracy requires either two baselines or a baseline long enough (\(\gtrsim1000\,\mathrm{km}\)) that the matter effect dominates. That is the design brief of the next generation of experiments, and it is why the ordering is not yet known.
The evidence
The experimental case belongs to Experiment: Neutrino Oscillations and is summarized here in the order in which it was made, because the order is itself instructive: a thirty-year anomaly, then two measurements that admitted no escape.
The solar deficit, 1968–1998. Davis's chlorine detector, \(615\,\mathrm{t}\) of perchloroethylene \(1478\,\mathrm{m}\) underground, counted about one third of the \(\nu_{e}\) capture rate the standard solar model predicted [Davis:1968] [Cleveland:1998]. Because the \(^{8}\mathrm{B}\) flux it was sensitive to depends on roughly the twenty-fifth power of the solar core temperature, the natural reading for twenty years was that the solar model was wrong [Bahcall:1968]. The gallium experiments removed that escape by reaching the \(pp\) neutrinos, whose flux is fixed by the measured solar luminosity almost independently of the model, and found a deficit there too [Hampel:1999] [Abdurashitov:2009]. This is Phenomenon 114.3.
The atmospheric zenith-angle dependence, 1998. Super-Kamiokande measured the \(\nu_{\mu}\) and \(\nu_{e}\) rates as functions of arrival direction, and found the muon neutrinos arriving from below — having crossed \(13000\,\mathrm{km}\) of Earth — depleted by about half relative to those from above, at the same energy, while the electron neutrinos showed no such dependence [Fukuda:1998]. Because the same detector measures both, and because the deficit is a function of \(L/E\) and not of \(L\) or \(E\) separately, no production model and no absorption mechanism reproduces it. This is Phenomenon 114.4, whose derivation there uses exactly Equation (103.52).
The neutral-current measurement, 2002. SNO measured, in one detector, the \(\nu_{e}\) flux by a charged-current reaction on deuterium and the total active flux by the flavour-blind neutral-current breakup of the deuteron. The first came to about a third of the second, and the second agreed with the solar model [Ahmad:2001] [Ahmad:2002]. The subtraction Equation (114.35) contains no solar-model input at all: the missing electron neutrinos arrive as \(\nu_{\mu}\) and \(\nu_{\tau}\), and Phenomenon 114.5 is the statement.
Reactor disappearance and reappearance, 2003. KamLAND observed electron antineutrinos from Japanese power reactors at an average baseline of \(180\,\mathrm{km}\) with a survival probability of about \(0.6\), and — decisively — found that probability to be an oscillatory function of \(L/E\) rather than a monotone suppression [Eguchi:2003]. A flux normalization error produces a deficit; only interference produces a wave. This is Phenomenon 114.6, and it converted the solar result from a statement about the Sun into a statement about neutrinos.
The third angle, 2012. Daya Bay and RENO compared near and far identical detectors at reactor complexes and measured \(\sin^{2}2\theta_{13}\approx0.09\) [An:2012] [Ahn:2012], which is Phenomenon 114.8. Many had expected it to vanish; that it does not is the precondition for ever observing \(\delta_{CP}\), since by Theorem 103.54 the \(CP\)-odd term is proportional to the leptonic Jarlskog invariant \(J_{\ell}=\tfrac{1}{8}\sin2\theta_{12}\sin2\theta_{23} \sin2\theta_{13}\cos\theta_{13}\sin\delta_{CP}\) and therefore to \(\sin\theta_{13}\).
Accelerator beams, 2006–. K2K [Ahn:2006] and MINOS [Michael:2006] confirmed the atmospheric parameters with a controlled beam over \(250\,\mathrm{km}\) and \(735\,\mathrm{km}\); T2K found \(\nu_{\mu}\to\nu_{e}\) appearance [Abe:2011] and now constrains \(\delta_{CP}\) [Abe:2020]; NOvA measures the same parameters over \(810\,\mathrm{km}\) [Acero:2022]; and OPERA observed \(\nu_{\tau}\) appearance directly [Agafonova:2015], closing the three-flavour picture by seeing the flavour the others infer.
Measured parameters and what is still unknown
| Parameter | Value | Principal source |
|---|---|---|
| $\Delta m^{2}_{21}c^{4}$ | \(7.53(18)\times 10^{-5}\,\mathrm{eV}^{2}\) | KamLAND reactor spectrum; solar |
| $\abs{\Delta m^{2}_{32}}c^{4}$ | \(2.455(28)\times 10^{-3}\,\mathrm{eV}^{2}\) | atmospheric; accelerator disappearance |
| $\sin^{2}\theta_{12}$ | \(0.307(13)\) | solar survival probability; KamLAND |
| $\sin^{2}\theta_{23}$ | \(0.546(21)\) | atmospheric and accelerator $\nu_{\mu}$ disappearance |
| $\sin^{2}\theta_{13}$ | \(0.0220(7)\) | short-baseline reactor near/far ratio |
| $\delta_{CP}$ | $\approx1.2\pi$, poorly bounded | $\nu_{\mu}\to\nu_{e}$ versus $\bar{\nu}_{\mu}\to\bar{\nu}_{e}$ |
Three questions remain open within the three-flavour picture itself, and it is worth being exact about what each is.
The mass ordering. The sign of \(\Delta m^{2}_{32}\) is unknown. Normal ordering means \(m_{1}<m_{2}<m_{3}\), with the close pair at the bottom; inverted means \(m_{3}<m_{1}<m_{2}\), with the close pair at the top. Global fits mildly prefer the normal ordering, at a level between two and three standard deviations that depends on which datasets are combined and is not a measurement. The routes to settling it are those of Remark 103.60.
The octant of \(\theta_{23}\). Oscillation of muon neutrinos to tau neutrinos depends on \(\sin^{2}2\theta_{23}\), which is symmetric under \(\theta_{23}\to\pi/2-\theta_{23}\). Since \(\sin^{2}2\theta_{23}=0.99\), the angle is close to \(\pi/4\) and the two solutions are close together; distinguishing them requires the appearance channel, where \(\abs{U_{\mu3}}^{2}=\sin^{2}\theta_{23}\) enters unsquared.
The phase \(\delta_{CP}\). T2K's data disfavour \(CP\) conservation at about the \(95\,\mathrm{\%}\) level [Abe:2020] and prefer a value near \(-\pi/2\); NOvA's are compatible with \(CP\) conservation [Acero:2022]. Combining them gives a preference and not a measurement, and this treatise records it as such. A non-zero \(\delta_{CP}\) would be leptonic \(CP\) violation — the leptonic counterpart of Section 103.5 — and would not, by itself, establish anything about the baryon asymmetry of Phenomenon 103.44, a point taken up in Section 103.7.4.
Neutrino mass
Oscillation measures squared-mass differences and nothing else. Three quite different experiments bound three different combinations of the three masses, and none of them measures a mass eigenvalue. Keeping the three separate is the point of this section, because they are routinely and wrongly quoted as though they were the same number.
The first is an incoherent sum of squares and is always real and positive; the second involves \(U_{ei}^{2}\) rather than \(\abs{U_{ei}}^{2}\) and therefore depends on the Majorana phases of Equation (103.48), so cancellations can drive it to zero; the third is a plain sum.
Direct kinematic limits
The only model-independent route to the mass scale is the shape of a beta spectrum near its endpoint, because there the neutrino is non-relativistic and its rest energy competes with its kinetic energy. The method is that of Section 101.2.3: a non-zero mass removes the last \(m_{\nu}c^{2}\) of electron energy and rounds the approach to the endpoint, changing the Kurie plot from a straight line to a curve that meets the axis vertically.
Measuring the electron spectrum of molecular tritium decay, \(^{3}\mathrm{H}\to{}^{3}\mathrm{He}^{+}+e^{-}+\bar{\nu}_{e}\), with endpoint \(Q=18.574\,\mathrm{keV}\), the KATRIN experiment finds
[Aker:2022], with \(m_{\beta}\) the combination Equation (103.64). The bound assumes only energy and momentum conservation and the measured molecular final-state distribution. Rests on Equation (103.64), Equation (101.11) and Definition 101.19.
Derivation. Derives Phenomenon 103.62. By Equation (101.11), the electron spectrum of an allowed transition is
\(T_{e}\) being the electron kinetic energy. For \(m_{\beta}=0\) this is \(\left(Q-T_{e}\right)^{2}\) near the endpoint, so the Kurie function of Definition 101.19 is linear and reaches zero at \(T_{e}=Q\); for \(m_{\beta}\neq0\) the spectrum terminates at \(T_{e}=Q-m_{\beta}c^{2}\) and approaches it with infinite slope. Two consequences fix the design.
First, the sensitivity is to \(m_{\beta}^{2}\), not to \(m_{\beta}\): the leading distortion of the rate integrated over the last interval \(\Delta\) below the endpoint is a relative shift of order \(m_{\beta}^{2}c^{4}/\Delta^{2}\). Halving the mass sensitivity therefore costs a factor of four in statistics, which is why progress in this field is slow and why an experiment reports a limit on \(m_{\beta}^{2}\) that can and does fluctuate negative.
Second, the region carrying the information is minute. Since \(\dd\Gamma/\dd T_{e}\propto\left(Q-T_{e}\right)^{2}\), the fraction of all decays occurring within \(\Delta\) of the endpoint is
which for \(\Delta=1\,\mathrm{eV}\) and \(Q=18.574\,\mathrm{keV}\) is \(1.6\times 10^{-13}\). An experiment sensitive to \(0.2\,\mathrm{eV}/c^{2}\) must therefore combine a source of enormous activity with a spectrometer of electronvolt resolution at \(18.6\,\mathrm{keV}\) — a relative resolution of \(5\times 10^{-5}\) — and reject every background event in the same window. That is what KATRIN's \(10\,\mathrm{m}\)-diameter magnetic adiabatic collimation spectrometer exists to do.
Tritium is chosen because \(Q\) is the smallest of any convenient superallowed emitter, maximizing Equation (103.68), and because its nuclear matrix element is simple. The residual model dependence is molecular, not nuclear: the daughter \(^{3}\mathrm{HeT}^{+}\) ion is left in a distribution of rotational, vibrational and electronic excitations spread over some \(1\,\mathrm{eV}\), and that distribution is calculated, not measured. It is the leading systematic of Equation (103.67).
∎Equation (103.67) is an order of magnitude weaker than the cosmological bound of Section 103.7.3 and is nevertheless the more valuable number, for the reason that runs through this whole treatise: it is model-independent. It assumes no cosmological model, no growth of structure, no assumption about whether the neutrino is its own antiparticle, and no nuclear matrix element. Its interpretation requires only Equation (103.64) and the measured mixing angles. When a model-dependent bound and a model-independent one disagree, it is the model that is in question, and an experiment that can say so is worth building even at lower sensitivity.
Neutrinoless double beta decay
Double beta decay is the slowest process ever measured, and its existence is a consequence of nuclear structure alone: the pairing term of Equation (107.42) splits the even-\(A\) mass parabola, so an even–even nuclide can lie below both its odd–odd neighbours while lying above the even–even nuclide two steps away. Single beta decay is then closed and double beta decay is open. That argument, the isotopes it selects and the \(Q^{11}\) phase space are Phenomenon 107.72. What belongs here is the lepton-number-violating variant and what its non-observation means.
Two-neutrino double beta decay is observed, with half-lives between \(10^{18}\) and \(10^{24}\) years [Navas:2024] — the range quoted in Phenomenon 107.72, which is the isotope-by-isotope statement. The neutrinoless mode, which would change lepton number by two units and is possible only if the neutrino is its own antiparticle, is not observed: the best searches set half-life limits
at \(90\,\mathrm{\%}\) confidence, from GERDA [Agostini:2020b] and KamLAND-Zen [Abe:2023] respectively. This is stated here as what it is — a search that has found nothing — and it leaves the Dirac or Majorana character of the neutrino undetermined. Rests on Phenomenon 107.72, Proposition 101.32 and Equation (103.65).
Derivation. Derives Phenomenon 103.64. Furry computed the transition \((A,Z)\to(A,Z+2)+2e^{-}\) in 1939 [Furry:1939]: the antineutrino emitted at one nucleon vertex is absorbed as a neutrino at the other. Two conditions must hold for the amplitude not to vanish, and each supplies a factor.
The neutrino must be its own antiparticle, since the particle emitted at the first vertex is what must be absorbed at the second. This is the Majorana condition, and it is what makes the process a test of Section 103.7.4 rather than of anything else.
The neutrino must have mass. The charged current emits a right-helicity antineutrino and absorbs a left-helicity neutrino (Proposition 101.32), so the two vertices demand opposite helicities of the same particle. The amplitude therefore requires the “wrong-helicity” component of the propagating state. Equation (101.32) gives the chirality weight of that component as \(\tfrac{1}{2}(1-\beta)\simeq m^{2}c^{4}/4E^{2}\), which is a probability, so the amplitude carries its square root, \(m_{i}c^{2}/2E\). The internal neutrino carries a virtual momentum of order the inverse internucleon spacing, \(\hbar c/r\approx100\,\mathrm{MeV}\), so the suppression factor is of order \(m_{i}c^{2}/100\,\mathrm{MeV}\), one factor per exchanged mass eigenstate; the factor of two, and the shape of the neutrino propagator between the two nucleons, are conventionally absorbed into the nuclear matrix element defined below rather than carried separately. Summing coherently over the three mass states with the mixing factors \(U_{ei}\) from each vertex, the amplitude is proportional to \(\sum_{i}U_{ei}^{2}m_{i}\) — squared, not modulus-squared, because both vertices are of the same type — which is \(m_{\beta\beta}\) of Equation (103.65). Hence
with \(G^{0\nu}\) a phase-space factor, carrying \(/\mathrm{yr}\), calculable exactly from the \(Q\) value and the Coulomb field, and \(M^{0\nu}\) the dimensionless nuclear matrix element.
Two structural remarks. First, Equation (103.70) is a rate proportional to a squared mass, so a half-life limit improved by a factor \(100\) improves the mass limit only by a factor \(10\) — the same quadratic wall that, in the derivation of Phenomenon 103.62, makes every halving of the kinematic mass sensitivity cost a factor of four in statistics. Second, the electron sum-energy spectrum of the neutrinoless mode is a line at \(Q\), sitting on the continuum of the two-neutrino mode, and the two-neutrino mode is therefore an irreducible background whose tail is controlled only by energy resolution. This is why the two leading techniques are a high-purity germanium diode, whose resolution at \(Q=2039\,\mathrm{keV}\) is about \(3\,\mathrm{keV}\), and a very large mass of xenon-loaded scintillator, which trades resolution for target mass.
∎Converting Equation (103.69) into a bound on \(m_{\beta\beta}\) requires \(M^{0\nu}\), and the many-body methods available — the interacting shell model, the quasiparticle random-phase approximation, the interacting boson model, energy density functionals — disagree with one another by a factor of two to three for every isotope of interest, as Phenomenon 107.72 records. Since the rate goes as the square, that is a factor of four to nine in half-life and a factor of two to three in the mass limit. The published bounds are therefore quoted as ranges:
the spread being theory, not statistics [Agostini:2020b] [Abe:2023]. A treatise of evidence must quote the range and not its most favourable end.
With the mixing parameters of Table 103.3, \(m_{\beta\beta}\) of Equation (103.65) is bounded, over the unknown Majorana phases and the unknown lightest mass, by
so the inverted ordering predicts a floor that the next generation of experiments can reach, and the normal ordering does not. Rests on Equations (103.48) and (103.65).
Derives Proposition 103.66. Write \(m_{\beta\beta}=\abs{c_{12}^{2}c_{13}^{2}m_{1} +s_{12}^{2}c_{13}^{2}m_{2}\ee^{\ii\alpha_{21}} +s_{13}^{2}m_{3}\ee^{\ii\alpha_{31}}}\). In the inverted ordering \(m_{3}\approx0\) and \(m_{1}c^{2}\approx m_{2}c^{2}\approx \sqrt{\abs{\Delta m^{2}_{32}}c^{4}}=50\,\mathrm{meV}\), so the third term is negligible and the first two, being nearly equal in mass, can cancel only to the extent that their coefficients differ: the minimum over \(\alpha_{21}\) is
using \(\cos2\theta_{12}=0.386\) from Table 103.3. The maximum, at \(\alpha_{21}=0\), is \(49\,\mathrm{meV}/c^{2}\). In the normal ordering with \(m_{1}\to0\) the three terms are \(0\), \(s_{12}^{2}c_{13}^{2}\sqrt{\Delta m^{2}_{21}c^{4}}/c^{2} =2.6\,\mathrm{meV}/c^{2}\) and \(s_{13}^{2}\sqrt{\Delta m^{2}_{31}c^{4}}/c^{2} =1.1\,\mathrm{meV}/c^{2}\), which are comparable and can cancel exactly for a suitable phase. There is therefore no floor.
The consequence for the programme is sharp and should be stated without optimism. Equation (103.71) is at the level of the inverted-ordering band, and an experiment covering that band completely would, if it saw nothing, exclude the inverted ordering provided the neutrino is a Majorana particle. It would not exclude the Majorana hypothesis, because the normal ordering permits \(m_{\beta\beta}=0\) exactly. The asymmetry is inherent: the experiment can discover, and can exclude one ordering, but can never refute the Majorana nature.
∎Cosmological bounds
Neutrinos decouple from the primordial plasma while relativistic, and survive as a relic background of number density fixed by the temperature. Once non-relativistic they contribute to the matter density like any other matter; but unlike cold dark matter they retain a large velocity dispersion, and therefore stream freely out of gravitational potential wells smaller than a characteristic scale, suppressing the growth of structure below it.
Three species of light neutrino of total mass \(\Sigma\) contribute
and suppress the matter power spectrum on scales below the free-streaming length by a relative amount \(\Delta P/P\approx-8\Omega_{\nu}/\Omega_{m}\). For \(\Sigma=0.12\,\mathrm{eV}/c^{2}\), \(\Omega_{\nu}h^{2}=1.3\times 10^{-3}\) and the suppression is at the per-cent level — measurable, and the basis of the bound. Rests on Equation (103.66).
Derives Proposition 103.67. The relic number density per species follows from decoupling in thermal equilibrium, the general treatment of which belongs to Section 48.3: a fermion species in equilibrium at temperature \(T_{\nu}\) has \(n_{\nu}=\tfrac{3}{4}\times\tfrac{2\zeta(3)}{\pi^{2}} \left(k_{B}T_{\nu}/\hbar c\right)^{3}\) per degree of freedom, and \(T_{\nu}=\left(4/11\right)^{1/3}T_{\gamma}\) because the electron–positron annihilation that follows decoupling heats the photons and not the neutrinos. With \(T_{\gamma,0}=2.7255\,\mathrm{K}\) this gives \(n_{\nu}=1.12\times 10^{8}\,/\mathrm{m}^{3}\) per species. Dividing the mass density \(\Sigma\,n_{\nu}\) by the critical density \(\rho_{c}=1.87834\times 10^{-26}\,\mathrm{kg}/\mathrm{m}^{3}\,h^{2}\) gives Equation (103.73). The naive arithmetic, \(\rho_{c}c^{2}/n_{\nu}\) per unit \(h^{2}\), comes to \(94.1\,\mathrm{eV}\); the standard constant \(93.14\,\mathrm{eV}\) includes the one-per-cent correction from the fact that decoupling is not instantaneous, which leaves the neutrinos slightly hotter than \(\left(4/11\right)^{1/3}T_{\gamma}\).
For the free streaming: a relic neutrino of mass \(m\) has, at redshift \(z\), a typical momentum \(p\approx3.15k_{B}T_{\nu}(1+z)/c\) and hence a speed \(v/c\approx3.15k_{B}T_{\nu}(1+z)/mc^{2}\). Today \(T_{\nu}=\left(4/11\right)^{1/3}\times2.7255\,\mathrm{K} =1.9454\,\mathrm{K}\), so \(k_{B}T_{\nu}=1.676\times 10^{-4}\,\mathrm{eV}\) and \(3.15k_{B}T_{\nu}=5.28\times 10^{-4}\,\mathrm{eV}\); for \(mc^{2}=0.05\,\mathrm{eV}\) this gives \(v/c=1.06\times 10^{-2}\), that is \(v=3.2\times 10^{3}\,\mathrm{km}/\mathrm{s}\). Over a Hubble time it therefore travels a comoving distance of order \(v/H\), and perturbations of smaller wavelength are erased because the neutrinos simply leave them. The suppression coefficient \(8\) follows from solving the equation for the growth of density perturbations, which belongs to Section 48.5.1, with a component that does not cluster; the essential point for this chapter is that the effect is proportional to \(\Omega_{\nu}\), hence to \(\Sigma\), and is therefore a measurement of the sum of the masses and of no other combination.
∎Combining the temperature and polarization anisotropies of the cosmic microwave background with the baryon acoustic oscillation scale measured in galaxy surveys,
[Aghanim:2020], within the six-parameter \(\Lambda\)CDM model whose parameters belong to Section 48.7. Rests on Proposition 103.67 and Phenomenon 48.15.
Derivation. Derives Phenomenon 103.68. The microwave background alone constrains \(\Sigma\) only weakly, because a neutrino of mass below about \(0.6\,\mathrm{eV}/c^{2}\) is still relativistic at recombination and its effect on the acoustic peaks is degenerate with other parameters; the constraint comes from combining it with a low-redshift measurement of the expansion history and of the clustering amplitude, which is what the baryon acoustic scale of Phenomenon 48.15 supplies. Extracting Equation (103.74) from those data is a global fit, not a closed-form derivation, and it is quoted here for comparison with Equation (103.67) rather than reproduced.
The comparison is the point. Equation (103.74) is a factor \(20\) stronger than the kinematic bound — and rests on the assumptions that the dark energy is a cosmological constant, that the primordial spectrum is a power law, and that no other light relic contributes. Each has been examined and none is established independently. Relaxing them typically weakens the bound by a factor of two to three, which is the honest uncertainty to attach to it.
∎The minimum values of \(\Sigma\) permitted by the measured splittings are, taking the lightest mass to zero,
computed from \(m_{2}c^{2}=\sqrt{\Delta m^{2}_{21}c^{4}}\) and \(m_{3}c^{2}=\sqrt{\Delta m^{2}_{31}c^{4}}\) in the normal case, and from \(m_{1}c^{2}=\sqrt{\abs{\Delta m^{2}_{31}}c^{4}}\), \(m_{2}c^{2}=\sqrt{\abs{\Delta m^{2}_{32}}c^{4}}\) in the inverted case. The inverted minimum sits just below Equation (103.74), so cosmology mildly disfavours the inverted ordering — and would exclude it if the bound tightened by a factor of two without a detection. That is a striking situation and it deserves a warning: a cosmological bound excluding a laboratory mass ordering would be an inference of exactly the model-dependent kind Remark 103.63 cautions against. The ordering should be settled by Remark 103.60, in an experiment, and this treatise will record it as settled only then.
Dirac or Majorana: an open question
Every charged fermion has a Dirac mass, because a Majorana mass term would violate electric charge conservation. The neutrino is the only known fermion for which both options are open, and no experiment has distinguished them.
For a left-chiral field \(\nu_{L}\) and, if it exists, a right-chiral partner \(\nu_{R}\), the Lorentz-invariant mass terms are
with \(\nu^{c}\) the charge-conjugate field of Equation (105.15). The first conserves lepton number and requires a new field; the second violates lepton number by two units and requires none.
The full analysis of these two terms, the degrees-of-freedom count that makes a Majorana field equivalent to one massive Weyl field, and Majorana's original symmetric theory [Majorana:1937] belong to Section 96.6.1 and Section 96.7.1, which reserve them. What belongs here is the experimental status, and it can be stated in four sentences.
No experiment distinguishes the two. Oscillation cannot, by Proposition 103.47; kinematic beta-decay measurements cannot, since Equation (103.64) is phase-independent; cosmology cannot, since Equation (103.66) is too. The only practical discriminator is neutrinoless double beta decay, and by Equation (103.69) it has found nothing.
A positive result would settle it whatever the mechanism. If the decay is observed, a Majorana mass term for the neutrino follows, even if the operator that drives the decay is something else entirely: the observed transition can always be closed into a \(\bar{\nu}\to\nu\) self-energy insertion, generating Equation (103.77) at some order. This is the Schechter–Valle argument [Schechter:1982]. It is a statement about the existence of the term, not about its size, which the argument leaves many orders of magnitude below the observed masses — a limitation worth stating, since the theorem is often quoted as though it were quantitative.
The most economical hypothesis has no evidence. If a heavy right-chiral singlet \(N\) exists with a Majorana mass \(M\) far above the electroweak scale, the light eigenvalue of the combined mass matrix is \(m_{\nu}\approx m_{D}^{2}/M\) — small because \(M\) is large. This is the seesaw [Minkowski:1977]; with \(m_{D}\) of the order of the top mass, \(Mc^{2}\sim10^{15}\,\mathrm{GeV}\). In the effective theory below \(M\) it is one realization of the unique dimension-five operator that can generate a neutrino mass from Standard Model fields [Weinberg:1979b], whose coefficient carries \(/\mathrm{J}\) and is therefore suppressed by one power of a new scale. That operator is a parametrization of ignorance and the seesaw is one model for its coefficient. There is at present no experimental evidence for either, and this treatise names them as hypotheses and nothing more.
Sterile neutrinos: null. A fourth neutrino not coupling to the \(Z\) evades Equation (103.15) and would appear as a fourth oscillation frequency. The LSND excess [Aguilar:2001] and the MiniBooNE low-energy excess have been read that way; the MicroBooNE liquid-argon measurement finds no corresponding excess in the electron-neutrino channel [Abratenko:2022], and the reactor and gallium rate anomalies have been partly resolved by recalculating fluxes. No sterile state is established and the anomalies have not all gone away; Section 114.6.4 is the section reserved for the state of each.
The heavy singlets of the seesaw, decaying out of equilibrium with a \(CP\)-violating asymmetry, would generate a lepton asymmetry that electroweak processes partly convert into the baryon asymmetry of Equation (103.45). The mechanism is attractive precisely because it needs no new ingredient beyond the one already invoked to explain the neutrino masses. It is nevertheless untested and, with the relevant scale above \(10^{9}\,\mathrm{GeV}\), untestable by any foreseeable experiment; and a measurement of \(\delta_{CP}\) would not establish it, because the phases entering the heavy-neutrino decays are not the phases entering Equation (103.47). It is listed in Section 133.5 as an unexplained observation, not as a solved problem.
Open questions in flavour
The Standard Model has nineteen free parameters, and thirteen of them — nine charged-fermion masses, three CKM angles and one phase — belong to flavour; the neutrino sector adds at least seven more. Every one is measured and none is predicted. This closing section names the specific questions that are open, distinguishing those where a measurement disagrees with a prediction from those where there is no prediction to disagree with.
The mass hierarchy. The charged-fermion masses span from \(0.511\,\mathrm{MeV}/c^{2}\) to \(172.57\,\mathrm{GeV}/c^{2}\), a factor \(3.4\times 10^{5}\) within one set of gauge quantum numbers, and the neutrino masses lie a further six orders of magnitude below the electron. Expressed as Yukawa couplings by Remark 103.21 they run from \(1\) down to \(10^{-6}\) for the charged leptons and to \(10^{-12}\) or less for neutrinos. There is no prediction here to disagree with; there is a pattern with no explanation, and it is Section 133.1.
Charged-lepton flavour violation: not observed. In the Standard Model with massive neutrinos, \(\mu\to e\gamma\) proceeds through a loop with the GIM suppression of Theorem 103.11, and because the internal masses are neutrino masses the suppression factor \(\left(\Delta m^{2}c^{4}\right)^{2} /\left(M_{W}c^{2}\right)^{4}\) — with \(\Delta m^{2}c^{4}=2.5\times 10^{-3}\,\mathrm{eV}^{2}\) and \(M_{W}c^{2}=80.4\,\mathrm{GeV}\), equal to \(1.5\times 10^{-49}\) — makes it utterly unobservable. Any observation would therefore be new physics, which is why the search is worth pursuing. The MEG II experiment finds
from its first dataset, and \(1.5\times 10^{-13}\) combining with the earlier MEG result, both at \(90\,\mathrm{\%}\) confidence [Afanaciev:2024]. Phenomenon 101.53 records the same null result in its charged-current context and quotes the same two values. Nothing is seen.
The \(B\)-decay anomalies: mostly resolved, and the resolution is instructive. For a decade the ratios \(R_{K}=\mathcal{B}\left(B\to K\mu^{+}\mu^{-}\right) /\mathcal{B}\left(B\to Ke^{+}e^{-}\right)\) and its \(K^{*}\) analogue were measured below the Standard Model value of unity, at the level of \(2.5\sigma\) each. Lepton universality is a sharp prediction: the ratio is one up to phase-space and radiative corrections of a per cent, because the gauge couplings do not distinguish the flavours. In 2022 LHCb identified a mis-estimated electron-identification efficiency, and the corrected measurements are
consistent with unity [Aaij:2023]. This treatise records the episode rather than quietly dropping it, because it is a clean example of an anomaly that was a systematic effect — and because the related \(b\to s\mu\mu\) branching fractions and angular observables remain about \(3\sigma\) below their Standard Model predictions, with hadronic uncertainties that are disputed. Universality holds; the absolute rates are not settled.
The muon anomalous magnetic moment. The muon's anomaly is sensitive to any flavour-dependent physics at high scale, and its Standard Model prediction is limited by the hadronic vacuum polarization, which belongs to Section 112.6.3. The status — including the disagreement between the data-driven and lattice evaluations of that contribution, which is larger than the effect being tested — belongs to Section 112.6.4. It is quoted here only to record that the one place where a flavour-sensitive precision observable may disagree with theory is a place where the theory prediction is itself contested.
Three generations, and no reason. The number is measured twice, by Equation (103.15) and by nucleosynthesis, and predicted by nothing.
First-row unitarity. Proposition 103.8: the most precisely tested relation in quark flavour physics is currently low by two to three standard deviations. It is the one entry in this list that is a genuine discrepancy between a prediction and a measurement rather than an absence of prediction, and it is most likely a systematic effect in one of the inputs.
Every item above is listed as open in What We Observe but Do Not Understand rather than resolved by a model this treatise would have no evidence to include. The models exist — horizontal symmetries relating the generations, additional gauge bosons distinguishing them, composite fermions, extra spatial dimensions in which the Yukawa hierarchy becomes a geometric one — and every one of them predicts something that has not been seen. Under editorial rule 1 they are named here, once, so that a reader knows they were considered, and are not developed.