Discrete Symmetries and CPT
Three transformations lie outside the connected Lorentz group of Minkowski Space and Its Symmetries: space inversion \(P\), charge conjugation \(C\) and time reversal \(T\). Each was assumed to be an exact symmetry of nature, and each turned out not to be. Parity fell first, proposed as testable by Lee and Yang [Lee:1956] and demolished by Wu's cobalt-60 experiment [Wu:1957], the subject of Experiment: Parity Violation; the combination \(CP\), offered by Landau as the replacement [Landau:1957], fell in turn with the \(2\pi\) decay of the long-lived kaon [Christenson:1964], the subject of Experiment: CP Violation; and direct violation of \(T\) was observed in the neutral kaon [Angelopoulos:1998] and in the neutral \(B\) system [Lees:2012]. This chapter collects those results and the theory that survives them.
What survives is the product. The \(CPT\) theorem of Lüders [Lueders:1954], Pauli [Pauli:1955] and Jost [Jost:1957] says that any local, Lorentz-invariant quantum field theory with a positive-energy spectrum is invariant under \(CPT\) — so particle and antiparticle must have equal masses and lifetimes and opposite magnetic moments. It is therefore the sharpest available test of the axioms of Axiomatic Quantum Field Theory, and the chapter reviews the measurements that probe it, from antihydrogen spectroscopy [Ahmadi:2017] [Ahmadi:2018] to the antiproton magnetic moment [Smorra:2017] and charge-to-mass ratio [Borchert:2022]. It closes on two honest failures: Sakharov's conditions for a baryon asymmetry [Sakharov:1967], which the observed \(CP\) violation falls short of by many orders of magnitude, and the strong \(CP\) problem, bounded but unexplained by the neutron electric dipole moment [Abel:2020].
The conventions are those fixed for quantum electrodynamics in Section 100.1.1, and are used without further comment: the metric is \(\eta_{\mu\nu}=\diag(+1,-1,-1,-1)\), coordinates are \(x^{\mu}=(ct,\vect{x})\), four-momenta are \(p^{\mu}=(E/c,\vect{p})\) with the mass shell \(p^{2}=m^{2}c^{2}\) of Equation (100.1), and \(\hbar\) and \(c\) appear explicitly in every formula. Exactly one symbol is reassigned against them, and the reassignment is declared here rather than left to the reader to discover: Notation 100.1 reserves \(\epsilon\), standing alone, for the positive infinitesimal of the Feynman contour, whereas from Section 105.3.1 onwards \(\epsilon\) is the \(CP\)-violating mixing parameter of the neutral kaon, which is what the symbol means everywhere in this subject. The contour prescription is accordingly written \(+\ii0\) in this chapter and never \(+\ii\epsilon\).
Operators on the Hilbert space are written in calligraphic type — \(\mathcal{P},\mathcal{C},\mathcal{T}\) — with a single exception: their product is written \(\Theta\) (Definition 105.33), the capital theta being universal for it. The Dirac matrices \(\gamma^{\mu}\) and \(\gamma^{5}\) are the ordinary italic ones of The Dirac Equation; the numerical \(4\times4\) matrices that implement \(P\), \(C\) and \(T\) on Dirac indices are written in upright roman, and carry the subscript D where the letter would otherwise collide with an operator — the charge-conjugation matrix \(\mathrm{C}_{\mathrm{D}}\) of Definition 100.9 and the time-reversal matrix \(\mathrm{T}_{\mathrm{D}}\), against the parity matrix \(\mathrm{S}_{P}\), whose letter collides with nothing. The two kinds must not be confused: \(\mathcal{C}\) is a unitary operator on a Fock space, \(\mathrm{C}_{\mathrm{D}}\) a matrix one can write down.
The three discrete operations
The Lorentz group \(L\) has four connected components (Proposition 39.8), and the three not containing the identity are reached from it by adjoining the two matrices
acting on \(x^{\mu}=(ct,\vect{x})\). Neither can be reached by any continuous motion, so whether Nature is invariant under them is not settled by the Lie algebra of Section 39.2.7: it is an experimental question, and Section 105.2 records the answer.
Charge conjugation is of a different kind. It is not a spacetime transformation at all: it acts on the internal labels of the fields, exchanging each particle with its antiparticle while leaving momentum and spin untouched. It has no classical counterpart in the sense that \(P\) and \(T\) do, and it exists only because the relativistic wave equations of The Dirac Equation force antiparticles into the theory — which is why the empirical discovery of \(C\) had to wait for the positron of Experiment: The Positron.
By Wigner's theorem (Theorem 94.3) each of the three, being a symmetry of the ray space, is implemented by an operator that is either unitary or antiunitary. The theorem does not say which. The answer, derived in Section 105.1.3, is that \(\mathcal{P}\) and \(\mathcal{C}\) may be taken unitary while \(\mathcal{T}\) must be antiunitary, and that single structural difference is responsible for most of what distinguishes \(T\) from its two companions: \(T\) carries no conserved quantum number, forbids no transition outright, and can be tested only by comparing a process with its motion reverse.
Parity
Parity in quantum mechanics
Wigner introduced parity as a conserved quantum number in 1927 [Wigner:1927], reading it off the reflection symmetry of atomic spectra. The argument is short enough to give in full, and it fixes the pattern that the field-theoretic version will repeat.
Let a spinless particle move in a central potential, so that the Hamiltonian is
Define the unitary operator \(\mathcal{P}\) on wavefunctions by
with \(\eta_{P}\) a phase. Then \(\mathcal{P}\hat{\vect{x}} \mathcal{P}^{-1}=-\hat{\vect{x}}\) and, because \(\hat{\vect{p}}=-\ii\hbar\nabla\) and \(\nabla\to-\nabla\) under \(\vect{x}\to-\vect{x}\), also \(\mathcal{P}\hat{\vect{p}} \mathcal{P}^{-1}=-\hat{\vect{p}}\). Both terms of Equation (105.2) are therefore invariant, \(\comm{\mathcal{P}}{H}=0\), and energy eigenstates may be chosen to be eigenstates of \(\mathcal{P}\). Since \(\mathcal{P}^{2}\) acts as the identity on \(\vect{x}\), we may choose \(\eta_{P}^{2}=1\) for a scalar wavefunction, so the eigenvalues are \(\pm1\): even and odd states.
Orbital angular momentum \(\hat{\vect{L}} =\hat{\vect{x}}\times\hat{\vect{p}}\) is a product of two odd vectors and is therefore even: \(\mathcal{P}\hat{\vect{L}} \mathcal{P}^{-1}=+\hat{\vect{L}}\). A quantity that behaves this way under \(P\) — like a vector under rotations but even under reflection — is called an axial (or pseudo-) vector, as against the polar vectors \(\vect{x}\) and \(\vect{p}\); the corresponding scalars are called pseudoscalar and scalar. The spherical harmonics inherit the eigenvalue
so a state of orbital angular momentum \(\ell\) has parity \((-1)^{\ell}\) times the intrinsic parities of its constituents.
In a parity-conserving theory, an electric-dipole transition connects only states of opposite parity. In particular no \(E1\) line joins two even or two odd atomic configurations. Rests on Equation (105.3).
Derives Proposition 105.1. The electric-dipole operator is \(\hat{\vect{d}}=-e\sum_{a} \hat{\vect{x}}_{a}\), a sum of polar vectors, so \(\mathcal{P}\hat{\vect{d}}\mathcal{P}^{-1}=-\hat{\vect{d}}\). If \(\mathcal{P}\ket{i}=\pi_{i}\ket{i}\) and \(\mathcal{P}\ket{f}=\pi_{f}\ket{f}\) with \(\pi_{i,f}=\pm1\), then inserting \(\mathcal{P}^{-1}\mathcal{P}=\identity\) twice gives
so the matrix element vanishes unless \(\pi_{i}\pi_{f}=-1\).
∎This is the selection rule that Laporte extracted empirically from iron and titanium spectra in 1925 [Laporte:1925], two years before Wigner explained it: the terms of a complex spectrum fall into two classes, and lines occur only between classes and never within one. It is the earliest quantitative evidence for parity as a physical quantum number, and it is worth noticing what kind of evidence it is — a systematic absence, which is what a symmetry always delivers.
Parity on the fields
In field theory \(\mathcal{P}\) is a unitary operator on the Fock space, defined by its action on the fields. For a spinless field the possibilities are two,
and the field is called scalar for \(\eta_{P}=+1\) and pseudoscalar for \(\eta_{P}=-1\). The number \(\eta_{P}\) is the intrinsic parity of the corresponding particle.
For the electromagnetic potential the transformation is not a matter of choice. The observable fields are \(\vect{E}=-\nabla\phi_{\mathrm{el}}-\pp_{t}\vect{A}\) and \(\vect{B}=\nabla\times\vect{A}\); under \(\vect{x}\to-\vect{x}\) one has \(\nabla\to-\nabla\), so the assignment
gives \(\vect{E}\to-\vect{E}\) and \(\vect{B}\to+\vect{B}\): the electric field is polar, the magnetic field axial. Both signs are forced by Equation (105.6) and neither is conventional. In compact notation \(\mathcal{P}A^{\mu}(x)\mathcal{P}^{-1} =\Lambda_{P}{}^{\mu}{}_{\nu}A^{\nu}(x_{P})\) with \(x_{P}=(ct,-\vect{x})\); the photon has intrinsic parity \(-1\), since the spatial polarization vector, which is what a physical photon carries, reverses.
For the Dirac field the transformation is likewise forced, up to one phase.
Let \(\psi\) satisfy \((\ii\hbar\gamma^{\mu}\pp_{\mu}-mc)\psi=0\). Then
maps solutions to solutions, and \(\gamma^{0}\) is the unique matrix (up to a factor) that does so. Rests on Equations (100.5) and (105.1).
Derives Proposition 105.2. Write \(\psi'(x_{P})=\mathrm{S}_{P}\psi(x)\) and demand that \(\psi'\) satisfy the Dirac equation in the reflected coordinates. Since \(\pp/\pp x_{P}^{0}=\pp_{0}\) and \(\pp/\pp x_{P}^{i}=-\pp_{i}\), the requirement is
Multiplying the original equation on the left by \(\mathrm{S}_{P}\) and inserting \(\mathrm{S}_{P}^{-1}\mathrm{S}_{P}\) before \(\psi\) gives
so the two agree if and only if \(\mathrm{S}_{P}\gamma^{0}\mathrm{S}_{P}^{-1}=\gamma^{0}\) and \(\mathrm{S}_{P}\gamma^{i}\mathrm{S}_{P}^{-1}=-\gamma^{i}\). The Clifford relation \(\acomm{\gamma^{\mu}}{\gamma^{\nu}}=2\eta^{\mu\nu}\identity\) gives \(\gamma^{0}\gamma^{0}\gamma^{0}=\gamma^{0}\) and \(\gamma^{0}\gamma^{i}\gamma^{0}=-\gamma^{i}\), since \((\gamma^{0})^{2}=\identity\), so \(\mathrm{S}_{P}=\gamma^{0}\) works. Uniqueness up to a factor follows because any other solution \(\mathrm{S}'\) makes \(\mathrm{S}'\gamma^{0}\) commute with all four \(\gamma^{\mu}\), and the Dirac algebra is irreducible, so by Schur's lemma (Linear Algebra and Representation Theory) \(\mathrm{S}'\gamma^{0}\) is a multiple of the identity.
∎The phase \(\eta_{P}\) is a new label, and it is worth saying exactly where it comes from. Wigner's classification in Particles as Poincaré Representations sorts one-particle states by the two Casimir invariants of the connected Poincaré group, mass and spin (Theorem 94.29); \(\mathcal{P}\) is not in that group, so the classification says nothing about it. Enlarging the group by \(\Lambda_{P}\) enlarges the label set by one sign per irreducible representation, and that sign is the intrinsic parity. It is therefore a representation label of the full Lorentz group, and it exists as a quantum number only in a sector where \(P\) is a symmetry — which is why the weak interaction does not assign one.
Applying Equation (105.7) twice returns \(\eta_{P}^{2}\psi\), so \(\eta_{P}^{2}=\pm1\) and \(\eta_{P}\in\set{\pm1,\pm\ii}\): a fermion's intrinsic parity need not be real. This is not a defect but a reflection of the fact that fermion number is superselected, so that a state of one fermion and a state of none are never superposed and the phase never interferes with anything. What is measurable is a relative parity, and one relative parity is fixed by the theory rather than by convention.
If the field \(\psi\) transforms as Equation (105.7) with phase \(\eta_{P}\), the charge-conjugate field \(\psi^{c}=\mathrm{C}_{\mathrm{D}}\bar{\psi}\transpose\) of Definition 100.9 transforms with phase \(-\eta_{P}^{*}\). Rests on Equation (105.7), Definition 100.9 and Equation (100.20).
Derives Proposition 105.3. From Equation (105.7), \(\mathcal{P}\bar{\psi}(x)\mathcal{P}^{-1} =\eta_{P}^{*}\bar{\psi}(x_{P})\gamma^{0}\), since \(\gamma^{0\dagger}=\gamma^{0}\) and \(\bar{\psi}=\psi^{\dagger} \gamma^{0}\). Hence
Equation (100.20) with \(\mu=0\) reads \(\mathrm{C}_{\mathrm{D}}\gamma^{0\transpose} \mathrm{C}_{\mathrm{D}}^{-1}=-\gamma^{0}\), so \(\mathrm{C}_{\mathrm{D}}\gamma^{0\transpose} =-\gamma^{0}\mathrm{C}_{\mathrm{D}}\) and the right-hand side is \(-\eta_{P}^{*}\gamma^{0}\psi^{c}(x_{P})\).
∎An electron–positron bound state with relative orbital angular momentum \(L\) has parity
by Proposition 105.3 and Equation (105.4). The ground states are therefore odd, \(J^{P}=0^{-}\) for para-positronium and \(1^{-}\) for ortho-positronium, which is why the \(2\gamma\) decay of the former proceeds to a state of odd parity and not, as one might carelessly guess, to an even one.
The bilinears, and where parity can hide
Everything that can appear in a local Lagrangian built from a Dirac field is a sum of the sixteen bilinears \(\bar{\psi}\Gamma\psi\) with \(\Gamma\) running over \(\identity\), \(\gamma^{5}\), \(\gamma^{\mu}\), \(\gamma^{\mu}\gamma^{5}\) and \(\sigma^{\mu\nu} =\tfrac{\ii}{2}\comm{\gamma^{\mu}}{\gamma^{\nu}}\), possibly carrying derivatives. Their behaviour under \(P\) follows from Equation (105.7) by one line of Clifford algebra each. Since \(\gamma^{0}\gamma^{\mu}\gamma^{0}=\gamma_{\mu}\) — meaning \(+\gamma^{0}\) for \(\mu=0\) and \(-\gamma^{i}\) for \(\mu=i\) — and \(\gamma^{0}\gamma^{5}\gamma^{0}=-\gamma^{5}\) because \(\gamma^{5}\) anticommutes with every \(\gamma^{\mu}\), one obtains
The four names — scalar \(S\), pseudoscalar \(\tilde{S}\), vector \(V^{\mu}\), axial vector \(A^{\mu}\) — are exactly these four lines, and Table 105.1 collects them together with the \(C\) and \(T\) columns derived below.
The interaction of quantum electrodynamics is \(-j^{\mu}A_{\mu}\) with \(j^{\mu}=qc\,\bar{\psi}\gamma^{\mu}\psi\) the vector current of Equation (100.6). Comparing Equation (105.11) with Equation (105.6), the two lower an index in the same way, so the contraction \(j^{\mu}A_{\mu}\) is invariant and \(P\) is an exact symmetry of QED — the statement made without proof in Section 100.1.4 and proved here. The same argument applies verbatim to the quark–gluon coupling of Quantum Chromodynamics, which is also a vector current, with one exception that is the subject of Section 105.7: the topological term permitted in the QCD Lagrangian is \(P\)-odd, and only experiment tells us that its coefficient is tiny.
What Equation (105.12) shows is where parity violation must live. If a Lagrangian contains both \(V^{\mu}\) and \(A^{\mu}\) coupled to the same object, then under \(P\) the relative sign of the two flips and the theory is not invariant. Parity is violated exactly when vector and axial currents interfere — which is why the discovery of Section 105.2.2 took the form it did.
Multiplicativity and the parity of the pion
Because \(\mathcal{P}\) is a unitary operator implementing a \(\Z_{2}\) operation, the parity of a composite state is the product of the intrinsic parities of its constituents times \((-1)^{L}\) for each relative orbital angular momentum. Parity is therefore a multiplicative quantum number, unlike charge or baryon number, which are additive; this is the general fact that a symmetry group \(\Z_{2}\) has one-dimensional representations labelled by a sign.
Intrinsic parities are measurable only relative to a convention, and the convention universally adopted is that the proton, the neutron and the \(\Lambda\) carry \(\eta_{P}=+1\). Everything else is then determined by experiment. The classical determination is the pion's, and it is worth giving in full because it uses nothing but the two counting rules just established plus the antisymmetry of identical fermions.
The reaction \(\pi^{-}+d\to n+n\), in which a negative pion is captured at rest from an atomic orbit of pionic deuterium, proceeds at a rate comparable to the competing radiative channel \(\pi^{-}+d\to n+n+\gamma\) [Panofsky:1951]. Its occurrence at all requires the pion to have intrinsic parity \(\eta_{P}(\pi)=-1\), given \(\eta_{P}(n)=+1\). Rests on Equation (105.4).
Derivation. Derives Phenomenon 105.5. A negative pion stopped in liquid deuterium is captured into a Bohr orbit of the pionic deuterium atom and cascades downward. Because the pion is some \(273\) times heavier than the electron, its Bohr radius is smaller by the same factor and the overlap with the deuteron is enormous; the cascade terminates in the \(1S\) state, from which nuclear capture occurs. The initial orbital angular momentum is therefore \(L_{i}=0\), and since the pion is spinless and the deuteron has spin \(1\), the initial total angular momentum is \(J=1\). The initial parity is
because the deuteron, a \(\,^{3}S_{1}\) bound state of two nucleons of parity \(+1\) with a small \(\,^{3}D_{1}\) admixture, has \(\eta_{P}(d)=+1\): both \(L=0\) and \(L=2\) are even.
The final state is two neutrons, identical fermions, so the total wavefunction must be antisymmetric under exchange. Exchanging the two neutrons multiplies the spatial part by \((-1)^{L}\) and the spin part by \((-1)^{S+1}\) — the triplet \(S=1\) is symmetric and the singlet \(S=0\) antisymmetric — so the requirement is
Enumerate the states with \(J=1\): \(\,^{1}S_{0}\) has \(J=0\); \(\,^{3}S_{1}\) has \(L+S=1\), odd, excluded; \(\,^{1}P_{1}\) has \(L+S=1\), excluded; \(\,^{3}D_{1}\) has \(L+S=3\), excluded; and \(\,^{3}P_{1}\) has \(L+S=2\), even, and \(J=1\). Exactly one state survives, \(\,^{3}P_{1}\), of parity \((-1)^{L}=(-1)^{1}=-1\).
Equating initial and final parity gives \(\eta_{P}(\pi)=-1\). The experimental input is that the reaction occurs: Panofsky, Aamodt and Hadley measured the ratio of the two capture channels and found the non-radiative one present at a comparable rate [Panofsky:1951], which cannot happen if the only accessible final state is forbidden by parity.
∎The same \(J^{P}=0^{-}\) assignment for the neutral pion follows from the observed \(\pi^{0}\to\gamma\gamma\) angular correlation, and the fact that all three pions form an isospin triplet with the same spin and parity ties them together. A pseudoscalar pion is the reason a two-pion state has parity \((+1)(-1)^{L}\) rather than \((-1)^{L}\), and it is what makes \(K\to2\pi\) and \(K\to3\pi\) distinguish parities in Section 105.2.1.
Charge conjugation
The operation on the Dirac field
The Dirac equation in an external electromagnetic field,
with \(D_{\mu}\) the covariant derivative of Equation (100.9), describes a particle of charge \(q\) and mass \(m\). Its solutions of negative frequency were interpreted by Dirac as antiparticles and observed as the positron (Experiment: The Positron). Charge conjugation is the statement that the equation with \(q\) replaced by \(-q\) has exactly as many solutions, in explicit correspondence.
Let \(\mathrm{C}_{\mathrm{D}}\) satisfy Equation (100.20), \(\mathrm{C}_{\mathrm{D}}(\gamma^{\mu})\transpose \mathrm{C}_{\mathrm{D}}^{-1}=-\gamma^{\mu}\). If \(\psi\) solves Equation (105.14), then
solves the same equation with \(q\to-q\), with the same mass. Rests on Equations (100.20) and (105.14).
Derives Proposition 105.6. Expand Equation (105.14): \(\ii\hbar\gamma^{\mu}\pp_{\mu}\psi-q\gamma^{\mu}A_{\mu}\psi-mc\psi=0\). Take the Hermitian conjugate, using \(\gamma^{\mu\dagger} =\gamma^{0}\gamma^{\mu}\gamma^{0}\) and \(A_{\mu}\) real:
Multiply on the right by \(\gamma^{0}\) and use \(\gamma^{\mu\dagger}\gamma^{0}=\gamma^{0}\gamma^{\mu}\) together with \(\bar{\psi}=\psi^{\dagger}\gamma^{0}\):
This is a row-spinor equation; transpose it to obtain a column,
and multiply on the left by \(\mathrm{C}_{\mathrm{D}}\), inserting \(\mathrm{C}_{\mathrm{D}}^{-1}\mathrm{C}_{\mathrm{D}}\) in front of each \(\bar{\psi}\transpose\). With \(\mathrm{C}_{\mathrm{D}}(\gamma^{\mu})\transpose \mathrm{C}_{\mathrm{D}}^{-1}=-\gamma^{\mu}\) and \(\psi^{c}=\mathrm{C}_{\mathrm{D}}\bar{\psi}\transpose\) this reads
which is Equation (105.14) with \(q\to-q\).
∎In the Dirac representation one may take \(\mathrm{C}_{\mathrm{D}}=\ii\gamma^{2}\gamma^{0}\), as recorded in Definition 100.9; then \(\mathrm{C}_{\mathrm{D}}\transpose=-\mathrm{C}_{\mathrm{D}}\) and \(\mathrm{C}_{\mathrm{D}}^{\dagger}=\mathrm{C}_{\mathrm{D}}^{-1} =-\mathrm{C}_{\mathrm{D}}\). Two features of Proposition 105.6 deserve emphasis. First, the mass is untouched: charge conjugation relates a particle to something of the same mass, which is a theorem of the free theory and, once interactions are switched on, becomes the much deeper statement of Section 105.5. Second, \(\psi^{c}\) is built from \(\bar{\psi}\transpose\), i.e. from \(\psi^{*}\), so the operation involves complex conjugation of the wavefunction while remaining, on the Fock space, a unitary operator: the antilinearity is absorbed into the exchange of creation and annihilation operators. In terms of modes, \(\mathcal{C}\) maps the operator that annihilates an electron of momentum \(\vect{p}\) and spin \(s\) into the one that annihilates a positron of the same \(\vect{p}\) and the same \(s\). Momentum and spin are spectators; only the charge-like labels move.
C-parity as a quantum number
Since \(\mathcal{C}\) exchanges a particle with a different particle, a charged state is never an eigenstate: \(\mathcal{C}\ket{e^{-}} \propto\ket{e^{+}}\), which is orthogonal to it. Only a state that is its own antiparticle can carry a \(C\) eigenvalue — the photon, the \(\pi^{0}\), the \(\eta\), positronium, a neutral non-strange meson–antimeson combination.
The photon's value is forced. The interaction \(-j^{\mu}A_{\mu}\) must be \(C\)-even (it is the coupling that generates all of atomic physics, and \(C\) is an exact symmetry of QED), while the current \(j^{\mu}=qc\,\bar{\psi}\gamma^{\mu}\psi\) is \(C\)-odd. To see the latter, apply \(\mathcal{C}\psi\mathcal{C}^{-1} =\eta_{C}\mathrm{C}_{\mathrm{D}}\bar{\psi}\transpose\) and \(\mathcal{C}\bar{\psi}\mathcal{C}^{-1} =-\eta_{C}^{*}\psi\transpose\mathrm{C}_{\mathrm{D}}^{-1}\) to a bilinear. The two fermion operators must be reordered, which costs a minus sign, and the result is
For \(\Gamma=\gamma^{\mu}\) the defining relation gives \(\mathrm{C}_{\mathrm{D}}^{-1}\gamma^{\mu}\mathrm{C}_{\mathrm{D}} =-(\gamma^{\mu})\transpose\), whose transpose is \(-\gamma^{\mu}\), so \(V^{\mu}\to-V^{\mu}\). Hence
and an \(n\)-photon state has \(C=(-1)^{n}\). The same computation applied to the other four \(\Gamma\)'s gives the \(C\) column of Table 105.1; the only entry that requires a moment's care is \(\gamma^{5}=\ii\gamma^{0}\gamma^{1}\gamma^{2}\gamma^{3}\), for which \(\mathrm{C}_{\mathrm{D}}^{-1}\gamma^{5}\mathrm{C}_{\mathrm{D}} =\ii(-\gamma^{0\transpose})(-\gamma^{1\transpose}) (-\gamma^{2\transpose})(-\gamma^{3\transpose}) =\ii\left(\gamma^{3}\gamma^{2}\gamma^{1}\gamma^{0}\right)\transpose =(\gamma^{5})\transpose\), because reversing four mutually anticommuting factors costs six transpositions and hence no sign. So the pseudoscalar bilinear is \(C\)-even.
The decay \(\pi^{0}\to\gamma\gamma\) is observed with a branching fraction of \(98.823(34)\,\mathrm{\%}\) [Navas:2024], so \(\eta_{C}(\pi^{0})=(-1)^{2}=+1\). The decay \(\pi^{0}\to3\gamma\) would require \(C=(-1)^{3}=-1\) and is forbidden. The experimental limit is a branching fraction below \(3.1\times10^{-8}\) at \(90\,\mathrm{\%}\) confidence [Navas:2024] — a suppression of more than seven orders of magnitude relative to the allowed channel, where phase space alone would cost only about two. The missing five orders are the symmetry.
For a fermion–antifermion pair of relative orbital angular momentum \(L\) and total spin \(S\), exchanging the two particles requires exchanging their positions, their spins, and their charge labels; the first costs \((-1)^{L}\), the second \((-1)^{S+1}\), the third is \(\mathcal{C}\) itself, and the total must be \(-1\) because the constituents are fermions. Hence
Para-positronium (\(L=0\), \(S=0\)) has \(C=+1\) and decays to two photons; ortho-positronium (\(L=0\), \(S=1\)) has \(C=-1\) and must decay to three. The observed rates differ by the extra factor of \(\alpha\) that the third photon costs, the singlet living \(125\,\mathrm{ps}\) and the triplet \(142\,\mathrm{ns}\), a ratio of about \(1100\) that no dynamical accident would supply.
Any amplitude in quantum electrodynamics consisting of a single closed fermion loop with an odd number of external photons vanishes identically [Furry:1937]. Rests on Theorem 100.10 and Equation (105.17).
This is proved as Theorem 100.10 by comparing the two orientations of the loop; the symmetry reading is one line. A closed fermion loop with \(n\) photons attached is a matrix element of an operator that is \(C\)-even (the vacuum is \(C\)-even, and so is the interaction), while the \(n\)-photon state has \(C=(-1)^{n}\); for odd \(n\) the two disagree and the amplitude must vanish. The practical value is large: Furry's theorem removes whole classes of diagrams before they are computed, and it is the reason light-by-light scattering starts at four photons and not at three.
G-parity
Charge conjugation is useless for a charged pion, since \(\mathcal{C}\ket{\pi^{+}}=\ket{\pi^{-}}\). But the strong interaction of Quantum Chromodynamics conserves isospin, and a rotation by \(\pi\) about the \(2\)-axis in isospin space maps \(\pi^{\pm}\) into \(\pi^{\mp}\) as well. Composing the two gives an operation under which the whole pion triplet is an eigenstate. Isospin is a dimensionless internal label and not an angular momentum, so its generators carry no \(\hbar\) and the rotation angle appears alone in the exponent:
For a neutral, non-strange multiplet of isospin \(I\) one has
because the isospin rotation acts on the \(I_{3}=0\) member of an isospin-\(I\) multiplet with eigenvalue \((-1)^{I}\), exactly as an ordinary rotation by \(\pi\) acts on \(Y_{I0}\). For the pion, \(\eta_{C}(\pi^{0})=+1\) and \(I=1\), so
\(G\)-parity is conserved by the strong interaction, and the selection rule is immediately useful: the \(\rho\) meson (\(I=1\), \(\eta_{C}=-1\), hence \(\eta_{G}=+1\)) decays strongly to two pions and not to three, while the \(\omega\) (\(I=0\), \(\eta_{C}=-1\), hence \(\eta_{G}=-1\)) decays to three and not to two. Both patterns are observed, and neither follows from energy or angular momentum. It fails, as it must, for the electromagnetic and weak interactions, which do not conserve isospin.
Time reversal
Why the operator must be antiunitary
Classically, motion reversal sends \(t\to-t\), leaves positions alone and reverses velocities; Newton's equation with a velocity-independent force is invariant. The quantum-mechanical implementation of the same idea is constrained far more tightly, and the constraint is Wigner's [Wigner:1932].
Let \(\mathcal{T}\) implement motion reversal, in the sense that
on a system with a Hamiltonian bounded below. Then \(\mathcal{T}\) is antiunitary. Rests on Theorem 94.3.
Derives Theorem 105.10. By Theorem 94.3 the operator is unitary and linear or antiunitary and antilinear; the argument is to exclude the first.
Two independent exclusions are available, and both are worth having. The first uses nothing beyond the canonical commutator of ordinary quantum mechanics, \(\comm{\hat{x}_{j}}{\hat{p}_{k}}=\ii\hbar\delta_{jk}\identity\). Conjugating with \(\mathcal{T}\) and using Equation (105.22),
If \(\mathcal{T}\) were linear the left-hand side would instead be \(\mathcal{T}(\ii\hbar\delta_{jk}\identity)\mathcal{T}^{-1} =+\ii\hbar\delta_{jk}\identity\), since a linear operator commutes with multiplication by a number. The two are inconsistent, so \(\mathcal{T}\) is antilinear: it satisfies \(\mathcal{T}(\ii\hbar)\mathcal{T}^{-1}=-\ii\hbar\), which reconciles them.
The second exclusion is the physically decisive one, because it invokes the spectrum rather than a commutator. Motion reversal must satisfy
since propagating forward for a time \(t\) and then reversing must equal reversing and then propagating backward. If \(\mathcal{T}\) were linear, the left-hand side would be \(\exp(-\ii(\mathcal{T}H\mathcal{T}^{-1})t/\hbar)\) and Equation (105.23) would force \(\mathcal{T}H\mathcal{T}^{-1}=-H\). Then for every eigenvalue \(E\) of \(H\) the value \(-E\) would also be an eigenvalue, the spectrum would be unbounded below, and there would be no stable ground state. If instead \(\mathcal{T}\) is antilinear, the \(\ii\) in the exponent changes sign as it is pulled through, the left-hand side is \(\exp(+\ii(\mathcal{T}H\mathcal{T}^{-1})t/\hbar)\), and Equation (105.23) gives
with the spectrum intact. Antiunitarity then follows from Theorem 94.3.
∎The same argument applied to space inversion produces no obstruction: \(\mathcal{P}\) leaves \(t\) alone, so Equation (105.23) is replaced by \(\mathcal{P}\ee^{-\ii Ht/\hbar}\mathcal{P}^{-1} =\ee^{-\ii Ht/\hbar}\), which a linear operator satisfies with \(\mathcal{P}H\mathcal{P}^{-1}=H\). The asymmetry between \(P\) and \(T\) in this chapter — one has a conserved quantum number and selection rules, the other has neither — traces back to precisely this line, and not to any experimental fact.
Let \(\mathcal{T}\) be antiunitary with \(\mathcal{T}\ket{a}=\eta_{T}\ket{a}\) for a nondegenerate state \(\ket{a}\). Then \(\eta_{T}\) carries no physical information: it can be set to \(1\) by a choice of phase for \(\ket{a}\). Rests on Theorem 105.10.
Derives Proposition 105.12. Replace \(\ket{a}\) by \(\ket{a'}=\ee^{\ii\theta}\ket{a}\). Antilinearity gives \(\mathcal{T}\ket{a'}=\ee^{-\ii\theta}\eta_{T}\ket{a} =\ee^{-2\ii\theta}\eta_{T}\ket{a'}\), so the eigenvalue is \(\eta_{T}\ee^{-2\ii\theta}\) and \(\theta=\tfrac12\arg\eta_{T}\) makes it unity.
∎There is a deeper way to say the same thing. A conserved quantum number comes from a Hermitian generator, by Noether's theorem or by the spectral theorem; an antiunitary operator has no Hermitian generator, because it is not part of any continuous one-parameter group. What \(T\) invariance gives instead is a relation between amplitudes, derived in Equation (105.25) below, and one genuine degeneracy theorem.
Kramers degeneracy
Let \(H\) be \(T\)-invariant, \(\mathcal{T}H\mathcal{T}^{-1}=H\), on a system of half-integer total angular momentum, for which \(\mathcal{T}^{2}=-\identity\). Then every energy eigenvalue is at least twofold degenerate, and no perturbation that is itself \(T\)-invariant can lift the degeneracy [Kramers:1930]. Rests on Theorem 105.10 and Equation (105.24).
Derives Theorem 105.13. On a system of angular momentum \(j\) the time-reversal operator may be written \(\mathcal{T}=\ee^{-\ii\pi \hat{J}_{y}/\hbar}K\) with \(K\) complex conjugation in the \(\hat{J}_{z}\) eigenbasis, since this is the unique antiunitary satisfying \(\mathcal{T}\hat{\vect{J}}\mathcal{T}^{-1} =-\hat{\vect{J}}\). Squaring, the \(K\)'s cancel and the rotation doubles: \(\mathcal{T}^{2}=\ee^{-2\ii\pi \hat{J}_{y}/\hbar}=(-1)^{2j}\identity\), which is \(-\identity\) for half-integer \(j\) (Remark 94.6 is the same sign seen in a neutron interferometer).
Let \(H\ket{\psi}=E\ket{\psi}\). Then \(H\mathcal{T}\ket{\psi}=\mathcal{T}H\ket{\psi}=E\mathcal{T}\ket{\psi}\) — note that \(E\) is real, so antilinearity does not disturb it — and \(\mathcal{T}\ket{\psi}\) has the same energy. It remains to show that it is a different state. Antiunitarity means \(\braket{\mathcal{T}\varphi}{\mathcal{T}\psi}=\braket{\psi}{\varphi}\) for all \(\varphi,\psi\). Take \(\varphi=\mathcal{T}\psi\):
The left side is \(\braket{-\psi}{\mathcal{T}\psi} =-\braket{\psi}{\mathcal{T}\psi}\), so \(\braket{\psi}{\mathcal{T}\psi}=0\): the two states are orthogonal. Finally, if \(V\) is a \(T\)-invariant perturbation, the same argument applies to \(H+V\) at every order, so the pair stays degenerate.
∎Kramers's own context was the paramagnetic rotation of light in crystals: an ion with an odd number of electrons has half-integer total angular momentum, and no electrostatic crystal field — however low its symmetry — can split its lowest level into singlets. Only a magnetic field can, because \(\vect{B}\) is \(T\)-odd and breaks the hypothesis. The theorem thus explains, without any detailed calculation, why odd-electron and even-electron ions behave completely differently in a crystal, and it remains the standard tool for classifying magnetic ions.
Motion reversal is not the arrow of time
It is worth stating plainly, because the confusion is common, that the \(T\) of this chapter has nothing to do with the second law of Classical Thermodynamics. The microscopic laws of electromagnetism, of the strong interaction and of gravity are exactly \(T\)-invariant, and the weak interaction violates \(T\) by a few parts in a thousand; yet entropy increases by factors of \(10^{20}\) in ordinary processes. The second law is a statement about the overwhelming majority of initial conditions compatible with a given macroscopic description, together with the coarse-graining that defines that description. It would hold with the same force in a world where \(T\) were an exact symmetry, and it is not measurably strengthened by the \(10^{-3}\) violation that the kaon exhibits. A theory of the arrow of time must explain a low-entropy initial condition; no amount of microscopic \(T\) violation does that.
Reciprocity, detailed balance, and why they prove less than they seem to
If \(\mathcal{T}H\mathcal{T}^{-1}=H\), the \(S\)-matrix satisfies
where \(\mathcal{T}i\) denotes the state \(i\) with all momenta and spins reversed. Rests on Theorem 105.10 and Equation (105.23).
Derives Proposition 105.14. First an identity used repeatedly below. For an antiunitary \(\mathcal{A}\) and any operator \(O\),
which follows by writing \(O\mathcal{A}\ket{\psi} =\mathcal{A}(\mathcal{A}^{-1}O\mathcal{A})\ket{\psi}\) and applying \(\braket{\mathcal{A}\varphi}{\mathcal{A}\chi}=\braket{\chi}{\varphi}\).
Now \(T\) invariance of \(H\) implies, by Equation (105.23) extended to the interacting evolution operator, that \(\mathcal{T}S\mathcal{T}^{-1}=S^{\dagger}\): reversing the motion turns the scattering operator into its inverse, which unitarity identifies with \(S^{\dagger}\). Substituting \(O=S\), \(\mathcal{A}=\mathcal{T}\), \(\varphi=i\), \(\psi=f\) into Equation (105.26) and using \(\mathcal{T}^{-1}S^{\dagger}\mathcal{T}=S\) gives \(\bra{\mathcal{T}i}S\ket{\mathcal{T}f} =\bra{f}S\ket{i}^{**}=\bra{f}S\ket{i}\).
∎Equation (105.25) is the correct statement of “detailed balance”: the amplitude for \(i\to f\) equals that for the reversed process \(\mathcal{T}f\to\mathcal{T}i\), with every momentum and every spin flipped. Two cautions attach to using it as a test.
First, at lowest order in a weak perturbation the test is empty. If \(S=\identity-\tfrac{\ii}{\hbar}\int\dd t\,H_{w}(t)\) to first order, then \(\abs{\bra{f}S\ket{i}}=\abs{\bra{i}S\ket{f}}\) follows from the hermiticity of \(H_{w}\) alone, with no reference to \(T\). A first-order process therefore cannot violate detailed balance whether or not \(T\) holds, and observing detailed balance in, say, a nuclear reaction and its inverse tests \(T\) only to the extent that higher orders matter.
Second, and more insidiously, a nonzero \(T\)-odd observable does not by itself establish \(T\) violation. The classic example is the triple correlation
in the decay of polarized neutrons, in which all three vectors reverse under motion reversal so that the product is odd. Yet a nonzero \(D\) can be generated by final-state interactions in a perfectly \(T\)-invariant theory. To see how, write an amplitude as a sum of two contributions carrying a \(CP\)-odd (“weak”) phase \(\phi\) and a \(CP\)-even (“strong”, or Coulomb) rescattering phase \(\delta\),
with \(a_{1,2}\) real. A \(T\)-odd correlation is proportional to \(\mathrm{Im}(\mathcal{M}_{1}\mathcal{M}_{2}^{*}) =a_{1}a_{2}\sin(\delta_{1}-\delta_{2}+\phi_{1}-\phi_{2})\), which is nonzero when \(\delta_{1}\ne\delta_{2}\) even if \(\phi_{1}=\phi_{2}\). In neutron decay the Coulomb interaction between the outgoing proton and electron supplies \(\delta_{1}-\delta_{2}\sim\alpha m_{e}c/p_{e}\), faking \(D\sim10^{-5}\); the measured value, \(D=(-1.2\pm2.0)\times10^{-4}\) [Navas:2024], is not yet sensitive to the fake, let alone to anything beyond it. The lesson generalizes: a genuine \(T\) test must either compare a process with its motion reverse (Section 105.4) or measure a static property of a nondegenerate state (Section 105.4.3), where no final state exists to rescatter.
Time reversal on the fields
On a Dirac field, antiunitarity plus the requirement that solutions map to solutions gives
with \(\mathrm{T}_{\mathrm{D}}=\gamma^{1}\gamma^{3}\) in the Dirac representation; the complex conjugate on \(\gamma^{\mu}\) is the signature of the antilinearity, exactly as in Theorem 105.10. For the gauge field the assignment is fixed by the classical behaviour of the sources: reversing all motions reverses currents but not charges, so
whence \(\vect{E}\to+\vect{E}\) and \(\vect{B}\to-\vect{B}\). That \(\vect{B}\) is \(T\)-odd and \(\vect{E}\) is \(T\)-even is used at three separate places below — in Kramers degeneracy, in the electric dipole moment argument, and in the sign of the magnetic moment under \(CPT\) — and it is worth remembering as a picture: a magnetic field is made by currents, and currents reverse.
Carrying Equation (105.29) through the bilinears as before completes Table 105.1.
| bilinear | $P$ | $C$ | $T$ | $\Theta=CPT$ |
|---|---|---|---|---|
| $S=\bar{\psi}\psi$ | $+S(x_{P})$ | $+S$ | $+S(x_{T})$ | $+S(-x)$ |
| $\tilde{S}=\bar{\psi}\gamma^{5}\psi$ | $-\tilde{S}(x_{P})$ | $+\tilde{S}$ | $-\tilde{S}(x_{T})$ | $+\tilde{S}(-x)$ |
| $V^{\mu}=\bar{\psi}\gamma^{\mu}\psi$ | $+V_{\mu}(x_{P})$ | $-V^{\mu}$ | $+V_{\mu}(x_{T})$ | $-V^{\mu}(-x)$ |
| $A^{\mu}=\bar{\psi}\gamma^{\mu}\gamma^{5}\psi$ | $-A_{\mu}(x_{P})$ | $+A^{\mu}$ | $+A_{\mu}(x_{T})$ | $-A^{\mu}(-x)$ |
| $T^{\mu\nu}=\bar{\psi}\sigma^{\mu\nu}\psi$ | $+T_{\mu\nu}(x_{P})$ | $-T^{\mu\nu}$ | $-T_{\mu\nu}(x_{T})$ | $+T^{\mu\nu}(-x)$ |
The last column will be derived and used in Section 105.5.1; it is written here because it costs nothing once the other three exist, and because seeing it early makes the shape of the \(CPT\) theorem obvious. Under \(\Theta=CPT\) a bilinear returns to itself, at the reflected point, with the sign \((-1)^{n}\) where \(n\) counts its Lorentz indices — and with no memory of whether it was scalar or pseudoscalar, vector or axial. That insensitivity is why the product survives when the factors do not.
Parity violation
The tau–theta puzzle and the Lee–Yang question
By 1955 two strange mesons were known that agreed in every property that could then be measured and disagreed in one that could not be explained. The \(\theta^{+}\) decayed to \(\pi^{+}\pi^{0}\); the \(\tau^{+}\) decayed to \(\pi^{+}\pi^{+}\pi^{-}\). Their masses agreed to better than a per cent and their lifetimes to better than the experimental errors, which for two unrelated particles would be an extraordinary coincidence. But the parities of the two final states are opposite.
A state of \(n\) pions with total orbital angular momentum content \(L_{\mathrm{tot}}\) (the sum of the relative orbital angular momenta) has parity
Rests on Phenomenon 105.5 and Equation (105.4).
Derives Proposition 105.15. Immediate from multiplicativity, the pseudoscalar assignment \(\eta_{P}(\pi)=-1\) of Phenomenon 105.5, and Equation (105.4) applied to each relative coordinate.
∎For the \(\theta^{+}\), the parent is spinless (its decay products are two spinless particles, and the angular distribution is isotropic), so \(L=0\) and \(\eta_{P}=(-1)^{2}=+1\). For the \(\tau^{+}\), if all relative angular momenta vanish then \(\eta_{P}=(-1)^{3}=-1\). Two particles of the same mass and the same lifetime were therefore decaying into final states of opposite parity — impossible if parity is conserved, unless the two are genuinely different particles that happen to be degenerate, or unless the \(\tau\) decay carries \(L_{\mathrm{tot}}=1\) and the parent has spin \(\ge1\).
The Dalitz plot
It was Dalitz's construction that closed the second escape [Dalitz:1953]. For a three-body decay of a particle at rest into three pions of kinetic energies \(T_{1},T_{2},T_{3}\) with \(T_{1}+T_{2}+T_{3}=Q\) fixed, plot the event as the point of an equilateral triangle of height \(Q\) whose perpendicular distances to the three sides are \(T_{1},T_{2},T_{3}\). The kinematically allowed region is a closed curve inside the triangle, and — this is the content of the construction — the density of points is uniform whenever the decay matrix element is constant.
Derivation. The three-body phase space for a parent of mass \(M\) decaying at rest is
Do the \(\vect{p}_{3}\) integral with the momentum delta function, write \(\dd^{3}p_{1}=p_{1}^{2}\dd p_{1}\dd\Omega_{1}\) and \(\dd^{3}p_{2}=p_{2}^{2}\dd p_{2}\,\dd\cos\theta_{12}\,\dd\varphi\), and use the energy delta function to fix \(\cos\theta_{12}\). Since \(E_{3}^{2}/c^{2}=\vect{p}_{1}^{2}+\vect{p}_{2}^{2} +2p_{1}p_{2}\cos\theta_{12}+m_{3}^{2}c^{2}\), one has \(\dd\cos\theta_{12}=E_{3}\dd E_{3}/(c^{2}p_{1}p_{2})\) at fixed \(p_{1},p_{2}\), and the three factors of \(E_{a}\) in the denominators cancel against \(p_{a}^{2}\dd p_{a}=E_{a}\dd E_{a}p_{a}/c^{2}\). Collecting,
up to the overall angular factors that integrate to a constant. Since \(T_{a}=E_{a}-m_{\pi}c^{2}\), the density in the \((T_{1},T_{2})\) plane — and hence in the Dalitz triangle, which is an affine image of it — is \(\abs{\mathcal{M}}^{2}\) and nothing else.
∎The plot therefore reads off the matrix element directly. If the decay proceeded with a unit of relative orbital angular momentum, the amplitude would vanish at the centre of the plot, where all three momenta are small, and the population there would be depleted; a \(J=1\) or \(J=2\) parent would leave a characteristic pattern of zeros along the boundary. The observed \(\tau\) plot was flat. The three pions were in a state of no relative orbital angular momentum, the parent had \(J=0\), and its parity was \(-1\): no escape.
Lee and Yang's question
Lee and Yang responded not with a mechanism but with an audit [Lee:1956]. They asked what the experimental evidence for parity conservation actually was, and found that all of it came from strong and electromagnetic processes — Laporte's rule, nuclear level schemes, the selection rules of Proposition 105.1. In these sectors the evidence was overwhelming, bounding any parity-violating admixture in the nuclear force at the level of \(10^{-7}\) in amplitude. In the weak interaction there was none at all: not one experiment had ever been designed to detect a parity-odd observable in a weak decay. The \(\theta\)–\(\tau\) puzzle was, on this reading, not a puzzle about two particles but a signal that the weak interaction does not conserve parity, and that \(\theta\) and \(\tau\) are one particle, the \(K^{+}\), decaying two ways.
Lee and Yang then did the thing that makes the paper a model of method: they listed the experiments that would settle it. Each of their proposals is the expectation value of a pseudoscalar, since a pseudoscalar is precisely a quantity that must vanish if \(P\) is a symmetry:
-
the correlation \(\avg{\vect{J}}\cdot\vect{p}_{e}\) between the polarization of an oriented nucleus and the momentum of the emitted electron — an axial vector dotted into a polar vector;
-
the longitudinal polarization \(\avg{\vect{s}}\cdot\vect{p}\) of a lepton emitted in a weak decay — the same structure;
-
in the chain \(\pi\to\mu\nu\), \(\mu\to e\nu\bar{\nu}\), the correlation between the muon direction and the electron direction, which is nonzero only if the muon is produced polarized.
All three were done within a year, and all three gave the maximum possible answer.
The experimental collapse of 1957
Nuclei of \(^{60}\)Co polarized in a magnetic field at a temperature of about \(0.01\,\mathrm{K}\) emit beta electrons preferentially opposite to the nuclear spin. The angular distribution is
with \(A_{\beta}\) consistent with \(-1\), and the asymmetry disappears as the sample warms and the polarization is lost [Wu:1957]. The apparatus is Experiment: Parity Violation; see Phenomenon 110.10. Rests on Equations (105.34) and (105.37).
Derivation. Derives Phenomenon 105.16. That \(A_{\beta}\ne0\) is impossible in a parity-conserving theory needs one line. \(\vect{J}\) is an axial vector and \(\vect{v}_{e}\) a polar vector, so their scalar product is a pseudoscalar and changes sign under \(\mathcal{P}\). If \(\comm{\mathcal{P}}{H}=0\) the rate is \(\mathcal{P}\)-invariant, so \(W(\theta)=W(\pi-\theta)\) and the coefficient of \(\cos\theta\) must vanish.
That the observed value is the maximum takes a little more. The transition \(^{60}\mathrm{Co}(5^{+})\to{}^{60}\mathrm{Ni}^{*}(4^{+})\) is allowed and pure Gamow–Teller: the leptons carry off one unit of angular momentum in their spins and none in orbital motion. If the emitted antineutrino is purely right-handed and the electron predominantly left-handed — the content of the \(V-A\) structure below — then conservation of angular momentum along the nuclear spin axis forces the electron to go backwards. Quantitatively, for a pure Gamow–Teller transition with \(\Delta J=-1\) the \(V-A\) prediction is \(A_{\beta}=-1\) exactly, and the factor \(v_{e}/c\) in Equation (105.33) is the longitudinal polarization of an electron of speed \(v_{e}\), derived at Equation (105.37).
∎Wu, Ambler, Hayward, Hoppes and Hudson polarized their source by adiabatic demagnetization of a cerium magnesium nitrate crystal into which a thin layer of \(^{60}\)Co had been grown, and monitored the polarization independently through the anisotropy of the gamma rays from the daughter. The two signals — gamma anisotropy and beta asymmetry — decayed together over some six minutes as the crystal warmed, which is the control that makes the result an observation of parity violation rather than of an instrumental drift.
Two further experiments were reported in the same issue of the same journal. Garwin, Lederman and Weinrich [Garwin:1957] and Friedman and Telegdi [Friedman:1957] used the chain \(\pi^{+}\to\mu^{+}\nu_{\mu}\) followed by \(\mu^{+}\to e^{+}\nu_{e}\bar{\nu}_{\mu}\). If parity is conserved the muon from a spinless pion cannot be longitudinally polarized, since \(\avg{\vect{s}_{\mu}}\cdot\vect{p}_{\mu}\) is a pseudoscalar; and even if it were, the electron from its decay could not remember the direction. Both correlations were found: the positrons came out preferentially along the muon spin, with an asymmetry that also measured the muon magnetic moment to be very close to the Dirac value \(g=2\) (the quantity refined to twelve digits in Experiment: The Electron Anomalous Magnetic Moment). Phenomenon 110.13 records the modern statement.
The neutrino emitted in the electron capture \(^{152}\mathrm{Eu}^{m}+e^{-}\to{}^{152}\mathrm{Sm}^{*}+\nu_{e}\) has helicity \(-1\) within the experimental accuracy [Goldhaber:1958]. The apparatus, which used resonant scattering of the daughter gamma ray to select neutrinos emitted backwards, is Experiment: Neutrino Helicity; see Phenomenon 98.2. Rests on Equations (105.34) and (105.37).
Derivation. Derives Phenomenon 105.17. The charged-current interaction Equation (105.34) contains the neutrino field only through \(\left(1-\gamma^{5}\right)\psi_{\nu} =2P_{L}\psi_{\nu}\), so a neutrino leaving a weak vertex is emitted in a state of definite chirality — the left-handed one — whatever the nucleus does. Chirality is not helicity for a massive particle, and the exact relation is Equation (105.37): the longitudinal polarization of a fermion emitted by a \(V-A\) current is \(-v/c\), so helicity \(-1\) is predicted only up to
The neutrino of this decay carries \(E\approx0.90\,\mathrm{MeV}\), and the laboratory bound on the neutrino mass is of order \(1\,\mathrm{eV}/c^{2}\) (Flavour Physics and Neutrinos), so the correction is at most \(6\times10^{-13}\). The prediction is therefore helicity \(-1\) to twelve significant figures, and the measurement — which resolves the sign and roughly its first digit — tests the chiral structure of Equation (105.34) and nothing about the neutrino mass. What the experiment adds to the theory is the sign: \(V-A\) with the opposite relative sign, \(V+A\), would predict \(+1\) just as sharply.
∎Goldhaber, Grodzins and Sunyar's measurement is the sharpest of the three, because helicity is not merely a nonzero pseudoscalar but a saturated one: \(\avg{\vect{s}}\cdot\hat{\vect{p}}=-\hbar/2\) is as large as it can be. Parity is not slightly violated in the weak interaction; it is violated maximally.
The V–A structure
The theoretical form that accommodates all of this was proposed by Feynman and Gell-Mann [Feynman:1958] and independently by Sudarshan and Marshak [Sudarshan:1958], with the two-component neutrino of Landau [Landau:1957], Lee and Yang [Lee:1957] and Salam [Salam:1957] as its kinematic core. The charged-current weak interaction at low energy is a product of two currents, each of the form \(V-A\):
with \(g_{A}/g_{V}=-1.2754\pm0.0013\) for the nucleon [Navas:2024] — the deviation from \(-1\) being a strong-interaction effect, since the quark-level current is pure \(V-A\).
Fermi's constant [Fermi:1934] is quoted throughout the literature in the natural-unit form
which is a statement about \(G_{F}/(\hbar c)^{3}\) and not about \(G_{F}\). The constant itself, as it appears in Equation (105.34) where the four fields carry \(\mathrm{m}^{-6}\) between them and \(\Ham_{w}\) is an energy density, has the SI value
obtained from Equation (105.35) by multiplying by \((\hbar c)^{3}=(3.16152677\times 10^{-26}\,\mathrm{J}\,\mathrm{m})^{3}\) and converting \(\mathrm{GeV}^{2}\) to \(\mathrm{J}^{2}\). Both numbers describe the same physics; the first is the one to compare with a textbook and the second is the one that makes Equation (105.34) dimensionally correct. The smallness of the weak interaction at low energy is the smallness of Equation (105.36) measured against \(\hbar c\) and a length: at a nuclear scale \(r\sim1\,\mathrm{fm}\) the ratio \(G_{F}/(\hbar c\,r^{2})\approx10^{-6}\), which is the familiar suppression.
Under \(\mathcal{P}\) the chiral projector \(P_{L}=\tfrac12(1-\gamma^{5})\) is mapped to \(P_{R}=\tfrac12(1+\gamma^{5})\), so the parity image of Equation (105.34) is an interaction with no overlap with the original. The parity-violating part of the leptonic current is as large as the parity-conserving part. Rests on Equations (105.7), (105.11) and (105.34).
Derives Proposition 105.19. From Equation (105.7), \(\mathcal{P}\left[\bar{\psi}_{1} \gamma^{\mu}P_{L}\psi_{2}\right]\mathcal{P}^{-1} =\bar{\psi}_{1}\gamma^{0}\gamma^{\mu} \tfrac12(1-\gamma^{5})\gamma^{0}\psi_{2}\) evaluated at \(x_{P}\). Because \(\gamma^{5}\) anticommutes with \(\gamma^{0}\), moving the \(\gamma^{0}\) through gives \(\gamma^{0}\gamma^{5}\gamma^{0}=-\gamma^{5}\) and hence \(P_{L}\to P_{R}\); the index is lowered as in Equation (105.11). The transformed interaction couples right-handed leptons and is orthogonal, in the space of possible couplings, to the one we started from. Writing \(V-A\) as \(V\) and \(A\) pieces of equal magnitude makes the same point: parity violation is proportional to the \(V\)–\(A\) interference term, which for coefficients of equal magnitude is as large as either term alone.
∎The physical content of the projector is helicity.
Let \(u(p,s)\) solve the free Dirac equation with \(\vect{p}\) along \(\hat{\vect{z}}\). Then
so a \(V-A\) current emits a fermion with longitudinal polarization \(-v/c\) and an antifermion with \(+v/c\); in the ultrarelativistic limit the helicity is \(\mp1\) exactly, and for a strictly massless fermion chirality and helicity coincide. Rests on Equations (100.5) and (105.34).
Derives Proposition 105.20. The helicity operator is \(h=\vect{\Sigma}\cdot\hat{\vect{p}}\) with \(\vect{\Sigma}=\diag(\vect{\sigma},\vect{\sigma})\). The Dirac equation \((\gamma^{\mu}p_{\mu}-mc)u=0\) gives, in the Dirac representation with \(\vect{p}=p\hat{\vect{z}}\), the lower two components in terms of the upper: \(u=N\bigl(\chi,\; c\,\vect{\sigma}\cdot\vect{p}\,\chi/(E+mc^{2})\bigr)\) with \(N^{2}=(E+mc^{2})/(2mc^{2})\) fixing \(\bar{u}u=1\). Take \(\chi\) the eigenvector of \(\sigma_{z}\) with eigenvalue \(+1\). Then \(\gamma^{5}u\), which exchanges the upper and lower two-spinors, gives
the factor \(2\) being the two equal cross terms between the upper and lower two-spinors. It is \(u^{\dagger}\gamma^{5}u\) and not the bilinear \(\bar{u}\gamma^{5}\gamma^{0}u\) that is wanted, the two differing by a sign: \(\gamma^{0}\gamma^{5}\gamma^{0}=-\gamma^{5}\), so \(\bar{u}\gamma^{5}\gamma^{0}u=-u^{\dagger}\gamma^{5}u\). Likewise \(u^{\dagger}u=N^{2}\left(1+c^{2}p^{2}/(E+mc^{2})^{2}\right) =E/(mc^{2})\), the bracket collapsing to \(2E/(E+mc^{2})\) because \(c^{2}p^{2}=(E-mc^{2})(E+mc^{2})\). The ratio is \(cp/E=v/c\), and since \(\avg{P_{L}}-\avg{P_{R}}=-\avg{\gamma^{5}}\) normalized to \(u^{\dagger}u\), Equation (105.37) follows. As \(m\to0\), \(E\to cp\) and the ratio tends to \(1\): \(P_{L}\) then projects exactly onto helicity \(-1\).
∎Equation (105.37) is the factor \(v_{e}/c\) in Equation (105.33), so the shape of Wu's angular distribution, the polarization of the beta electron, and the helicity of the neutrino are three faces of a single structure. The neutrino, being (very nearly) massless, saturates it. The theory of the current itself — why \(V-A\), and how it arises from an \(\SU(2)\) gauge interaction with a left-handed doublet — belongs to Weak Interactions and Electroweak Unification and the Higgs Boson.
Parity violation as a precision tool
Once parity violation is established it stops being a discovery and becomes an instrument, for a reason worth stating carefully. A weak amplitude is far smaller than an electromagnetic one at ordinary energies, so a weak rate is invisible under the electromagnetic background. But a parity-odd observable receives no contribution at all from the electromagnetic amplitude squared: its leading term is the interference between the weak and electromagnetic amplitudes, which is linear in the small quantity. Measuring a parity-violating asymmetry therefore buys a factor \(\mathcal{M}_{\gamma}/\mathcal{M}_{Z}\) of sensitivity relative to measuring a weak rate.
For electron scattering at squared momentum transfer \(Q^{2}\), the ratio of the neutral-current to the photon-exchange amplitude is
so the parity-violating asymmetry \(A_{PV}=(\sigma_{R}-\sigma_{L})/(\sigma_{R}+\sigma_{L})\) is of order \(10^{-4}\) at \(Q^{2}c^{2}=1\,\mathrm{GeV}^{2}\) and of order \(10^{-7}\) at the \(Q^{2}\) of a low-energy fixed-target experiment. Rests on Equations (105.34) and (105.35).
Derives Proposition 105.21. The photon-exchange amplitude carries \(4\pi\alpha\hbar c/Q^{2}\) from the propagator and coupling; the \(Z\)-exchange amplitude at \(Q^{2}\ll m_{Z}^{2}c^{2}\) is contact-like and carries \(G_{F}/(\sqrt{2}(\hbar c)^{3})\) in the same normalization. The ratio is Equation (105.38), and inserting \(G_{F}/(\hbar c)^{3} =1.1664\times 10^{-5}\,/\mathrm{GeV}^{2}\) together with \(\alpha=1/137.036\) gives \(1.1664\times10^{-5}/(4\sqrt{2}\pi\times7.2974\times10^{-3}) =8.99\times10^{-5}\) per \(\mathrm{GeV}^{2}\).
∎Deep inelastic scattering: Prescott 1978
Prescott and collaborators scattered longitudinally polarized electrons of \(16.2\text{–}22.2\,\mathrm{GeV}\) from deuterium at SLAC and reversed the beam helicity at random, pulse to pulse [Prescott:1978]. They measured
in magnitude exactly the estimate of Equation (105.38). The electroweak analysis writes
with \(a_{1}\propto1-\tfrac{20}{9}\sin^{2}\theta_{W}\) and \(a_{2}\propto1-4\sin^{2}\theta_{W}\). The second coefficient nearly vanishes near \(\sin^{2}\theta_{W}\approx0.23\), so it is the \(y\)-dependence of the asymmetry — the fact that it is nearly flat — that fixes the weak mixing angle rather than its magnitude. Prescott obtained \(\sin^{2}\theta_{W}=0.20\pm0.03\), in agreement with the value inferred from neutrino scattering, and thereby confirmed the \(\SU(2)\times\U(1)\) structure of Electroweak Unification and the Higgs Boson against the alternatives then in play. See Phenomenon 110.17.
Atomic parity violation
In an atom, the exchange of a \(Z\) between an electron and the nucleus adds to the Hamiltonian a term that is parity-odd, mixing \(S\) and \(P\) states and giving a normally forbidden \(E1\) transition a small amplitude. The nuclear-spin-independent part is governed by the weak charge
and the induced \(E1\) amplitude grows roughly as \(Z^{3}\): one power from \(Q_{W}\approx-N\propto Z\), and two from the relativistic enhancement of the electron density at the nucleus. Heavy atoms are therefore enormously favoured, and caesium (\(Z=55\)) is the practical optimum because its single valence electron makes the atomic structure calculable to a per cent.
Wood and collaborators measured the parity-violating \(6S\to7S\) amplitude in caesium and extracted [Wood:1997]
the two errors being experimental and atomic-theoretical. This is a measurement of the weak interaction made in a tabletop apparatus at an energy of a few electronvolts, and it agrees with the value obtained at LEP at \(91\,\mathrm{GeV}\) — the same weak mixing angle at two energies about ten decades apart, since \(91\,\mathrm{GeV}\) is \(3\times10^{10}\) times a few electronvolts. The agreement is not automatic: \(\sin^{2}\theta_{W}\) runs with the scale at which it is probed, and the two determinations agree only once that running is applied. See Phenomenon 110.16.
The proton's weak charge, and a nuclear skin
Two modern measurements exploit the near-cancellation in Equation (105.41). For a proton, \(Z=1\) and \(N=0\), so \(Q_{W}^{p}=1-4\sin^{2}\theta_{W}\approx0.0719\): a small number, and a fractional measurement of it is a much finer measurement of \(\sin^{2}\theta_{W}\). Differentiating, \(\delta(\sin^{2}\theta_{W})=\tfrac14 Q_{W}^{p}\, \delta Q_{W}^{p}/Q_{W}^{p}\), so a \(6\,\mathrm{\%}\) determination of \(Q_{W}^{p}\) gives \(\delta\sin^{2}\theta_{W}\approx0.001\). The Qweak experiment measured an asymmetry of \(-226.5\pm9.3\) parts per billion at \(Q^{2}c^{2}=0.0248\,\mathrm{GeV}^{2}\) and obtained [Androic:2018]
in agreement with the Standard Model. See Phenomenon 110.19.
For a heavy nucleus the same cancellation acts the other way. Since \(Q_{W}^{p}\approx0.07\) while \(Q_{W}^{n}=-1\), the \(Z\) couples almost exclusively to neutrons, and the parity-violating asymmetry in elastic electron scattering measures the neutron distribution while ordinary electron scattering measures the proton distribution. The difference of the two radii is the neutron skin. PREX-2 measured \(A_{PV}=550\pm16\pm8\) parts per billion on \(^{208}\)Pb and obtained [Adhikari:2021]
The skin thickness is controlled by the density dependence of the nuclear symmetry energy, which is the same quantity that sets the radius of a neutron star, so a parity-violating asymmetry measured on a lead target in a Virginia accelerator hall constrains the equation of state used in Compact Stars and Relativistic Astrophysics. It is a fair example of what a symmetry-violating observable buys: an otherwise inaccessible distribution, extracted cleanly because the electromagnetic background cannot contribute to it.
CP violation
The neutral kaon system
Landau's proposal, made immediately after the collapse of parity, was that the true symmetry is the product \(CP\) [Landau:1957]. It is an attractive suggestion: \(C\) exchanges a left-handed neutrino for a left-handed antineutrino, which does not exist; \(P\) exchanges it for a right-handed neutrino, which does not exist either; but \(CP\) exchanges it for a right-handed antineutrino, which does. Maximal violation of \(P\) and of \(C\) separately is exactly compatible with exact \(CP\), and for seven years it appeared to be the truth.
The system in which it failed is the one Gell-Mann and Pais had analysed for a different purpose in 1955 [GellMann:1955], and it remains the most instructive two-state problem in physics.
Two states, two bases
The strong interaction produces states of definite strangeness. Write \(\ket{K^{0}}\) for the state of strangeness \(S=+1\) (quark content \(d\bar{s}\)) and \(\ket{\bar{K}^{0}}\) for \(S=-1\) (\(\bar{d}s\)); they are distinct particles, since strangeness is conserved by the strong interaction that makes them, and each is the antiparticle of the other. Fix phases by
which is consistent because \((\mathcal{CP})^{2}=\identity\) on these states. (The opposite sign convention, \(\mathcal{CP}\ket{K^{0}}=-\ket{\bar{K}^{0}}\), is equally common and changes nothing physical; it interchanges the roles of the two combinations below.) The \(CP\) eigenstates are then
with \(CP=+1\) and \(CP=-1\) respectively.
The weak interaction, which destroys them, does not conserve strangeness, so it is \(\ket{K_{1}}\) and \(\ket{K_{2}}\) — not \(\ket{K^{0}}\) and \(\ket{\bar{K}^{0}}\) — that have definite decay properties if \(CP\) holds. And the two available final states have opposite \(CP\):
For \(\pi\pi\) from a spinless parent, \(L=0\), so \(\eta_{P}=(-1)^{2}(-1)^{0}=+1\) by Equation (105.31) and \(\eta_{C}=+1\) for \(\pi^{0}\pi^{0}\) or, by Equation (105.18) applied to the charge-conjugate pair, for \(\pi^{+}\pi^{-}\) in an \(L=0\) state. For \(3\pi^{0}\), \(\eta_{P}=(-1)^{3}=-1\) and \(\eta_{C}=+1\). Hence if \(CP\) is exact, \(K_{1}\to\pi\pi\) is allowed and \(K_{2}\to\pi\pi\) is forbidden, and \(K_{2}\) must decay to three pions or semileptonically.
That is a strong prediction, because three pions barely fit: with \(m_{K}c^{2}=497.611(13)\,\mathrm{MeV}\) [Navas:2024] and \(3m_{\pi}c^{2}\approx405\,\mathrm{MeV}\), the released kinetic energy is only about \(90\,\mathrm{MeV}\), and the resulting phase-space suppression must make \(K_{2}\) long-lived. Gell-Mann and Pais predicted exactly this: a second neutral kaon, of the same mass to within the weak interaction's accuracy, with a lifetime orders of magnitude longer. Lande, Booth, Impeduglia, Lederman and Chinowsky found it at the Brookhaven Cosmotron [Lande:1956]. The measured lifetimes are
a ratio of \(571\) [Navas:2024].
The effective Hamiltonian
To do better than the symmetry argument one needs the dynamics of an unstable two-state system, and the only honest way to obtain it is to start from the full unitary problem and eliminate what is not being watched. This is the Wigner–Weisskopf construction, and it is worth carrying out because it is also the prototype of the open-system formalism of Open Quantum Systems and Decoherence and because the non-Hermitian matrix that comes out of it is otherwise mysterious.
Let the Hilbert space be spanned by \(\ket{K^{0}},\ket{\bar{K}^{0}}\) and a continuum of decay states \(\ket{n}\) of energy \(E_{n}\), with a weak perturbation \(H_{w}\) connecting the two sectors. On time scales long compared with the memory time of the decay channels — of order \(\hbar\) divided by the width of the accessible continuum — the amplitudes \(a(t),b(t)\) of the two kaon states obey
where \(\mathbf{M}\) and \(\vect{\Gamma}\) are Hermitian \(2\times2\) matrices,
\(\mathrm{P}\) denoting the principal value. \(\mathbf{M}\) has the SI dimension of energy and \(\vect{\Gamma}\) of inverse time. Rests on Proposition 17.42.
Derives Proposition 105.22. Write \(\ket{\psi(t)}=a\ket{K^{0}}+b\ket{\bar{K}^{0}} +\sum_{n}c_{n}\ket{n}\) and project the Schrödinger equation \(\ii\hbar\pp_{t}\ket{\psi}=(H_{0}+H_{w})\ket{\psi}\) onto each sector. For the continuum,
with \(\psi_{1}=a\), \(\psi_{2}=b\). Integrate with \(c_{n}(0)=0\):
Substituting into the equation for the two-dimensional sector gives an exact integro-differential equation with a memory kernel. The Wigner–Weisskopf approximation is to note that the kernel is sharply peaked in \(t-t'\) on the scale \(\hbar/\Delta E\) set by the width of the continuum, which for a hadronic decay is far shorter than the kaon lifetime, so that \(\psi_{j}(t')\) may be replaced by \(\psi_{j}(t)\ee^{\ii m_{K}c^{2}(t-t')/\hbar}\) and the upper limit sent to infinity. The remaining integral is
using the Sokhotski–Plemelj identity, Proposition 17.42. Collecting the real and imaginary parts reproduces Equations (105.50) and (105.51). That \(\mathbf{M}\) and \(\vect{\Gamma}\) are separately Hermitian is manifest: each is built from \(\bra{i}H_{w}\ket{n}\overline{\bra{j}H_{w}\ket{n}}\) with a real weight, since \(H_{w}\) is Hermitian.
∎Two structural consequences follow at once. First, \(\vect{\Gamma}\) is positive semidefinite, being of the form \(\sum_{n}v_{n}v_{n}^{\dagger}\) with \(v_{n,i}=\bra{i}H_{w}\ket{n}\) weighted by a positive delta function; hence
and probability leaks out of the two-dimensional space at exactly the rate at which decay products appear. Second, \(\mathbf{H}\) is not Hermitian, so its eigenvectors need not be orthogonal — and in the presence of \(CP\) violation they are not. This is not a pathology; it is what an open subsystem looks like.
What the discrete symmetries impose
With the phase convention Equation (105.45):
-
\(CPT\) invariance implies \(H_{11}=H_{22}\), that is \(M_{11}=M_{22}\) and \(\Gamma_{11}=\Gamma_{22}\);
-
\(T\) invariance implies \(\mathrm{Im}\left(M_{12} \hbar\Gamma_{12}^{*}\right)=0\), equivalently \(\abs{H_{12}}=\abs{H_{21}}\);
-
\(CP\) invariance implies both.
Rests on Equation (105.45), Equation (105.26) and Proposition 105.22.
Derives Proposition 105.23. (i) Let \(\Theta\) be the antiunitary \(CPT\) operator, so that \(\Theta\ket{K^{0}}=\ket{\bar{K}^{0}}\) up to a phase and \(\Theta H_{w}\Theta^{-1}=H_{w}\), \(\Theta H_{0}\Theta^{-1}=H_{0}\). Apply Equation (105.26) to \(M_{11}=\bra{K^{0}}\mathbf{M}\ket{K^{0}}\): since \(\mathbf{M}\) is built from \(H_{w}\), \(H_{0}\) and real weights, \(\Theta\mathbf{M}\Theta^{-1} =\mathbf{M}\), and Equation (105.26) gives \(\bra{\bar{K}^{0}}\mathbf{M}\ket{\bar{K}^{0}} =\bra{K^{0}}\mathbf{M}\ket{K^{0}}^{*}\). But \(\mathbf{M}\) is Hermitian, so its diagonal elements are real, and \(M_{22}=M_{11}\). The same argument on \(\vect{\Gamma}\) gives \(\Gamma_{22}=\Gamma_{11}\). Note that \(\Theta\) maps the state to its antiparticle without reversing momentum (the kaon is at rest), which is exactly what makes this constraint a relation between the two diagonal entries.
(ii) \(T\) invariance means \(\mathcal{T}H_{w}\mathcal{T}^{-1}=H_{w}\) with \(\mathcal{T}\) antiunitary and, for kaons at rest, \(\mathcal{T}\ket{K^{0}}=\ee^{\ii\alpha}\ket{K^{0}}\). Applying Equation (105.26) to the off-diagonal element gives \(M_{12}=\ee^{-2\ii\alpha}M_{12}^{*}\) and likewise \(\Gamma_{12}=\ee^{-2\ii\alpha}\Gamma_{12}^{*}\): both off-diagonal elements carry the same phase \(\alpha\), which the phase freedom of \(\ket{K^{0}}\) can then remove. The convention-independent content is that their relative phase vanishes, \(\mathrm{Im}(M_{12}\hbar\Gamma_{12}^{*})=0\), whereupon \(H_{21}=H_{12}^{*}\) with \(H_{12}\) of a removable phase, so \(\abs{H_{12}}=\abs{H_{21}}\).
(iii) \(CP=T^{-1}\cdot CPT\) up to the phases already fixed, so it imposes the union of (i) and (ii).
∎The three constraints are logically independent, and that is the architecture of the whole subject: the diagonal of \(\mathbf{H}\) tests \(CPT\), the moduli of its off-diagonal elements test \(T\), and \(CP\) tests both at once. Experiment finds (i) satisfied to the accuracy of Section 105.5.3 and (ii) violated at the \(10^{-3}\) level.
Diagonalization, and the definition of epsilon
Assume \(CPT\), so \(H_{11}=H_{22}=:H_{0}\) and
The characteristic equation gives eigenvalues \(\lambda_{\pm}=H_{0}\pm\sqrt{H_{12}H_{21}}\) with right eigenvectors proportional to \((\sqrt{H_{12}},\pm\sqrt{H_{21}})\). Define the complex number \(\epsilon\) by
so that the two physical states are
with eigenvalues written as
Equation (105.54) is the whole of \(CP\) violation in the mixing: \(\epsilon=0\) if and only if \(H_{12}=H_{21}\), which by Proposition 105.23 holds if and only if \(T\) (and, given \(CPT\), \(CP\)) is a symmetry of the mass matrix. It is also manifest that \(\ket{K_{S}}\) and \(\ket{K_{L}}\) are not orthogonal: \(\braket{K_{S}}{K_{L}}=2\mathrm{Re}\,\epsilon/(1+\abs{\epsilon}^{2})\ne0\).
To first order in the small asymmetry, write Equation (105.54) as \(\epsilon=(H_{12}-H_{21})/[(H_{12}+H_{21})+2\sqrt{H_{12}H_{21}}]\) and use \(H_{12}+H_{21}\approx2\sqrt{H_{12}H_{21}}\), so that the whole denominator is \(4\sqrt{H_{12}H_{21}}=2(\lambda_{S}-\lambda_{L})\). The sign of that difference is not free, and it is where a factor of \(-1\) is most easily lost: Equation (105.57) gives \(\lambda_{S}-\lambda_{L}=-\Delta m\,c^{2} -\tfrac{\ii\hbar}{2}\left(\Gamma_{S}-\Gamma_{L}\right)\), with both terms negative, because \(K_{S}\) is at once the lighter state and the broader one. With \(H_{21}=M_{12}^{*}-\tfrac{\ii\hbar}{2}\Gamma_{12}^{*}\) the numerator is \(H_{12}-H_{21}=2\ii\,\mathrm{Im}\, M_{12} +\hbar\,\mathrm{Im}\,\Gamma_{12}\), whence
where \(\Gamma_{L}\ll\Gamma_{S}\) has been used. Both numerator terms vanish in a \(T\)-invariant theory, as they must.
Equation (105.58) makes a prediction that is one of the most satisfying numerical accidents in physics. The \(\pi\pi\) channel dominates \(\Gamma_{S}\) and hence \(\Gamma_{12}\), and it carries essentially the phase of \(M_{12}\), so \(\mathrm{Im}\,\Gamma_{12}\) is negligible next to \(\mathrm{Im}\, M_{12}\). The phase of \(\epsilon\) is then fixed by the denominator alone, up to the \(180^{\circ}\) carried by the sign of \(\mathrm{Im}\, M_{12}\):
With the measured [Navas:2024]
this gives \(\arg\epsilon=\arctan(0.948)=43.5^{\circ}\), and the measured value is \((43.5\pm0.5)^{\circ}\) — which also settles the branch, since \(+43.5^{\circ}\) rather than \(-136.5^{\circ}\) requires \(\mathrm{Im}\, M_{12}<0\) in the phase convention of Equation (105.45). The agreement is a check on the whole construction, and the near-equality \(\Delta m\,c^{2} \approx\hbar\Gamma_{S}/2\) that produces it is, so far as anyone knows, a coincidence.
Strangeness oscillation
Because \(\ket{K^{0}}\) is not an eigenstate of \(\mathbf{H}\), a beam produced as pure \(K^{0}\) develops a \(\bar{K}^{0}\) component. Neglecting \(\epsilon\) (a \(10^{-3}\) correction) and using \(\ket{K^{0}}=(\ket{K_{S}}+\ket{K_{L}})/\sqrt{2}\),
so that, projecting back onto the strangeness basis,
The oscillation frequency is \(\Delta m\,c^{2}/\hbar=5.29\times 10^{9}\,/\mathrm{s}\), comparable to \(1/\tau_{S}\), so the interference term is observable over the first few \(\tau_{S}\) and then dies, leaving equal \(K^{0}\) and \(\bar{K}^{0}\) content in the surviving \(K_{L}\) beam.
It is worth pausing on what Equation (105.61) makes measurable. A mass difference of \(3.5\,\mu\mathrm{eV}/c^{2}\) is being read off a particle of mass \(497.6\,\mathrm{MeV}/c^{2}\):
No spectrometer resolves that. The interferometer does, because the observable is a phase accumulated over a macroscopic flight path, and this is the same trick that makes the neutral-meson systems the sharpest \(CPT\) laboratories in Section 105.5.3.
Regeneration
A \(K_{L}\) beam traversing matter emerges with a \(K_{S}\) component [Pais:1955]. The reason is that the strong interaction sees strangeness, not \(CP\): the forward scattering amplitude \(f\) for \(K^{0}\) on a nucleus differs from the amplitude \(\bar{f}\) for \(\bar{K}^{0}\), because \(\bar{K}^{0}\) can convert a nucleon into a hyperon while \(K^{0}\) cannot. Writing the transmitted state through a slab of \(N\) scatterers per unit volume and thickness \(\ell\), each strangeness component picks up its own index of refraction, and the regenerated amplitude is
proportional to \(f-\bar{f}\) and therefore vanishing if the medium cannot distinguish the two. Coherent regeneration is the standard tool for producing a \(K_{S}\) beam far from the target; it is also the background that any claim of \(K_{L}\to\pi\pi\) must exclude, and excluding it was the principal experimental burden of Experiment: CP Violation.
The 1964 discovery
Christenson, Cronin, Fitch and Turlay put a two-arm spectrometer at the end of a \(17\,\mathrm{m}\) helium-filled decay volume at the Brookhaven AGS, far enough downstream that every \(K_{S}\) had decayed, and looked for two-body decays of the surviving beam [Christenson:1964]. They found 45 events with the reconstructed invariant mass of the kaon and with the reconstructed parent momentum collinear with the beam, over a background of about 12. The branching fraction was about \(2\times10^{-3}\); the modern value is \((1.967\pm0.010)\times10^{-3}\) [Navas:2024]. The collinearity is what kills regeneration — a regenerated \(K_{S}\) from the residual gas would produce a non-collinear parent — and the substance of the apparatus and its controls is Experiment: CP Violation; see Phenomenon 111.8.
The immediate reading of the result is Equation (105.56): the long-lived state is not the \(CP\) eigenstate \(\ket{K_{2}}\) but carries an admixture \(\epsilon\ket{K_{1}}\), and the \(CP\)-even component decays to two pions. Quantitatively,
and if the only source of \(CP\) violation were the mixing, both would equal \(\epsilon\). The measured magnitude is
consistent with Equation (105.59).
Indirect and direct violation
The distinction that took thirty-five years to settle is whether \(CP\) violation lives only in the mass matrix — in the state — or also in the decay amplitudes. Decompose the two-pion final state by isospin. Two pions in a relative \(S\) wave from a spinless parent must have a symmetric isospin wavefunction, because the spatial part is symmetric and pions are bosons; of the three combinations \(I=0,1,2\) available to two isospin-1 particles, \(I=1\) is antisymmetric and is therefore forbidden. Write the amplitudes for the two allowed channels as
where \(\delta_{I}\) is the \(\pi\pi\) strong rescattering phase — real and identical for the two, since the strong interaction conserves \(CP\) — and any \(CP\) violation in the decay resides in \(\mathrm{Im}\, A_{I}\). The relative sign of the two conjugate amplitudes is fixed by Equation (105.45).
The observed states are recovered from the isospin eigenstates by the Clebsch–Gordan coefficients of \(1\otimes1\) in the \(I_{3}=0\) sector. Writing \(\ket{m_{1},m_{2}}\) for the product state of two pions of third components \(m_{1},m_{2}\), those coefficients are
the third combination \(\ket{1}=\left(\ket{+,-}-\ket{-,+}\right)/\sqrt{2}\) being the antisymmetric one already excluded. Since \(\ket{\pi^{+}\pi^{-}}\) in an \(S\) wave is the normalized symmetric combination \(\left(\ket{+,-}+\ket{-,+}\right)/\sqrt{2}\), while \(\ket{\pi^{0}\pi^{0}}=\ket{0,0}\), projecting gives
The two are orthogonal and normalized, as they must be. The weight of \(A_{2}\) relative to \(A_{0}\) is \(1/\sqrt{2}\) in the charged channel and \(-\sqrt{2}\) in the neutral one, the second being \(-2\) times the first, and that single ratio is the entire origin of the factor \(-2\) below. Carrying Equation (105.70) through gives
with
Everything in Equation (105.72) is a statement about the decay amplitudes and none of it about the mass matrix, so \(\epsilon'\ne0\) is \(CP\) violation in the decay: direct \(CP\) violation. Note the structure of the bracket — a difference of two phases. If the two isospin amplitudes carried a common \(CP\)-odd phase it could be rotated away, and \(\epsilon'\) would vanish; direct \(CP\) violation requires two interfering amplitudes with different weak phases, a point proved in general at Proposition 105.39.
The observable is the double ratio.
Rests on Equation (105.71).
Derives Proposition 105.24. From Equation (105.71), \(\eta_{00}/\eta_{+-}=(1-2\epsilon'/\epsilon)/(1+\epsilon'/\epsilon)\). With \(\abs{\epsilon'/\epsilon}\sim10^{-3}\), expand to first order: \(\eta_{00}/\eta_{+-}\approx1-3\epsilon'/\epsilon\), whose squared modulus is \(1-6\mathrm{Re}(\epsilon'/\epsilon)\).
∎The double ratio is designed so that every quantity that is common to the four channels — kaon flux, detector solid angle, most of the efficiency — cancels. NA48 at CERN [Fanti:1999] and KTeV at Fermilab [AlaviHarati:1999] measured it with independent systematics, and the world average is [Navas:2024]
See Phenomenon 111.11. A hypothetical superweak \(\Delta S=2\) interaction acting only in the mass matrix reproduces \(\epsilon\) exactly and predicts \(\epsilon'=0\); Equation (105.74) excludes it. That \(\epsilon'/\epsilon\) is itself of order \(10^{-3}\) rather than order unity is largely the \(\Delta I=\tfrac12\) rule, the empirical fact that \(\abs{A_{0}/A_{2}}\approx22\), which enhances \(\epsilon\) and suppresses the \(I=2\) amplitude appearing in Equation (105.72).
An absolute definition of matter
The semileptonic decays give the cleanest statement of what \(CP\) violation means. The rule \(\Delta S=\Delta Q\), which holds because the strange quark decays to an up quark by emitting a \(W^{-}\), allows \(K^{0}\to\pi^{-}e^{+}\nu_{e}\) and forbids \(K^{0}\to\pi^{+}e^{-}\bar{\nu}_{e}\); the conjugate statements hold for \(\bar{K}^{0}\). From Equation (105.56) the \(K^{0}\) content of \(K_{L}\) carries amplitude \((1+\epsilon)\) and the \(\bar{K}^{0}\) content \(-(1-\epsilon)\), so
Numerically \(2\abs{\epsilon}\cos43.5^{\circ}=3.23\times10^{-3}\), against the measured \((3.32\pm0.06)\times10^{-3}\) [Navas:2024]; see Phenomenon 111.9. The significance is not the number but its status. Instruct a correspondent, with whom no sample and no convention has been shared, to prepare a long-lived neutral kaon beam and count the charge of the emitted lepton: the commoner sign defines the positron, and therefore matter, absolutely. Before 1964 no such instruction existed — \(C\) and \(P\) violation alone leave the labelling conventional, since one can compensate a mirror reflection by an exchange of matter and antimatter. Only \(CP\) violation makes “matter” an objective term.
The CKM phase
Kobayashi and Maskawa asked, in 1973, what the minimal modification of the Standard Model would be that admits \(CP\) violation at all [Kobayashi:1973]. Their answer is a counting argument, and it predicted a third generation of quarks from a single measured asymmetry of \(2\times10^{-3}\).
Let there be \(N\) generations of quarks. After diagonalizing the mass matrices, the charged-current interaction is governed by a unitary \(N\times N\) matrix \(V\), of which
parameters are physical. There is no \(CP\)-violating phase for \(N\le2\) and exactly one for \(N=3\). Rests on Equation (104.46).
Derives Theorem 105.25. The Yukawa couplings of Electroweak Unification and the Higgs Boson give arbitrary complex mass matrices \(M_{u}\) and \(M_{d}\); each is brought to positive diagonal form by a bi-unitary transformation \(M=U_{L}^{\dagger}M_{\mathrm{diag}}U_{R}\), and the charged current, which couples \(\bar{u}_{L}\gamma^{\mu}d_{L}\), retains the combination \(V=U_{L}^{u\dagger}U_{L}^{d}\), a general \(N\times N\) unitary matrix.
A unitary \(N\times N\) matrix depends on \(N^{2}\) real parameters (the condition \(V^{\dagger}V=\identity\) imposes \(N^{2}\) real conditions on \(2N^{2}\) real entries). Of these, an orthogonal \(N\times N\) matrix would use \(N(N-1)/2\); the remainder, \(N^{2}-N(N-1)/2=N(N+1)/2\), are phases.
Not all phases are physical. The quark fields may be rephased, \(u_{j}\to\ee^{\ii\alpha_{j}}u_{j}\) and \(d_{k}\to\ee^{\ii\beta_{k}}d_{k}\), which changes \(V_{jk}\to\ee^{-\ii\alpha_{j}}V_{jk}\ee^{\ii\beta_{k}}\) without touching the diagonal mass terms. That is \(2N\) phases, but the overall common phase \(\alpha_{j}=\beta_{k}=\theta\) leaves \(V\) unchanged, so only \(2N-1\) are usable. Hence
For \(N=2\) this is zero: the two-generation mixing matrix is a real rotation through the Cabibbo angle [Cabibbo:1963], and no phase survives. For \(N=3\) it is one.
∎The argument's force is that it is a counting theorem, not a model. Given that \(CP\) violation exists and that the Standard Model gauge structure is right, three generations are the minimum. In 1973 three quarks were known, the fourth was conjectural, and the fifth and sixth were found in 1977 and 1995.
For any \(i\ne k\) and \(j\ne l\),
defining a single real number \(J\) independent of which quartet is chosen [Jarlskog:1985].
The quartet in Equation (105.77) is invariant under the rephasing \(V_{jk}\to\ee^{-\ii\alpha_{j}}V_{jk}\ee^{\ii\beta_{k}}\), since each index appears once with a phase and once with its conjugate; \(J\) is therefore a physical observable and not a parametrization artefact. Unitarity makes all nine quartets equal up to sign, which is what Equation (105.77) records. The measured value is [Navas:2024] [Charles:2005]
a small number whose smallness is a fact about the mixing angles and not about the phase, which is of order unity.
\(CP\) is violated in the quark sector if and only if
is nonzero. In particular \(CP\) is conserved if any two quarks of the same charge are degenerate, or if \(J=0\). Rests on Definition 105.26 and Theorem 105.25.
Derivation. Derives Proposition 105.27. The left-hand side is invariant under the field redefinitions that leave the physics alone, and it vanishes identically if the two Hermitian matrices commute, which happens exactly when they are simultaneously diagonalizable — i.e. when \(V\) can be taken real. Evaluating the determinant in the basis where \(M_{u}M_{u}^{\dagger}\) is diagonal, each entry of the commutator is \((m_{u_{i}}^{2}-m_{u_{j}}^{2})c^{4}\) times the corresponding entry of \(V\,\diag(m_{d}^{2}c^{4})\,V^{\dagger}\), and expanding the \(3\times3\) determinant collects the six mass differences together with the imaginary part of a quartet, which is \(J\) by Definition 105.26. That last step is where the sign on the right is fixed, and it is fixed by a convention worth stating: \(V\) is the matrix whose \((u,d)\) entry is \(V_{ud}\), so that in the up-diagonal basis \(M_{d}M_{d}^{\dagger}c^{4}=V\diag(m_{d}^{2}c^{4})V^{\dagger}\), and \(J\) is the quartet of Definition 105.26; with the opposite convention for \(V\) both \(J\) and the sign reverse together, and the statement is unchanged.
The factors of \(c^{4}\) in Equation (105.79) are not decoration. \(M_{u}M_{u}^{\dagger}\) is a squared mass, \(\mathrm{kg}^{2}\), so \(M_{u}M_{u}^{\dagger}c^{4}\) is a squared energy; each entry of the commutator of two such matrices carries \(\mathrm{J}^{4}\) and the \(3\times3\) determinant \(\mathrm{J}^{12}\), which is also what the right-hand side's six factors of \(\mathrm{J}^{2}\) come to. Written without them the left-hand side would be \(\mathrm{kg}^{12}\) and the two sides would not have the same dimension at all.
∎If two up-type quarks had equal mass, the freedom to rotate between them could be used to make \(V\) real; the phase would be unphysical. \(CP\) violation in the Standard Model therefore requires three non-degenerate generations, and the extreme hierarchy of the observed masses is precisely what makes Equation (105.79) numerically tiny — the fact that returns, with consequences, in Section 105.6.2.
Unitarity of \(V\) gives six orthogonality relations, of which
has three terms of comparable magnitude and so draws a non-degenerate triangle in the complex plane. Its area is \(J/2\): for any three complex numbers summing to zero, the triangle they close has area \(\tfrac12\abs{\mathrm{Im}(z_{1}z_{2}^{*})}\), and that imaginary part is a Jarlskog quartet. The angles are conventionally \(\alpha,\beta,\gamma\), and the programme of flavour physics is to measure the sides and the angles by independent means and ask whether they close. They do, at the few-per-cent level, which is the quantitative statement that one phase accounts for all observed \(CP\) violation in the quark sector; the detailed over-determination belongs to Flavour Physics and Neutrinos.
CP violation beyond the kaon
The B factories
If one phase controls everything, the kaon's \(10^{-3}\) effect and a possible order-unity effect elsewhere are the same physics. The Standard Model predicts the latter in \(B^{0}\to J/\psi\,K^{0}_{S}\), where the interference is between decay after mixing and decay without mixing, and no small ratio of amplitudes suppresses it.
For a final state \(f\) that is a \(CP\) eigenstate accessible to both \(B^{0}\) and \(\bar{B}^{0}\), define \(\lambda=\left(q/p\right)\left(\bar{\mathcal{A}}_{f} /\mathcal{A}_{f}\right)\), with \(q/p\) the mixing parameter analogous to \((1-\epsilon)/(1+\epsilon)\). Then
with \(S_{f}=2\mathrm{Im}\,\lambda/(1+\abs{\lambda}^{2})\) and \(C_{f}=(1-\abs{\lambda}^{2})/(1+\abs{\lambda}^{2})\). Rests on Equations (105.53), (105.55) and (105.56).
Derivation. Derives Proposition 105.28. The \(B\) system has \(\Delta\Gamma\approx0\), so the two mass eigenstates share a width and the time evolution of an initially pure \(B^{0}\) is \(\ket{B^{0}(t)}=g_{+}(t)\ket{B^{0}}+(q/p)g_{-}(t)\ket{\bar{B}^{0}}\) with \(g_{\pm}(t)=\ee^{-\Gamma t/2} \ee^{-\ii\bar{m}c^{2}t/\hbar}\) times \(\cos\) or \(+\ii\sin\) of \(\Delta m_{d}c^{2}t/(2\hbar)\) respectively, \(\bar{m}\) being the mean of the two masses, obtained by diagonalizing Equation (105.53) exactly as in Equations (105.55) and (105.56).
That sign in \(g_{-}\) is not free, and it is the one place the derivation can silently go wrong. The eigenvector \(p\ket{B^{0}}+q\ket{\bar{B}^{0}}\) is the light state, exactly as \(\ket{K_{S}}\) is in Equation (105.55), so writing \(\Delta m_{d}=m_{H}-m_{L}>0\) gives \(g_{\pm}=\tfrac12\left(\ee^{-\ii\lambda_{L}t/\hbar} \pm\ee^{-\ii\lambda_{H}t/\hbar}\right)\) and hence \(+\ii\sin\) and not \(-\ii\sin\); taking the other sign reverses the coefficient of \(\sin\Delta m_{d}c^{2}t/\hbar\) in Equation (105.81), which is the whole measurement. Then \(\mathcal{A}(B^{0}(t)\to f) =\mathcal{A}_{f}\left[g_{+}+\lambda g_{-}\right]\) and \(\mathcal{A}(\bar{B}^{0}(t)\to f) =(p/q)\mathcal{A}_{f}\left[g_{-}+\lambda g_{+}\right]\). Squaring, using \(\abs{g_{\pm}}^{2}=\tfrac12\ee^{-\Gamma t} (1\pm\cos\Delta m_{d}c^{2}t/\hbar)\) and \(g_{+}^{*}g_{-}=\tfrac{\ii}{2}\ee^{-\Gamma t} \sin\Delta m_{d}c^{2}t/\hbar\), and forming the asymmetry gives Equation (105.81).
∎For \(f=J/\psi\,K^{0}_{S}\) the decay proceeds through a single dominant amplitude, so \(\abs{\lambda}=1\), \(C_{f}=0\), and the phase of \(\lambda\) is the mixing phase \(\ee^{-2\ii\beta}\) of Equation (105.80); hence
The measurement requires knowing which meson was which at production, and the \(B\) factories solved this by producing the pair coherently from the \(\Upsilon(4S)\): the two mesons are in a \(C\)-odd, \(L=1\) state, so until one decays there is exactly one \(B^{0}\) and one \(\bar{B}^{0}\), and a flavour-specific decay of one tags the other. Belle [Abe:2001] and BaBar [Aubert:2001] reported simultaneously, with \(\sin2\beta=0.99\pm0.14\pm0.06\) and \(0.59\pm0.14\pm0.05\) respectively — two values differing by about two standard deviations, which is a useful reminder that first measurements scatter. The world average has since settled at \(\sin2\beta=0.708\pm0.011\) [Navas:2024], and \(\Delta m_{d}c^{2}=3.334(13)\times 10^{-4}\,\mathrm{eV}\). See Phenomenon 111.13.
The effect is of order unity. \(CP\) violation is therefore not intrinsically small: it is small in the kaon because the kaon system happens to sit where the interfering amplitudes are unequal, and large in this \(B\) channel because they are not.
Charm
The up-type sector was the last to yield. LHCb measured the difference of the time-integrated \(CP\) asymmetries of two singly Cabibbo-suppressed charm decays [Aaij:2019],
nonzero at \(5.3\) standard deviations. The difference is taken because production and detection asymmetries cancel in it. The effect requires interference of tree and loop (“penguin”) amplitudes with different strong phases, per Proposition 105.39, and the long-distance strong phases are not calculable, so the measurement establishes the phenomenon without sharply testing the CKM prediction. See Phenomenon 111.15.
Leptons
The mixing matrix of the leptons admits the same counting as Theorem 105.25, so with three neutrino generations there is one Dirac phase \(\delta_{CP}\). T2K, comparing \(\nu_{\mu}\to\nu_{e}\) with \(\bar{\nu}_{\mu}\to\bar{\nu}_{e}\) over a \(295\,\mathrm{km}\) baseline, finds the data preferring values of \(\delta_{CP}\) near \(-\pi/2\) and excluding \(CP\) conservation (\(\delta_{CP}=0\) or \(\pi\)) at the \(95\,\mathrm{\%}\) confidence level [Abe:2020]. This is a constraint and not a discovery: the exclusion is a two-sigma statement resting on a comparison of two event samples of a few tens of events each, and it depends on the mass ordering, which is not yet settled. The apparatus and the oscillation formalism are Experiment: Neutrino Oscillations; the honest summary is that leptonic \(CP\) violation is likely and unproven.
Direct tests of time reversal
Everything in Section 105.3 tests \(CP\). Given the \(CPT\) theorem, \(CP\) violation is \(T\) violation, and one might think nothing remains to be measured. But that inference assumes the theorem whose hypotheses are the deepest in the subject, and a test of \(T\) that does not assume \(CPT\) is a genuinely independent piece of information — indeed it is the only way to separate the three symmetries experimentally. This section describes the two measurements that compare a process directly with its motion reverse, and the static observable that needs no process at all.
The CPLEAR asymmetry
Kabir's observation is that the neutral kaon offers a transition and its exact reverse [Kabir:1970]. The process \(K^{0}\to\bar{K}^{0}\) and the process \(\bar{K}^{0}\to K^{0}\) are related by exchanging initial and final state; both particles are at rest, so no momenta or spins need reversing, and motion reversal is nothing but the exchange. Define
If \(T\) is a symmetry, Equation (105.25) forces \(A_{T}=0\) identically.
To first order in \(\epsilon\),
independent of \(t\). Rests on Equations (105.55), (105.56) and (105.84).
Derives Proposition 105.29. Invert Equations (105.55) and (105.56) to express \(\ket{K^{0}}\) and \(\ket{\bar{K}^{0}}\) in terms of \(\ket{K_{S}},\ket{K_{L}}\). Propagating and projecting, the two transition amplitudes are
and the reverse with \((1-\epsilon)/(1+\epsilon)\) replaced by its inverse; the time-dependent bracket is common. Hence, writing \(r=\abs{(1+\epsilon)/(1-\epsilon)}^{2}\approx1+4\,\mathrm{Re}\,\epsilon\) and therefore \(r^{-1}\approx1-4\,\mathrm{Re}\,\epsilon\) to the same order, so that the numerator is \(r-r^{-1}\approx8\,\mathrm{Re}\, \epsilon\) while the denominator is \(r+r^{-1}\approx2\):
and the time dependence cancels between numerator and denominator.
∎Numerically \(4\abs{\epsilon}\cos43.5^{\circ}=6.46\times10^{-3}\). The CPLEAR collaboration at CERN's Low Energy Antiproton Ring measured
the first direct observation of a time-reversal asymmetry [Angelopoulos:1998]. The method is worth stating because the tagging is the whole difficulty. CPLEAR produced neutral kaons in \(p\bar{p}\to K^{\mp}\pi^{\pm}K^{0}/\bar{K}^{0}\) at rest; strangeness conservation in the strong production ties the sign of the accompanying charged kaon to the strangeness of the neutral one, so the initial state is tagged by a charge measured in the same event. The final strangeness is tagged by the charge of the lepton in the semileptonic decay, using \(\Delta S=\Delta Q\) as in Equation (105.75).
The objection, and the answer
The measurement has been criticized on the ground that it is not free of \(CPT\) assumptions, and the criticism is partly right. The final-state tag uses the semileptonic amplitudes, and if those amplitudes themselves violated \(CPT\) — or if \(\Delta S=\Delta Q\) failed — the observed asymmetry could be produced without any \(T\) violation in the mixing. Two answers are available. First, the same data set constrains the offending parameters: the \(K_{L}\) and \(K_{S}\) semileptonic charge asymmetries and their time dependence bound the \(CPT\)-violating and \(\Delta S\ne\Delta Q\) amplitudes at a level well below the measured \(A_{T}\), so the interpretation survives within the standard phenomenological parametrization [Navas:2024]. Second — and this is the honest statement — no measurement of a decaying system can be completely free of assumptions about the decay, because the observation is made through the decay products. A \(T\) test with no such entanglement requires a stable system, which is what Section 105.4.3 provides at the cost of measuring something else.
Time reversal in the $B$ system
The BaBar measurement is the cleanest of its kind, and it works by exploiting entanglement as a state-preparation device [Lees:2012].
The \(\Upsilon(4S)\) decays to \(B^{0}\bar{B}^{0}\) in a \(C\)-odd, \(L=1\) state, so before either decays the pair is in the antisymmetric combination
the second form following because an antisymmetric combination of a two-dimensional space is basis-independent up to a phase; \(B_{\pm}\) denote the states that do and do not decay to the \(CP\) eigenstate \(J/\psi K^{0}_{S}\). This is the Einstein–Podolsky–Rosen correlation of Entanglement and Bell Tests, used here not to test locality but as a preparation: observing the decay of one meson at time \(t_{1}\) to a flavour-specific state \(\ell^{\pm}X\) projects the other onto a definite flavour at that instant, and observing a decay to \(J/\psi K^{0}_{S}\) projects it onto a definite \(CP\) label.
The comparison is then available in both time orders. Events in which the flavour tag comes first and the \(CP\) tag second measure the transition \(B^{0}\to B_{-}\); events in the opposite order measure \(B_{-}\to B^{0}\). These are a process and its motion reverse: the initial and final states are exchanged, and nothing is conjugated. That is what distinguishes the measurement from a \(CP\) test, which would compare \(B^{0}\to B_{-}\) with \(\bar{B}^{0}\to B_{-}\), exchanging particle for antiparticle rather than initial for final. BaBar extracted all three sets of parameters — \(T\)-odd, \(CP\)-odd and \(CPT\)-odd — from one fit to the same events, and found the \(T\)-violating parameters nonzero at \(14\) standard deviations, with
while the \(CPT\)-odd parameters were consistent with zero.
Two caveats keep the claim honest. The \(B\) mesons have a finite width, so the “reversed” transition is not an exact motion reverse of the original — the two occur over different intervals of proper time, and the equality of the rates is derived within the same effective two-state formalism as Proposition 105.22. And the projection interpretation of the tag is exact only if the tagging decay is flavour-specific, which is an assumption of the same family as \(\Delta S=\Delta Q\) above. What the measurement does establish, and what no earlier one did, is that the \(T\)-odd and the \(CP\)-odd parameters are separately determined from a single data set and are separately nonzero.
Electric dipole moments as $T$ tests
The observables above compare two processes. A permanent electric dipole moment compares nothing: it is a static property of a single state, and its existence would violate \(P\) and \(T\) at once. Purcell and Ramsey made the point in 1950, at a time when parity was universally assumed, with an argument that is really about method [Purcell:1950]: a symmetry that has never been tested in a given sector is an assumption, and the way to test it is to measure the quantity it forbids.
Let \(\ket{\psi}\) be a nondegenerate state of a system with definite total angular momentum \(j\), and let \(\hat{\vect{d}}=\int\dd^{3}x\,\vect{x}\,\rho(\vect{x})\) be the electric dipole operator. If \(\avg{\hat{\vect{d}}}\ne0\) then the Hamiltonian violates both \(P\) and \(T\). Rests on Equations (105.6), (105.17) and (105.30).
Derives Theorem 105.30. Only one fact about the state is needed, and it can be had without any general machinery. Write \(\ket{\psi}=\ket{j,m}\), the quantization axis chosen along \(\avg{\hat{\vect{J}}}\). The rotation \(U(\varphi)=\ee^{-\ii\varphi\hat{J}_{z}/\hbar}\) about that axis multiplies \(\ket{\psi}\) by the phase \(\ee^{-\ii m\varphi}\) and so changes no expectation value whatever; but on a vector operator it acts as \(U^{\dagger}\hat{V}_{i}U=R_{ij}(\varphi)\hat{V}_{j}\). Hence \(\avg{\hat{\vect{V}}}=R(\varphi)\avg{\hat{\vect{V}}}\) for every \(\varphi\), and a vector fixed by every rotation about \(\hat{\vect{z}}\) has no transverse component: every vector expectation value in such a state points along \(\avg{\hat{\vect{J}}}\). (The general statement, of which this is the diagonal case, is the projection theorem carried by the Wigner–Eckart theorem.) Fixing the constant so that \(d\) is the value of \(\avg{\hat{d}_{z}}\) in the stretched state \(m=j\),
with \(d\) real because \(\hat{\vect{d}}\) is Hermitian: the dipole, if it exists, must point along the spin, because the spin is the only vector the state possesses. Nondegeneracy is what licenses the first step — it is what gives the state a definite \(m\) with no admixture of opposite parity.
The interaction with a static electric field is \(H_{d}=-\avg{\hat{\vect{d}}}\cdot\vect{E} \propto-d\,\avg{\hat{\vect{J}}}\cdot\vect{E}\). By Equation (105.6) \(\vect{E}\) is polar and \(\hat{\vect{J}}\) axial; by Equation (105.30) \(\vect{E}\) is \(T\)-even and \(\hat{\vect{J}}\) is \(T\)-odd; and by Equation (105.17) \(\vect{E}\) is \(C\)-odd, while \(\hat{\vect{J}}\), being an angular momentum, is \(C\)-even. So \(\hat{\vect{J}}\cdot\vect{E}\) is odd under each of the three separately. The dipole operator supplies the missing sign under \(C\): the charge density \(\rho\) reverses under charge conjugation, so \(\hat{\vect{d}}\) is \(C\)-odd and a particle and its antiparticle carry opposite dipoles. The interaction \(-\hat{\vect{d}}\cdot\vect{E}\) is therefore \(C\)-even, \(P\)-odd and \(T\)-odd. A nonzero \(d\) means the Hamiltonian contains a term with those properties, i.e. \(P\) and \(T\) are violated (and, by the \(CPT\) theorem, \(CP\) with them — the three signs multiply to \(CPT\)-even, as Theorem 105.34 requires of any term in a Hermitian local Lagrangian).
∎The hypothesis of nondegeneracy is essential and is often omitted. The hydrogen atom in its \(n=2\) level, where \(2S_{1/2}\) and \(2P_{1/2}\) are degenerate up to the Lamb shift, exhibits a linear Stark effect: the field mixes states of opposite parity and the energy shifts linearly in \(\abs{\vect{E}}\), exactly as a permanent dipole would. No symmetry is violated, because the state that acquires the dipole is not an eigenstate of parity to begin with. The same mechanism is what makes polar molecules useful in EDM searches: a molecule such as ThO has closely spaced levels of opposite parity, which a modest laboratory field fully polarizes, producing an enormous effective internal field acting on the electron — in ThO of order \(7.8\times 10^{12}\,\mathrm{V}/\mathrm{m}\), which is the \(78\,\mathrm{GV}/\mathrm{cm}\) of the ACME literature, and enormously larger than any field a laboratory can apply directly. The molecule is a lever, not a violation.
How the measurement is made
The technique is Ramsey's method of separated oscillatory fields [Ramsey:1950]. Neutrons in a storage cell see parallel magnetic and electric fields; the Larmor precession frequency is
the sign depending on whether \(\vect{E}\) is parallel or antiparallel to \(\vect{B}\). Reversing \(\vect{E}\) and taking the difference isolates the dipole:
The statistical sensitivity of a Ramsey measurement with coherence time \(T\) on \(N\) neutrons is
with \(\alpha_{v}\lesssim1\) the fringe visibility. Inserting the PSI apparatus's parameters — \(E\approx1.1\times 10^{6}\,\mathrm{V}/\mathrm{m}\), \(T=180\,\mathrm{s}\), and of order \(10^{7}\) ultracold neutrons accumulated over the run — gives \(\sigma(d_{n})\sim8\times 10^{-47}\,\mathrm{C}\,\mathrm{m}\), which is the order of the published limit and confirms that the experiment is statistics-limited rather than limited by any subtlety.
Seven decades of measurement find no permanent electric dipole moment of the neutron. The current bound, from ultracold neutrons at the Paul Scherrer Institute, is
at \(90\,\mathrm{\%}\) confidence [Abel:2020], improving on [Baker:2006]. The same number is customarily quoted as \(1.8\times10^{-26}\) in units of the elementary charge times the centimetre; the conversion is \(1\,e\,\mathrm{cm} =1.602176634\times 10^{-21}\,\mathrm{C}\,\mathrm{m}\). See Phenomenon 111.19. Rests on Theorem 105.30, Equation (105.91) and Equation (105.92).
Derivation. Derives Phenomenon 105.32. That the observable is forbidden at all is Theorem 105.30; what is derived here is the number, since a bound is a statement about a resolved frequency and Equation (105.91) is what converts the two. At the PSI storage field \(E\approx1.1\times 10^{6}\,\mathrm{V}/\mathrm{m}\), a dipole equal to Equation (105.93) would shift the Ramsey resonance by
against a Larmor frequency \(2\mu_{n}B/h\approx29\,\mathrm{Hz}\) at the working field \(B\approx1\,\mu\mathrm{T}\): the experiment is a determination of a frequency ratio to \(7\times10^{-9}\), made twice with \(\vect{E}\) reversed. That such a resolution is available follows from Equation (105.92): with \(T=180\,\mathrm{s}\) and of order \(10^{7}\) ultracold neutrons the statistical limit is \(8.4\times 10^{-47}\,\mathrm{C}\,\mathrm{m}\), within a factor of three of the published \(90\,\mathrm{\%}\) bound, which is the quantitative content of the claim that the measurement is statistics-limited. The bound is therefore not a theoretical estimate at all: it is the frequency resolution of a Ramsey interferometer, divided by \(4E/h\).
∎The number is hard to feel, so it is worth converting once. A dipole of \(2.9\times 10^{-47}\,\mathrm{C}\,\mathrm{m}\) corresponds to separating one elementary charge from its opposite by \(1.8\times 10^{-28}\,\mathrm{m}\) — about \(10^{13}\) times smaller than the neutron itself, whose charge radius is of order \(0.8\,\mathrm{fm}\). Scaled up so that the neutron were the size of the Earth, the charge separation would be a few micrometres.
The electron is bounded even harder, in units of its own size, by experiments on polar molecules. ACME, using the metastable \(H\) state of ThO, obtained \(\abs{d_{e}}<1.8\times 10^{-50}\,\mathrm{C}\,\mathrm{m}\) (\(1.1\times10^{-29}\,e\,\mathrm{cm}\)) at \(90\,\mathrm{\%}\) confidence [Andreev:2018], and a trapped molecular-ion experiment on HfF\(^{+}\) has since reached \(\abs{d_{e}}<6.6\times 10^{-51}\,\mathrm{C}\,\mathrm{m}\) (\(4.1\times10^{-30}\,e\,\mathrm{cm}\)) [Roussy:2023].
What makes these null results powerful is the size of the Standard Model background. In the quark sector the CKM phase can generate a neutron EDM only at three loops, and the chiral suppression is severe: the estimate is of order \(10^{-32}\,e\,\mathrm{cm}\), six orders below Equation (105.93). For the electron the CKM contribution requires four loops and is of order \(10^{-38}\,e\,\mathrm{cm}\), nine orders below the current limit. Every improvement in these experiments therefore probes a range in which the Standard Model predicts nothing observable, so a nonzero result would be unambiguous — and a null result excludes new sources of \(CP\) violation over a wide range without any background to subtract. The one Standard Model contribution that is not negligible is the strong \(CP\) term of Section 105.7, which is why Equation (105.93) is simultaneously the sharpest constraint on \(\bar{\theta}\).
The CPT theorem
Statement and hypotheses
\(\Theta:=\mathcal{C}\mathcal{P}\mathcal{T}\), an antiunitary operator (the product of two unitaries and one antiunitary), acting on spacetime arguments by \(x\mapsto-x\) and exchanging particles with antiparticles.
Let a quantum field theory satisfy
-
Lorentz invariance: the fields transform under a representation of the proper orthochronous group \(L^{\uparrow}_{+}\) of Definition 39.5, and the dynamics is invariant under it and under translations;
-
locality: fields at spacelike separation commute (integer spin) or anticommute (half-integer spin);
-
the spectrum condition: the energy-momentum spectrum lies in the closed forward light cone, with a unique Poincaré-invariant vacuum;
-
hermiticity: the Hamiltonian is Hermitian, so the evolution is unitary.
Then there exists an antiunitary \(\Theta\) with \(\Theta\ket{0}=\ket{0}\) that commutes with the Hamiltonian: \(CPT\) is an exact symmetry [Lueders:1954] [Pauli:1955] [Jost:1957] [Luders:1957] [Bell:1955]. Rests on Definition 39.5, Lemma 105.35 and Equation (100.2).
Two proofs exist and they are of quite different kinds. The first works at the level of a Lagrangian and is short enough to give in full; the second is the theorem proper, since it assumes nothing about Lagrangians, and its machinery is not yet available in this book.
The Lagrangian-level proof
Let \(O^{\mu_{1}\cdots\mu_{n}}(x)\) be a local monomial built from Dirac bilinears, gauge fields and derivatives, carrying \(n\) uncontracted Lorentz indices. Then
Derives Lemma 105.35. For the five Dirac bilinears the rule is the last column of Table 105.1, obtained by multiplying the three preceding columns; the dagger appears because \(\mathcal{T}\) is antiunitary and reverses the order of operators in a product. Inspection confirms the count: the scalar and pseudoscalar (\(n=0\)) are even, the vector and axial vector (\(n=1\)) are odd, the tensor (\(n=2\)) is even — and, which is the essential point, \(CPT\) does not distinguish scalar from pseudoscalar nor vector from axial, so only \(n\) enters.
For the gauge field, \(\mathcal{P}A^{\mu}\mathcal{P}^{-1}\) and \(\mathcal{T}A^{\mu}\mathcal{T}^{-1}\) each lower the index (Equations (105.6) and (105.30)), so the product returns \(A^{\mu}\) with sign \(+1\), while \(\mathcal{C}\) contributes \(-1\) by Equation (105.17): total \((-1)^{1}\), as Equation (105.94) requires for \(n=1\).
For a derivative, let \(O'(x)=\Theta O(x)\Theta^{-1}=s\,O^{\dagger}(-x)\) with \(s\) the sign already established. Differentiating, \(\pp_{\mu}\left[O^{\dagger}(-x)\right] =-\left(\pp_{\mu}O\right)^{\dagger}(-x)\) by the chain rule, so \(\pp_{\mu}\) contributes one extra factor \(-1\) and one extra index, preserving Equation (105.94). A product of such objects multiplies the signs and adds the index counts, so the rule is closed under multiplication.
∎Proof of Theorem 105.34, Lagrangian level. Derives Theorem 105.34. A Lorentz-scalar Lagrangian density is a sum of monomials in which every index is contracted, either with the metric \(\eta_{\mu\nu}\) or with the Levi-Civita symbol \(\varepsilon^{\mu\nu\rho\sigma}\). In either case indices are used in pairs — \(\eta\) takes two, and \(\varepsilon\) takes four — so the total number of Lorentz indices appearing in any scalar monomial is even. By Lemma 105.35 the sign \((-1)^{n}\) is therefore \(+1\), and
the last step by hypothesis (iv), hermiticity of \(\Lag\). The action Equation (100.2) is then invariant,
because the Jacobian of \(x\to-x\) in four dimensions is \(+1\). Hence \(\comm{\Theta}{H}=0\).
Two hypotheses have been used silently and must be flagged. First, reordering the fermion operators inside a normal-ordered monomial — which the dagger in Equation (105.94) requires — produces a sign that is \(+1\) only if half-integer-spin fields anticommute; had a spinor field been quantized with commutators the argument would fail at exactly this step. This is hypothesis (ii), and its appearance here is the connection between the \(CPT\) theorem and spin–statistics noticed by Pauli [Pauli:1940] and made precise by Lüders and Zumino [Lueders:1958]. Second, the fields have been assumed to carry finite-dimensional representations built from vectors and spinors, which is what makes the index count exhaustive.
∎The Lagrangian argument is complete and correct as far as it goes, and it is what one uses in practice: given a candidate interaction, count its indices. What it does not do is prove the theorem, because it presupposes a Lagrangian, a perturbative expansion, and a set of fields transforming in specified representations. The theorem proper assumes none of these.
The axiomatic proof, honestly
Jost's proof [Jost:1957], and the canonical account in Streater and Wightman [Streater:1964], work directly with the vacuum expectation values
the Wightman functions [Wightman:1956]. The structure of the argument is the following, and it is worth seeing even though the machinery is reserved.
The spectrum condition (iii) says the energy-momentum of every state lies in the forward cone. Fourier-transforming Equation (105.96) in the differences \(\xi_{j}=x_{j}-x_{j+1}\), the support of the transform is confined to the forward cone, which by the Paley–Wiener theorem — imported here, not proved in this book, and not among the results of Fourier Analysis and Integral Transforms — makes \(W_{n}\) the boundary value of a function analytic in the forward tube, the domain in which each \(\xi_{j}\) has an imaginary part in the open forward cone. Lorentz covariance (i) then extends the domain: the extended tube is the union of the images of the forward tube under the complex proper Lorentz group \(L_{+}(\C)\), and \(W_{n}\) continues analytically to all of it.
The decisive fact is a statement about that group and nothing else. \(L_{+}(\C)\) is connected, and it contains the total inversion \(-\identity\), which corresponds to \(x\mapsto-x\). Over the reals, \(-\identity\) lies in \(L^{\downarrow}_{+}\), a different component from the identity (Equation (39.9)), and cannot be reached; over the complexes it can, by a continuous path through complex “rotations” by an imaginary angle. Analytic continuation along that path therefore gives
and Jost's theorem states that Equation (105.97) is equivalent to weak local commutativity at the Jost points — the totally spacelike configurations — which is a weakened form of hypothesis (ii). Equation (105.97) is the \(CPT\) condition, and from it the operator \(\Theta\) is reconstructed.
This is also the clean answer to the question a reader should be asking: why \(CPT\) and not \(CP\), or \(T\)? Because \(-\identity\) is in \(L_{+}(\C)\) while \(\Lambda_{P}\) and \(\Lambda_{T}\) separately have determinant \(-1\) and are in no component of \(L_{+}(\C)\) at all. The theorem protects exactly the combination that the complexified group can reach, and nothing else. No amount of ingenuity extends it to \(CP\).
The axiomatic CPT theorem: Jost's proof requires the Wightman axioms, the analyticity of the Wightman functions in the forward tube, the extended tube and the edge-of-the-wedge theorem. None of that machinery is developed here; the axiomatic-QFT chapter reserves it, and the proof belongs in Appendix A alongside it. What is proved in this chapter is the Lagrangian-level theorem, which is weaker in that it assumes a Lagrangian, a field content and a perturbative expansion.
What the theorem forbids
The consequences are three, and each is a short argument from antiunitarity. All three use the identity Equation (105.26); the reader who has not read that line should read it now, because everything below is a corollary of it.
A particle and its antiparticle have equal masses. Rests on Definition 105.33, Theorem 105.34 and Equation (105.26).
Derives Theorem 105.36. Let \(\ket{a}\) be a one-particle state at rest with spin projection \(s\), so \(H\ket{a}=m_{a}c^{2}\ket{a}\) and hence \(\bra{a}H\ket{a}=m_{a}c^{2}\), a real number. By Definition 105.33, \(\Theta\ket{a}=\ket{\bar{a}}\) up to a phase (the momentum is zero and so is unaffected; the spin projection reverses, which does not matter since the rest energy does not depend on it). Apply Equation (105.26) with \(\mathcal{A}=\Theta\), \(O=H\), \(\varphi=\psi=\ket{a}\), using \(\Theta^{-1}H\Theta=H\) from Theorem 105.34:
The phase cancels between bra and ket.
∎A particle and its antiparticle carry electric charges of equal magnitude and opposite sign, and magnetic moments of equal magnitude and opposite sign. Rests on Lemma 105.35 and Equation (105.26).
Derives Theorem 105.37. Both statements follow from the transformation of the current operator. By Lemma 105.35 with \(n=1\), \(\Theta j^{\mu}(x)\Theta^{-1}=-j^{\mu}(-x)\), the dagger being irrelevant since \(j^{\mu}\) is Hermitian.
For the charge, \(Q=\tfrac1c\int\dd^{3}x\,j^{0}(x)\), so \(\Theta Q\Theta^{-1}=-Q\) (the spatial integration measure is invariant under \(\vect{x}\to-\vect{x}\), and the time argument is a label). Then Equation (105.26) gives \(q_{\bar{a}}=\bra{\Theta a}Q\ket{\Theta a} =\bra{a}\Theta^{-1}Q\Theta\ket{a}^{*}=-q_{a}\).
For the magnetic moment, the operator is \(\hat{\vect{\mu}}=\tfrac12\int\dd^{3}x\, \vect{x}\times\vect{j}(x)\). Under \(\Theta\) the current picks up a minus sign and its spatial argument is reflected; substituting \(\vect{x}\to-\vect{x}\) in the integral, the two minus signs — one from \(\vect{x}\) and one from \(\vect{j}\) — cancel, so
Define \(\mu_{a}\) as the expectation of \(\hat{\mu}_{z}\) in the state of maximal spin projection along \(z\). Since \(\Theta\ket{a,s_{z}=+}=\ket{\bar{a},s_{z}=-}\) — \(\mathcal{T}\) and \(\mathcal{P}\) together reverse the spin, \(\mathcal{P}\) leaving it alone and \(\mathcal{T}\) reversing it — Equation (105.26) gives
and rotational invariance gives \(\bra{\bar{a},-}\hat{\mu}_{z}\ket{\bar{a},-}=-\mu_{\bar{a}}\). Hence \(\mu_{\bar{a}}=-\mu_{a}\).
∎Theorem 105.37 is exactly what the electron–positron comparison of Experiment: The Electron Anomalous Magnetic Moment tests, and it discharges the prediction recorded there at Phenomenon 112.13: the two \(g\) factors must agree in magnitude and the moments must point oppositely relative to the spin. They agree to about two parts in \(10^{12}\) [VanDyck:1987].
A particle and its antiparticle have equal total decay widths, and hence equal lifetimes — although their individual partial widths need not be equal. Rests on Theorem 105.34 and Equation (105.26).
Derives Theorem 105.38. Write \(S=\identity+\ii T\) with \(T\) the transition operator; unitarity \(S^{\dagger}S=\identity\) gives the optical relation
whose right-hand side is, up to kinematic factors, the total width \(\Gamma_{i}\). The left-hand side is \(2\,\mathrm{Im}\, T_{ii}\), since \(T_{ii}-T_{ii}^{*}=2\ii\,\mathrm{Im}\, T_{ii}\), and it is positive — a sum of squares — which fixes the sign once and for all. Now \(CPT\) invariance means \(\Theta S\Theta^{-1}=S\), and Equation (105.26) converts this into
an equality between the amplitude for \(i\to f\) and that for the \(CP\)-conjugated, motion-reversed process. Setting \(f=i\) and summing Equation (105.98) over all channels,
so the total width of the antiparticle equals that of the particle. The individual terms \(\abs{T_{fi}}^{2}\) in Equation (105.98) need not match their conjugates, because Equation (105.99) pairs \(i\to f\) with \(\bar{f}\to\bar{i}\) and not with \(\bar{i}\to\bar{f}\); unitarity forces the differences to cancel in the sum.
∎The last sentence is the whole reason direct \(CP\) violation is subtle, and it is worth extracting as a separate statement.
Let \(\mathcal{A}=a_{1}\ee^{\ii\delta_{1}}\ee^{\ii\phi_{1}} +a_{2}\ee^{\ii\delta_{2}}\ee^{\ii\phi_{2}}\) with \(a_{k}\) real, \(\phi_{k}\) \(CP\)-odd (weak) phases and \(\delta_{k}\) \(CP\)-even (rescattering) phases, so that the conjugate amplitude is \(\bar{\mathcal{A}}=a_{1}\ee^{\ii\delta_{1}}\ee^{-\ii\phi_{1}} +a_{2}\ee^{\ii\delta_{2}}\ee^{-\ii\phi_{2}}\). Then
A direct asymmetry therefore requires two interfering amplitudes, a difference of weak phases, and a difference of strong phases. Rests on Equation (105.28) and Theorem 105.38.
Derives Proposition 105.39. \(\abs{\mathcal{A}}^{2}=a_{1}^{2}+a_{2}^{2} +2a_{1}a_{2}\cos(\delta_{1}-\delta_{2}+\phi_{1}-\phi_{2})\), and \(\abs{\bar{\mathcal{A}}}^{2}\) is the same with \(\phi_{1}-\phi_{2}\to-(\phi_{1}-\phi_{2})\). Subtracting and using \(\cos(u+v)-\cos(u-v)=-2\sin u\sin v\) gives Equation (105.100).
∎This explains a great deal at once: why \(\epsilon'\) in Equation (105.72) is a difference of two phases; why the strong rescattering phases \(\delta_{0},\delta_{2}\) appear there at all; why the charm asymmetry of Equation (105.83) cannot be predicted, its strong phases being incalculable; and why an inclusive total rate can never show a \(CP\) asymmetry, since summing over final states washes out the strong phases and Theorem 105.38 takes over.
Consequences and loopholes
The companion theorem
Hypothesis (ii) of Theorem 105.34 assigns commutators to integer spin and anticommutators to half-integer spin. That assignment is not a convention: it is forced by the same hypotheses, and the result is the spin–statistics theorem [Pauli:1940]. The two theorems are so closely linked that Lüders and Zumino proved a converse — \(CPT\) invariance, plus the usual axioms, implies the correct connection between spin and statistics [Lueders:1958].
The essential computation is visible already for free fields. For a real scalar field the commutator at arbitrary separation is
the Pauli–Jordan function, in which \(k^{\mu}=(\omega_{k}/c,\vect{k})\) is on shell, \(\omega_{k}=c\sqrt{\vect{k}^{2}+m^{2}c^{2}/\hbar^{2}}\), so that \(k\cdot x=k_{\mu}x^{\mu}\) is a pure number and \(\Delta\) carries \(\mathrm{s}/\mathrm{m}^{3}\). The powers of \(\hbar\) and \(c\) in front are then forced: the canonical commutator \(\comm{\phi(t,\vect{x})}{\pi(t,\vect{y})} =\ii\hbar\,\delta^{3}(\vect{x}-\vect{y})\identity\) with \(\pi=c^{-2}\pp_{t}\phi\) makes \(\phi^{2}\) an energy per unit length, \(\mathrm{J}/\mathrm{m}\), and \(\hbar c^{2}\times \mathrm{s}/\mathrm{m}^{3}\) is exactly that. With \(\ii\hbar c\) in place of \(\ii\hbar c^{2}\) the two sides would differ by a factor of \(c\).
\(\Delta\) is Lorentz invariant and odd, so it vanishes for spacelike separation — a spacelike interval can be reversed by a proper orthochronous Lorentz transformation, and an odd invariant function of a reversible argument is zero. Had the same field been quantized with anticommutators, the object appearing would be the even combination \(\Delta_{1}\), which is nonzero everywhere outside the light cone: microcausality would fail. For the Dirac field the situation is exactly reversed: the anticommutator \(\acomm{\psi(x)}{\bar{\psi}(y)} \propto(\ii\hbar\gamma^{\mu}\pp_{\mu}+mc)\Delta(x-y)\) vanishes at spacelike separation because \(\Delta\) does, while the commutator would involve \(\Delta_{1}\) and would not. A second, independent strand of the argument is that a scalar field quantized with anticommutators has no ground state, its Hamiltonian being unbounded below, while a spinor field quantized with commutators produces negative-norm states. Either way the wrong choice destroys one of the hypotheses that Theorem 105.34 needs.
The spin–statistics theorem for arbitrary spin: the free-field computation sketched here settles spin 0 and spin 1/2, but the general statement requires the analyticity of the two-point Wightman function and the same extended-tube machinery as the axiomatic CPT proof. It belongs in Appendix A next to it, and the identical-particle chapter of Part IX states the result it needs.
The physical content is the exclusion principle of Identical Particles, and it is worth remembering how much rests on it: the stability of matter, the periodic table, the degeneracy pressure of Compact Stars and Relativistic Astrophysics. That a statement about analytic continuation in complexified Lorentz transformations underwrites the structure of the periodic table is one of the stranger facts in physics, and it is not a coincidence — both follow from locality and positivity of the energy.
Parametrizing a violation
If \(CPT\) were violated, how would one report the fact? The answer matters because experiments in wholly different systems — an antihydrogen atom, a kaon beam, a trapped antiproton — must be compared, and each measures something different. The standard device is an effective field theory in which every Lorentz- and \(CPT\)-violating operator that can be built from the observed fields is added to the Standard Model Lagrangian with an arbitrary coefficient [Colladay:1998]. A term such as \(b_{\mu}\bar{\psi}\gamma^{\mu}\gamma^{5}\psi\), with \(b_{\mu}\) a fixed background four-vector rather than a field, is \(CPT\)-odd by Lemma 105.35 — it has one uncontracted index, since \(b_{\mu}\) does not transform — and each experiment then bounds some combination of the coefficients. The tabulated bounds run to hundreds of entries [Kostelecky:2011].
It is important to be clear about the status of this construction. It is not a theory for which there is evidence; it is a bookkeeping scheme for null results, and its value is precisely that it lets a limit from one system be compared with a limit from another. The coefficients are all consistent with zero.
The relation to Lorentz invariance
In an interacting local quantum field theory satisfying the usual axioms, violation of \(CPT\) implies violation of Lorentz invariance. The converse is false: there are Lorentz-violating theories that preserve \(CPT\) [Greenberg:2002]. Rests on Theorem 105.34 and Lemma 105.35.
The asymmetry is easy to see in the parametrization above. An operator with an odd number of uncontracted indices, contracted with a fixed background vector, is both \(CPT\)-odd and Lorentz-violating; an operator with an even number, contracted with a fixed background tensor, is \(CPT\)-even but still Lorentz-violating. Hence \(CPT\)-violating \(\subset\) Lorentz-violating, strictly. Two consequences follow for experiment. A \(CPT\) test is automatically a Lorentz test, and in some channels the sharpest one available. And a search for Lorentz violation cannot be replaced by a \(CPT\) test, because the \(CPT\)-even sector is invisible to it.
Why these are tests of the axioms
The hypotheses of Theorem 105.34 are, essentially verbatim, the Wightman axioms of Axiomatic Quantum Field Theory. That is what gives a \(CPT\) measurement its unusual standing. A discrepancy between the electron's and the positron's \(g\) factor would not be evidence for some new particle or some unmeasured coupling; it would mean that one of locality, Lorentz invariance, or the positivity of the energy spectrum is false. There are very few measurements in physics whose failure would be that structural, and it is the reason the null results below are worth the effort they cost.
Experimental tests
| system | quantity compared | fractional limit | source |
|---|---|---|---|
| $K^{0}$ and $\bar{K}^{0}$ | rest energy | $6\times10^{-19}$ | [Navas:2024] |
| antihydrogen and hydrogen | $1S$–$2S$ transition frequency | $2\times10^{-12}$ | [Ahmadi:2018] |
| $\bar{p}$ and $p$ | magnetic moment | $1.5\times10^{-9}$ | [Smorra:2017] |
| $\bar{p}$ and $p$ | charge-to-mass ratio | $1.6\times10^{-11}$ | [Borchert:2022] |
| $e^{+}$ and $e^{-}$ | $g$ factor | $2\times10^{-12}$ | [VanDyck:1987] |
The kaon bound, and how to read it
The neutral kaon entry is the tightest number in Table 105.2 and the one most often misquoted. What is measured is a combination of the semileptonic charge asymmetries of \(K_{L}\) and \(K_{S}\) together with the phase of \(\eta_{+-}\), which bounds the \(CPT\)-violating diagonal difference \(M_{11}-M_{22}\) of Proposition 105.23. The bound on that energy is of order \(3\times 10^{-10}\,\mathrm{eV}\) — roughly \(10^{-4}\) of the \(K_{L}\)–\(K_{S}\) mass difference Equation (105.60), which is itself a \(3.5\,\mu\mathrm{eV}\) splitting. Dividing by \(m_{K}c^{2}=497.6\,\mathrm{MeV}\) turns it into \(6\times10^{-19}\). The exponent therefore reflects the largeness of the denominator as much as the sharpness of the measurement, and comparing it with the \(10^{-12}\) of an atomic experiment is comparing two different things. Both statements are true; neither is “the best \(CPT\) test” without a choice of what to divide by.
Antihydrogen
The ALPHA collaboration at CERN traps antihydrogen, formed by mixing antiprotons from the Antiproton Decelerator with positrons, in a magnetic minimum, and performs laser spectroscopy on the trapped sample. The two-photon \(1S\)–\(2S\) transition, a natural linewidth of about \(1\,\mathrm{Hz}\) on a frequency of \(2.466\times 10^{15}\,\mathrm{Hz}\), is the sharpest observable in hydrogen. ALPHA first observed the transition in antihydrogen [Ahmadi:2017] and then measured its frequency, in zero field, in agreement with hydrogen at the level of \(2\times10^{-12}\) [Ahmadi:2018]. The comparison tests Theorem 105.36 and Theorem 105.37 jointly, since the transition frequency depends on the electron and proton masses and charges in combination.
The antiproton
Two BASE measurements bound the antiproton's properties directly. The magnetic moment, measured on a single trapped antiproton by the continuous Stern–Gerlach effect, is \(\mu_{\bar{p}}=-2.7928473441(42)\mu_{N}\), agreeing with the proton's in magnitude to \(1.5\) parts in \(10^{9}\) and opposite in sign [Smorra:2017] — exactly Theorem 105.37. The charge-to-mass ratio, measured by comparing the cyclotron frequencies of an antiproton and a negative hydrogen ion in the same trap, agrees with the proton's to \(16\) parts in \(10^{12}\) [Borchert:2022].
What $CPT$ does not say
The ALPHA-g experiment released antihydrogen from a vertical trap and found that it falls, with a measured acceleration \(a_{g}=(0.75\pm0.13\pm0.16)\,g\) downward [Anderson:2023]. This is a fine measurement and it is not a test of \(CPT\). The theorem relates inertial masses, charges and moments; it says nothing about gravitational coupling, and a hypothetical antigravity would violate the equivalence principle of The Equivalence Principle and Classical Tests without disturbing any hypothesis of Theorem 105.34. Reporting it as a \(CPT\) test, as several popular accounts did, confuses two independent structures. Every genuine \(CPT\) test performed to date is consistent with exact invariance.
The baryon asymmetry
Sakharov's conditions
The universe contains matter and essentially no antimatter. If it began in a state symmetric under \(C\) and \(CP\), with equal numbers of baryons and antibaryons, then getting from there to here requires three things, identified by Sakharov in a two-page paper of 1967 [Sakharov:1967]. All three are necessary, and each necessity is a theorem.
Let a system begin with \(\avg{B}=0\), where \(B\) is baryon number. Then \(\avg{B}=0\) at all later times if any one of the following holds:
-
\(B\) is conserved, \(\comm{B}{H}=0\);
-
\(C\) and \(CP\) are symmetries of \(H\) and of the initial state;
-
the system is in thermal equilibrium.
Generating a baryon asymmetry therefore requires baryon number violation, \(C\) and \(CP\) violation, and a departure from thermal equilibrium. Rests on Theorem 105.34 and Equation (105.34).
Derives Theorem 105.41. (i) is immediate: a conserved quantity that starts at zero stays there.
(ii) Let \(\rho(t)\) be the density operator and suppose \(\mathcal{C}H\mathcal{C}^{-1}=H\) and \(\mathcal{C}\rho(0)\mathcal{C}^{-1}=\rho(0)\), whence \(\mathcal{C}\rho(t)\mathcal{C}^{-1}=\rho(t)\) for all \(t\) since the evolution commutes with \(\mathcal{C}\). Baryon number is \(C\)-odd, \(\mathcal{C}B\mathcal{C}^{-1}=-B\). Then, using cyclicity of the trace and the unitarity of \(\mathcal{C}\),
The same computation with \(\mathcal{CP}\) in place of \(\mathcal{C}\) gives the same conclusion, and both are needed: a theory may violate \(C\) maximally and still conserve \(CP\) — the \(V-A\) interaction of Equation (105.34) is the standard example — in which case the \(CP\)-version of the argument still forces \(\avg{B}=0\). It is the pair of symmetries that must fail.
(iii) is the deepest of the three, and it is where \(CPT\) enters. In thermal equilibrium at temperature \(T\) the density operator is \(\rho=\ee^{-H/k_{B}T}/Z\); there is no chemical potential for \(B\), because by hypothesis (i) has failed and \(B\) is not conserved, so it cannot label a conserved charge. Now \(\avg{B}=\tr(\rho B)\) is real, being the trace of a product of two Hermitian operators. For an antiunitary \(\Theta\) and any operator \(O\), choosing the orthonormal basis \(\set{\Theta e_{n}}\) to evaluate the trace gives
Apply this with \(O=\ee^{-H/k_{B}T}B\). By Theorem 105.34 \(\Theta H\Theta^{-1}=H\), and \(B\) is \(CPT\)-odd (it is \(C\)-odd and \(P\)- and \(T\)-even), so \(\Theta\left(\ee^{-H/k_{B}T}B\right)\Theta^{-1} =-\ee^{-H/k_{B}T}B\). Hence
the last step because the quantity is real; therefore it vanishes.
∎Part (iii) deserves a comment, because it is the point at which the two halves of this chapter meet. The reason equilibrium forbids a baryon asymmetry is that particles and antiparticles have the same mass, which is Theorem 105.36: with equal masses and no conserved charge to bias them, the equilibrium populations are equal whatever the interactions. The \(CPT\) theorem, which guarantees the symmetry of the world, is thereby also the obstacle to explaining its asymmetry. Any mechanism must arrange for the universe to be out of equilibrium while the baryon-number-violating processes are active.
The Standard Model's baryon-number violation
Condition (i) looks hopeless at first sight, since baryon number is an exact symmetry of the classical Standard Model Lagrangian. It is not a symmetry of the quantum theory. The baryon and lepton currents are anomalous: the divergence that vanishes classically acquires a contribution from the non-invariance of the fermion measure [Adler:1969] [Bell:1969] [tHooft:1976a], with the integrated statement
where \(n_{f}=3\) is the number of generations and \(n_{\mathrm{CS}}\in\Z\) is the Chern–Simons number of the \(\SU(2)\) gauge field, an integer for a transition between distinct vacua. Baryon number is violated; \(B-L\) is not. The barrier between adjacent vacua is the sphaleron, a static unstable solution whose energy is set by the \(\SU(2)_{L}\) coupling \(g\) and the \(W\) mass,
Here \(g^{2}=4\pi\alpha/\sin^{2}\theta_{W}=0.397\), taking \(\sin^{2}\theta_{W}=0.231\) and \(\alpha=1/137.036\), so that \(g=0.630\) and \(8\pi/g^{2}=63.4\); with \(m_{W}c^{2}=80.4\,\mathrm{GeV}\) the bare combination \((8\pi/g^{2})m_{W}c^{2}\) is only \(5.1\,\mathrm{TeV}\). The factor \(B\) is a dimensionless shape factor fixed numerically by the field equations: it depends on nothing but the ratio of the Higgs self-coupling to \(g^{2}\), runs from about \(1.5\) to \(2.7\) across the whole range of that ratio, and is close to \(1.8\) at the observed Higgs mass. That is where the \(9\,\mathrm{TeV}\) comes from, and quoting the formula without \(B\) — as is often done — understates the barrier by nearly a factor of two. At zero temperature the tunnelling rate then carries the factor \(\exp(-4\pi/\alpha_{W})\sim10^{-173}\), with \(\alpha_{W}=g^{2}/4\pi=0.0316\), and nothing happens. Above the electroweak scale, \(k_{B}T\gtrsim100\,\mathrm{GeV}\), the barrier can be crossed thermally and the rate per unit volume is unsuppressed, \(\Gamma/V\sim(\alpha_{W}k_{B}T)^{4}/(\hbar c)^{3}\hbar\) [Kuzmin:1985]. The Standard Model therefore satisfies Sakharov's first condition, in the early universe and nowhere else.
The triangle anomaly and the Chern–Simons current: the derivation that the baryon current's divergence is proportional to the topological density, whether by the triangle diagram or by the non-invariance of the path-integral measure, is not carried out in this chapter. It belongs to the path-integral chapter of this part, which reserves the anomaly, and to Appendix A. What is used here is only the integrated statement, that a change of Chern-Simons number by one unit changes baryon number by the number of generations.
Why the Standard Model fails
The number to be explained
The asymmetry is quantified by the baryon-to-photon ratio.
using \(\Omega_{b}h^{2}=0.02237\pm0.00015\) from Planck [Aghanim:2020]. Rests on Equation (68.5) and Phenomenon 51.3.
Derivation. Derives Proposition 105.42. The photon number density of a blackbody at temperature \(T\) is
at \(T_{0}=2.7255\,\mathrm{K}\). The baryon number density is \(\Omega_{b}\rho_{c}/\bar{m}\) with \(\rho_{c}=3H_{0}^{2}/(8\pi G) =1.87834\times 10^{-26}\,\mathrm{kg}/\mathrm{m}^{3}\,h^{2}\) and \(\bar{m}\) the mean mass per baryon, close to the proton mass \(1.67262\times 10^{-27}\,\mathrm{kg}\); hence \(n_{B}=11.23\,/\mathrm{m}^{3}\,\Omega_{b}h^{2}\). The ratio is \(11.23/4.107\times10^{8}=2.74\times10^{-8}\) per unit of \(\Omega_{b}h^{2}\).
∎The same number follows independently from the light-element abundances produced in primordial nucleosynthesis, which depend on \(\eta\) through the neutron-to-proton ratio at freeze-out; the two determinations agree, which is what makes Equation (105.105) a measurement rather than a fit. No antimatter region has been found: a matter–antimatter boundary anywhere within the observable universe would radiate annihilation gamma rays at a level long since excluded, and the antinuclei searched for in the cosmic radiation (Cosmic Rays and Astroparticle Physics) are consistent with secondary production. See Phenomenon 111.17.
The CKM phase is far too small
By Proposition 105.27, the strength of Standard Model \(CP\) violation is measured by the invariant Equation (105.79). Made dimensionless at the temperature \(k_{B}T\approx100\,\mathrm{GeV}\) where sphalerons are active, it is
Inserting the running quark masses at that scale — \(m_{t}c^{2} \approx170\,\mathrm{GeV}\), \(m_{b}c^{2}\approx2.9\,\mathrm{GeV}\), \(m_{c}c^{2}\approx0.6\,\mathrm{GeV}\), \(m_{s}c^{2}\approx50\,\mathrm{MeV}\), \(m_{d}c^{2}\approx3\,\mathrm{MeV}\), \(m_{u}c^{2}\approx1.5\,\mathrm{MeV}\) — and Equation (105.78) gives
The observed \(\eta\) is \(6\times10^{-10}\). Detailed calculations of what the Standard Model would actually produce, which involve more than this crude ratio, arrive at asymmetries of order \(10^{-20}\) or smaller [Gavela:1994]: the shortfall is at least eight orders of magnitude and on the standard estimates ten or more. The reason is structural and worth stating: \(CP\) violation in the Standard Model requires all six quark masses to be distinct (Proposition 105.27), and at a temperature of \(100\,\mathrm{GeV}\) five of the six are negligible compared with the temperature, so the effect is suppressed by twelve powers of small mass ratios. The phase itself is of order unity; it is the masses that kill it.
And the third condition fails too
Even a larger phase would not help, because the Standard Model does not depart from thermal equilibrium at the electroweak scale. A departure would require a first-order phase transition, proceeding by the nucleation of bubbles of the broken phase whose walls sweep through the plasma; and it would have to be strongly first order, since after the transition the sphaleron processes must switch off fast enough not to erase the asymmetry just created. The washout condition is
in the customary units, \(v\) being the Higgs expectation value at the critical temperature — a requirement that the broken-phase sphaleron energy Equation (105.104) exceed the thermal energy by enough that the Boltzmann factor quenches the rate.
Lattice studies settled the question. The electroweak transition is first order only for a light Higgs; the line of first-order transitions ends at a critical endpoint near \(m_{H}c^{2}\approx72\text{–}80\,\mathrm{GeV}\), and beyond it the change from the symmetric to the broken phase is a smooth crossover with no bubbles, no walls and no departure from equilibrium [Kajantie:1996]. The measured Higgs mass is \(m_{H}c^{2}=125.20 \pm 0.11\,\mathrm{GeV}\) [Navas:2024], far above the endpoint (Electroweak Unification and the Higgs Boson). Two of Sakharov's three conditions therefore fail in the Standard Model, one by ten orders of magnitude and one outright.
This is recorded as an open problem, not a solved one: see Section 133.5 in What We Observe but Do Not Understand. The measurement — Equation (105.105) — is secure; the mechanism is unknown. It is worth being precise about what is and is not established, because the literature contains many proposals and no evidence for any of them. What is established: baryon number is not conserved in the Standard Model at high temperature; \(CP\) is violated by the CKM phase; the observed asymmetry exceeds what those two can produce by a very large factor; and the electroweak transition is a crossover. What is not established: anything about where the missing asymmetry came from.
The strong CP problem
The theta term
Section 105.1.1 showed that the quark–gluon coupling, being a vector current, conserves \(P\), \(C\) and \(T\). That argument is incomplete, because it enumerated only the terms one would think to write. There is one more term allowed by every symmetry that QCD is supposed to have — gauge invariance, Lorentz invariance, locality and renormalizability — and it violates \(P\) and \(T\).
Let \(q(x)\) denote the topological charge density of the gluon field, normalized so that
for any finite-action field configuration, \(n\) being its winding number; \(q\) therefore carries the SI dimension \(/\mathrm{m}^{3}/\mathrm{s}\). The QCD Lagrangian density may contain
with \(\theta\) a dimensionless number. Its contribution to the action is \(S_{\theta}=\theta\hbar n\), so it enters the path-integral weight \(\exp(\ii S/\hbar)\) of Path-Integral Quantization only through the phase \(\ee^{\ii\theta n}\), and \(\theta\) is periodic with period \(2\pi\).
In terms of the gluon field strengths, \(q\) is proportional to \(\sum_{a}\vect{E}^{a}\cdot\vect{B}^{a}\), the colour analogue of the electromagnetic invariant \(\vect{E}\cdot\vect{B}\); the constant of proportionality is fixed by Equation (105.109) and involves \(g_{s}^{2}\) and \(\hbar\) but nothing else. That representation is all that is needed for the symmetry properties.
\(\Lag_{\theta}\) is odd under \(P\) and under \(T\), even under \(C\), and therefore odd under \(CP\) and even under \(CPT\). Rests on Equations (105.6), (105.17) and (105.30).
Derives Proposition 105.44. Use the abelian analogue, which carries the same index structure. By Equation (105.6), \(\vect{E}\) is a polar vector and \(\vect{B}\) an axial vector, so \(\vect{E}\cdot\vect{B}\) is a pseudoscalar: \(P\)-odd. By Equation (105.30), \(\vect{E}\) is \(T\)-even and \(\vect{B}\) is \(T\)-odd, so the product is \(T\)-odd. Under \(C\), both fields reverse (Equation (105.17)), so the product is even. Hence \(CP\)-odd and, as Theorem 105.34 requires of any term in a Hermitian local Lagrangian, \(CPT\)-even. The same signs hold for the non-abelian \(\sum_{a}\vect{E}^{a}\cdot\vect{B}^{a}\), since the colour index is inert under all three operations.
∎Two facts then make the term a problem rather than a curiosity.
First, \(q\) is a total derivative: \(q=\pp_{\mu}K^{\mu}\) for the Chern–Simons current \(K^{\mu}\). A total derivative contributes nothing in perturbation theory, which is why \(\theta\) never appears in a Feynman diagram; but \(K^{\mu}\) is not gauge invariant, and for field configurations with nontrivial topology at infinity — the instantons of [tHooft:1976a] [Callan:1976] [Jackiw:1976] — the surface term is the nonzero integer of Equation (105.109). The term is therefore invisible perturbatively and physical nonperturbatively, which is the worst combination for anyone hoping to argue it away.
Second, \(\theta\) is not by itself observable. A chiral rotation of the quark fields, \(\psi_{j}\to\ee^{\ii\alpha\gamma^{5}}\psi_{j}\), changes the phase of the quark mass matrix and, because the fermion measure is not invariant under it [Adler:1969] [Bell:1969], shifts \(\theta\) by a compensating amount. Only the combination
is invariant, and it is \(\bar{\theta}\) that is measurable. The strong and Yukawa sectors are thereby tied together: a quantity of the gluon sector can be cancelled against a quantity of the Higgs sector, and no symmetry of the Standard Model relates them.
There is exactly one way the problem could evaporate. If any quark were massless, \(\det M_{q}=0\), \(\arg\det M_{q}\) would be undefined, and the chiral rotation could be used to set \(\bar{\theta}=0\) with no cost. That escape is closed by measurement: lattice determinations give \(m_{u}c^{2}=2.16 \pm 0.07\,\mathrm{MeV}\) at the customary renormalization point [Navas:2024], tens of standard deviations from zero.
A nonzero \(\bar{\theta}\) generates a neutron electric dipole moment
equivalently \((1\)–\(4)\times10^{-16}\,\bar{\theta}\;e\, \mathrm{cm}\) [Baluni:1979] [Crewther:1979]. Rests on Proposition 105.44, Equation (105.111) and Theorem 105.30.
Derivation. Derives Proposition 105.45. The parametric form follows from dimensional analysis plus chiral counting, and is worth doing because it explains the size. A neutron EDM requires the violation of \(P\) and \(T\), which \(\bar{\theta}\) supplies; it requires a charge, which supplies \(e\); and it requires a length, for which the only candidates are the nucleon Compton wavelength \(\hbar/(m_{N}c)=0.21\,\mathrm{fm}\) and the pion's. Chiral symmetry supplies a further suppression: in the limit of vanishing quark masses the \(\theta\) term is unobservable, so the answer must be proportional to the reduced quark mass \(m_{*}=m_{u}m_{d}/(m_{u}+m_{d})\approx1.5\,\mathrm{MeV}/c^{2}\) measured against the nucleon mass. Hence
that is \(3\times10^{-17}\,\bar{\theta}\;e\,\mathrm{cm}\). The chiral-loop calculations of Baluni [Baluni:1979] and of Crewther, Di Vecchia, Veneziano and Witten [Crewther:1979] supply the coefficient, and it is between about three and thirteen times larger than this estimate — \(1.6\times 10^{-37}\,\mathrm{C}\,\mathrm{m}/5\times 10^{-38}\,\mathrm{C}\,\mathrm{m}=3.2\) at one end and \(12.8\) at the other — because the dominant contribution is a chiral logarithm from the pion cloud rather than a contact term; modern lattice determinations also fall inside Equation (105.112). That spread — a factor of a few in the coefficient — is the dominant uncertainty in the bound on \(\bar{\theta}\) below, and it is why the bound is quoted as an order of magnitude and not as a number.
∎The experimental bound and its status
Combining Equation (105.93) — the ultracold-neutron measurement at the Paul Scherrer Institute, whose Ramsey technique is described at Equation (105.91) — with Equation (105.112) gives
so \(\abs{\bar{\theta}}\lesssim10^{-10}\), with the factor-of-a-few uncertainty inherited from Proposition 105.45.
This is the strong \(CP\) problem, and it is worth stating exactly what kind of problem it is. \(\bar{\theta}\) is a free dimensionless parameter of the Standard Model. It is periodic with period \(2\pi\), so there is no sense in which small values are “near a boundary” of its range; a value of order unity is the natural expectation for a number about which nothing else is known. It is measured to be smaller than \(10^{-10}\). No symmetry of the Standard Model requires it to vanish — the one candidate, a massless quark, is excluded by measurement — and no dynamical mechanism within the Standard Model relaxes it. The problem is therefore an observation: a parameter that could have been anything is measured to be extraordinarily small, and nothing in the accepted theory explains why. It is recorded as such in Section 133.1.
The Peccei–Quinn proposal, and its status
The best-known response is to make \(\bar{\theta}\) dynamical. Peccei and Quinn observed that if the Standard Model is extended by a global \(\U(1)\) symmetry that is spontaneously broken and itself anomalous, then \(\bar{\theta}\) is replaced by a field whose potential — generated by the same instanton effects that make \(\theta\) physical — is minimized at zero [Peccei:1977]. The spontaneously broken symmetry implies a light pseudo-Goldstone boson, the axion, as Weinberg [Weinberg:1978] and Wilczek [Wilczek:1978] pointed out at once. Its mass is tied to the symmetry-breaking scale \(f_{a}\) by
so that a very weakly coupled axion is a very light one.
The chapter states the position plainly. No axion has been observed. The haloscope searches, of which ADMX is the most sensitive, look for the resonant conversion of a galactic-halo axion into a microwave photon in a strong magnetic field; the reported results are exclusions, covering a band of couplings in a narrow mass window around \(2.66\text{–}2.81\,\mu\mathrm{eV}\) [Braine:2020]. Helioscopes and light-shining-through-walls experiments likewise report null results. The Peccei–Quinn mechanism is a hypothesis with an attractive structure and no experimental support, and Section 133.7 lists it accordingly. What this chapter records as physics is the measurement, Equation (105.113): a dimensionless parameter of the Standard Model, free to be of order unity, is smaller than \(10^{-10}\), and no one knows why.